Infectious disease prediction method and system based on consistency constraint and zero inflation perception

CN122734720APending Publication Date: 2026-09-11BEIJING CENT FOR DISEASE PREVENTION & CONTROL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611148650.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-30
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

然而,上述方法均以市、省级数据为主要应用场景,当预测粒度细化至街道级别时,面临严峻挑战

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122734720A_ABST
    Figure CN122734720A_ABST
Patent Text Reader

Abstract

This invention provides a method and system for infectious disease prediction based on consistency constraints and zero-inflation perception, belonging to the field of infectious disease prediction technology. Addressing the challenges of extremely sparse and zero-inflation-distributed street-level infectious disease incidence data, gradient submersion due to significant differences in magnitude across levels, and naturally heterogeneous covariates at each level, this invention constructs independent long short-term memory network encoders for the city, district, and street levels. The street-level encoder output has a zero-inflation-gated output structure, decoupling the prediction into two parallel subtasks: zero-value probability judgment and positive-value intensity estimation. After concatenating the predicted values ​​from each level into a unified prediction vector, a joint loss including the prediction error loss of each level and the soft consistency regularization loss is calculated. All network parameters are then updated synchronously through backpropagation. This invention achieves end-to-end collaborative optimization of fine-grained hierarchical prediction accuracy and cross-level summative consistency, and can directly output logically consistent three-level prediction results for the city, district, and street levels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of infectious disease prediction and processing technology, specifically to an infectious disease prediction method and system based on consistency constraints and zero-inflation perception. Background Technology

[0002] Hierarchical Time Series Forecasting (HTSF) is a class of forecasting methods specifically designed for data with hierarchical aggregation structures. Its goal is to simultaneously produce forecasts at multiple granular levels that satisfy additive consistency (e.g., the city-wide forecast equals the sum of all district forecasts). Existing HTSF methods are mainly divided into two categories: one is a two-stage strategy of "predict first, then coordinate," including bottom-up, top-down, and optimal coordination methods, where base forecasts are generated independently, and the coordination step only corrects the final output; the other is the end-to-end deep hierarchical model that has emerged in recent years, directly embedding the hierarchical structure into neural networks for joint forecasting. However, the above methods are primarily applied to city and provincial data, and face significant challenges when the forecast granularity is refined to the street level.

[0003] The main problems with existing technologies are as follows: First, street-level incidence data exhibits extreme sparsity and zero inflation, and conventional assumptions of continuous or count distribution cannot accurately characterize its dual-process generation mechanism, leading to the failure of basic prediction models. Second, the daily incidence rate at the city level can reach hundreds of cases, while at the street level it is mostly zero or single digits. If a global model with shared parameters is used for joint modeling, the scale-sensitive loss function will cause the city-level sequence to dominate gradient updates, and the weak gradient signals of the sparse street-level sequence will be "submerged" by the gradient, making it almost impossible to learn the pattern effectively. Third, the city level can obtain indicators such as airport passenger flow and macro policies, while the street level relies on fine features such as population density. The covariates at different levels are naturally heterogeneous. Existing methods either force alignment to a unified dimension, losing level-specific information, or abandon cross-level information borrowing, thus losing macro-level situation constraints.

[0004] Therefore, there is an urgent need for a technical solution that can adapt to the distribution characteristics of extremely fine-grained data and simultaneously optimize the prediction accuracy and consistency of results at each level from end to end. Summary of the Invention

[0005] To address the technical problems in related technologies, this invention provides a method and system for predicting infectious diseases based on consistency constraints and zero-inflation perception.

[0006] To achieve the above objectives, the technical solution adopted by the present invention includes:

[0007] According to a first aspect of the present invention, an infectious disease prediction method based on consistency constraints and zero-inflation perception is provided, comprising the following steps:

[0008] Obtain historical incidence rate sequence data for multiple administrative levels and covariate data for each administrative level;

[0009] Independent long short-term memory network encoders were constructed at the city, district, and street levels, respectively. Covariate data and historical incidence sequence data at each level were input into the corresponding level of the long short-term memory network encoder to extract the temporal features at each level.

[0010] For the city and district levels, the predicted values ​​are generated by linear mapping based on the temporal characteristics of the corresponding levels. For the street level, the predicted values ​​are generated by a zero-inflation gating output structure based on the temporal characteristics of the street level. The zero-inflation gating output structure calculates in parallel the gating probability vector representing the probability of non-zero cases in each street and the intensity vector representing the expected number of non-zero cases. The street-level predicted values ​​are obtained by multiplying the gating probability vector and the intensity vector element by element.

[0011] The city-level, district-level, and street-level predicted values ​​are concatenated into a unified prediction vector.

[0012] Based on the unified prediction vector and the true value vector, the joint loss is calculated. The joint loss includes city-level prediction error loss, district-level prediction error loss, street-level prediction error loss, and soft consistency regularization loss used to penalize prediction values ​​that deviate from the hierarchical summation consistency constraint.

[0013] Based on the joint loss, the parameters of the long short-term memory network encoder and the zero-inflation gated output structure at each level are synchronously updated through backpropagation until training is complete;

[0014] The input data at each level at the time to be predicted are input into the trained long short-term memory network encoders and zero-inflation gating output structures at each level, and the prediction results at the city, district, and street levels are output simultaneously.

[0015] Optionally, in the step of obtaining historical incidence rate sequence data at multiple administrative levels and covariate data for each administrative level:

[0016] Set time The city-level observation value The district-level observation vector is The street-level observation vector is ,in, For the set of natural numbers, , The total number of district-level units, The total number of street-level units;

[0017] For the street Its covariate vector is denoted as ,in, Dimensions for street-level features;

[0018] City-level covariates are denoted as District-level covariates are denoted as ,in, and These represent the dimensions of city-level and district-level characteristics, respectively.

[0019] Optionally, the step of constructing independent long short-term memory network encoders at the city, district, and street levels, and inputting the covariate data and historical incidence sequence data at each level into the corresponding level of the long short-term memory network encoder:

[0020] For the street , covariate vector Compared with observed values The street is pieced together in time input vector ;

[0021] All The streets in the time window to The input vectors are stacked according to time steps to obtain the street-level prediction sub-model at time step [time]. input matrix ,in, This refers to the length of the historical observation window;

[0022] City-level input matrix and zone-level input matrix The covariates and observed values ​​at the city and district levels were constructed in the same manner.

[0023] Optionally, the long short-term memory network encoder at each level includes the following components:

[0024] The first layer of the Long Short-Term Memory network has a hidden state dimension of 64 and returns the complete time series.

[0025] The second layer is a Long Short-Term Memory network with a hidden state dimension of 32, which only returns the output of the last time step.

[0026] Dropout layer with a dropout rate of 0.2;

[0027] The first fully connected layer with an output dimension of 128 and using the ReLU activation function;

[0028] The second fully connected layer has an output dimension of 64 and uses the ReLU activation function.

[0029] Let the output vector of the second fully connected layer be... .

[0030] Optionally, in the step of generating city-level and district-level predicted values ​​through linear mapping based on the time-series characteristics of the corresponding levels:

[0031] City-level forecast values ​​are obtained through formulas Calculate, where, As a city-level weight, This is a city-level biased item;

[0032] District-level predicted values ​​are obtained through the formula Calculate, where, This is a district-level weight matrix. This is the zone-level bias vector.

[0033] Optionally, in the step of parallel computation of the gated probability vector and intensity vector in the zero-inflation gated output structure:

[0034] Gated probability vector Through formula Calculate, where, The sigmoid function maps the gate probability to the interval (0,1). and These are the weights and biases of the gated branches, respectively;

[0035] Intensity vector Through formula Calculate, where, To ensure that the output intensity is non-negative, and These represent the weights and biases of the strength branches, respectively.

[0036] Street-level predicted values ​​are obtained through a formula Calculate, where, This represents the Hadamard product.

[0037] Optionally, the soft consistency regularization loss is constructed as follows:

[0038] Let the aggregation matrix be... Its first row contains all 1s, indicating that the total number of city-level values ​​equals the sum of the values ​​of all street-level values; for the... Each district , matrix number The elements of the row satisfy If and only if the street If it belongs to this district, then it is 0;

[0039] Define constraint matrix ,in, for An identity matrix of order 1;

[0040] Based on the constraint matrix Define the orthogonal projection matrix ;

[0041] The soft consistency regularization loss Including the first and second items:

[0042] The first item is the consistency deviation penalty, calculated as follows: This is used to penalize the degree to which the predicted value deviates from the hierarchical summation consistency constraint, where... This represents the sample size for a single training batch.

[0043] The second item is the projection prediction error, which is calculated as follows: This is used to ensure that the model does not sacrifice prediction accuracy while satisfying the trend of consistency constraints.

[0044] Optionally, the street-level prediction error loss Including the first, second, and third sub-terms of the weighted sum:

[0045] The first sub-item is the principal loss, which uses the negative log-likelihood of the Tweedie distribution, and is expressed as: ,in, power parameter The range of values ​​is ;

[0046] The second sub-item is the gated loss, which uses binary cross-entropy loss, labeled by whether the true value is zero. The calculation method is as follows: ,in, For the first The gating probability vector corresponding to each sample This is an indicator function; it takes the value 1 when the true value is greater than zero, and 0 otherwise. This is the gating loss weighting coefficient;

[0047] The third sub-item is the non-zero sample amplitude loss, which calculates the mean squared error only for sample points where the true value is non-zero. The calculation method is as follows: ,in, For the first In the nth sample The intensity prediction value corresponding to each street To prevent division by zero constant, The amplitude loss weights are for non-zero samples.

[0048] Optionally, the joint loss Including city-level forecast errors District-level prediction error Street-level prediction error and soft consistency regularization loss :

[0049]

[0050] Among them, the city-level forecast error The mean square error is used, and the calculation method is as follows: ;

[0051] District-level prediction error The mean square error is used, and the calculation method is as follows: .

[0052] According to a second aspect of the present invention, an infectious disease prediction system based on consistency constraints and zero-inflation sensing is also provided, for executing the infectious disease prediction method based on consistency constraints and zero-inflation sensing described in any of the technical solutions of the first aspect of the present invention, the system comprising:

[0053] The city-level long short-term memory network prediction module is used to receive city-level input data and generate city-level prediction values;

[0054] The district-level long short-term memory network prediction module is used to receive district-level input data and generate district-level predicted values;

[0055] A street-level long short-term memory network prediction module, comprising a zero-inflation gated output submodule, wherein the zero-inflation gated output submodule is used to compute in parallel a gated probability vector representing the probability of non-zero cases occurring in each street and an intensity vector representing the expected number of non-zero cases, and multiplies the gated probability vector and the intensity vector element by element to obtain the street-level prediction value;

[0056] The hierarchical consistency constraint module is used to transform hierarchical summation constraints into soft consistency regularization loss;

[0057] The joint training module is used to combine the city-level prediction error loss, district-level prediction error loss, street-level prediction error loss and the soft consistency regularization loss into a total loss, and synchronously update the parameters of all modules through backpropagation.

[0058] The prediction output module is used to input the input data at each level at the time to be predicted into the trained modules and simultaneously output the prediction results at the city, district, and street levels.

[0059] Beneficial effects:

[0060] 1. First, this invention can effectively improve the accuracy of fine-grained hierarchical prediction. Specifically, this invention addresses the characteristics of extremely sparse and zero-inflation distribution of street-level incidence data by designing a dual-branch gating output structure in the street-level prediction sub-model. This explicitly decouples the prediction process into two parallel sub-tasks: "zero-value probability judgment" and "positive-value intensity estimation," making the model naturally adaptable to the generation mechanism of zero-inflation count data.

[0061] Second, this invention avoids gradient submersion across different levels, ensuring effective learning of sparse sequence features. Specifically, this invention constructs LSTM encoders with independent parameters for city, district, and street levels, fundamentally blocking the gradient submersion path caused by significant differences in data volume, allowing the weak gradient signals of street-level sparse sequences to be effectively propagated back within their independent encoders. Experimental results show that the multi-task global model LSTM achieves a MASE of 0.95 at the street level, close to the naive benchmark. Furthermore, this invention, through its strategy of hierarchical independent modeling followed by joint training, achieves a MASE of 0.45 at the street level while maintaining leading prediction accuracy at the city and district levels, verifying the effectiveness of this design in balancing cross-level learning effects.

[0062] Third, this invention is compatible with multi-level heterogeneous covariates, avoiding information loss and noise introduction. Specifically, the encoders at each level of this invention are independent and do not share parameters. Each level of the network can use its own set of exclusive covariates that are actually available at that level (e.g., macro indicators such as airport passenger flow and influenza positivity rate at the city level, and fine features such as population density at the street level), without the need for cross-level feature dimension alignment or down-filling. Experiments have shown that forcibly down-filling city-level covariates to the street level results in a baseline model prediction accuracy that is even lower than the seasonal naive baseline; however, this invention does not require such operations and can effectively transfer cross-level information through joint training while preserving the integrity of the original features at each level.

[0063] Fourth, this invention enables end-to-end joint optimization, overcoming the fragmented training defects of two-stage methods. Specifically, this invention transforms hierarchical aggregation constraints into soft consistency regularization loss, which, together with the prediction error loss of each level, constitutes a joint loss function. The parameters of all network levels are updated synchronously through a single backpropagation. Compared to the traditional two-stage method of "predict first, then coordinate," the training process of this invention ensures that the model parameters of each level jointly serve the global objective of "optimal accuracy at each level while maintaining overall consistency," achieving simultaneous optimization of prediction accuracy and consistency constraints.

[0064] Fifth, this invention can output logically consistent multi-granularity prediction results, directly supporting hierarchical collaborative decision-making. Specifically, after end-to-end training with soft consistency regularization constraints, the city-level, district-level, and street-level prediction results output by this invention naturally tend to satisfy an additive consistency relationship. They can be directly used by disease control departments for material allocation, resource deployment, and risk assessment according to administrative levels without any post-event adjustments, providing logically consistent and readily available quantitative evidence for hierarchical collaborative decision-making.

[0065] 2. Other beneficial effects or advantages of the present invention will be described in detail in the specific embodiments. Attached Figure Description

[0066] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0067] in:

[0068] Figure 1 This is an exemplary embodiment of the present invention, showing the MASE distribution of each model at the city, district, and street levels;

[0069] Figure 2 This is an overall architecture diagram of an infectious disease prediction system based on consistency constraints and zero-expansion perception, provided by an exemplary embodiment of the present invention. Detailed Implementation

[0070] To facilitate a clearer and more accurate understanding of the technical solution of the present invention by those skilled in the art, the existing related technologies will be further described below.

[0071] In the field of public health surveillance, the prediction of infectious disease incidence is a crucial basis for resource allocation and risk assessment. Since disease control decision-making requires coordination across multiple administrative levels, such as city, district, and street, the predicted values ​​at different levels must logically satisfy additive consistency; that is, the city-wide predicted value should equal the sum of the predicted values ​​for each district, and the predicted values ​​for each district should equal the sum of the predicted values ​​for its subordinate streets. Therefore, a prediction scheme is needed that can simultaneously produce multi-level prediction results that are mutually consistent.

[0072] Hierarchical time series forecasting is a class of forecasting methods specifically designed for data with hierarchical aggregation structures. In recent years, with the advancement of the concept of "precision public health," the spatial granularity of forecasting has been gradually refined from the traditional provincial and municipal levels to the street level. However, this refinement of granularity has brought three new technical challenges: First, street-level incidence data exhibits extremely sparse and zero-inflation characteristics. Zero-inflation means that the frequency of zero values ​​far exceeds the range that standard count distributions (such as Poisson or negative binomial distributions) can explain. This means that in streets with small populations or low transmission activity, the daily incidence rate is consistently zero, with only occasional non-zero records during outbreaks. Second, the data magnitude varies significantly across different levels. For example, the daily incidence rate at the city level can reach hundreds of cases, while at the street level it is mostly zero or single digits, a difference of several orders of magnitude. This magnitude difference makes it difficult to effectively capture the patterns of low-magnitude signals during joint modeling. Third, the available auxiliary predictive variables (covariates) at each level are naturally heterogeneous. At the city level, indicators such as airport passenger flow and macroeconomic policies are available, while at the street level, more detailed data such as local population density and community pedestrian traffic are relied upon; the two cannot be directly unified. However, existing technological solutions have limitations in addressing these challenges. Specifically:

[0073] First, the extreme sparsity of street-level data is incompatible with the distribution assumptions of existing models, causing the basic prediction model to fail.

[0074] When the spatial granularity of prediction is refined to the street level, case data exhibits extreme sparsity and zero-inflation characteristics, meaning that the number of cases is zero most of the time, with non-zero values ​​only appearing during outbreaks. As Douwes-Schultz et al. (2025) pointed out, ignoring this zero-inflation characteristic and directly applying standard counting models will lead to biased parameter estimations and a significant underestimation of prediction uncertainty. Existing HTSF methods, whether the underlying predictor or the coordination strategy, typically rely on continuous or general counting distributions, failing to specifically design for this data characteristic of extremely high zero-value frequency. This results in several problems: street-level data lacks effective historical signals, causing the model to fall into a "cold start" trap; simultaneously, conventional distribution assumptions cannot accurately characterize the dual-process generation mechanism of zero and non-zero values, resulting in inherently large biases in the underlying predictions. This bias propagates upwards through aggregation constraints, ultimately affecting the prediction quality at all levels. Therefore, how to construct a basic predictor compatible with the zero-inflation characteristics of the underlying data is the primary problem that the HTSF framework needs to solve in this scenario.

[0075] Second, "gradient flooding" occurs during cross-level joint modeling, which hinders the effective learning of sparse signals at lower levels.

[0076] In the HTSF framework, an intuitive way to achieve cross-level information sharing is to simultaneously fit city, district, and street data using a global model with shared parameters. However, the magnitude differences between different level sequences are enormous: the daily incidence rate at the city level can reach hundreds, while at the street level it is mostly zero or single digits. When using scale-sensitive loss functions (such as mean squared error, MSE), as Chen et al. (2018) pointed out in multi-task learning, the loss value and gradient update direction will be completely dominated by the large-value city-level sequences. The direct consequence of this "gradient drowning" phenomenon is that the model parameters mainly learn the smoothing trend of high-level sequences. For the low-level sparse sequences that truly need fine characterization, their weak gradient signals are masked during backpropagation, making it almost impossible to effectively learn the patterns at that level. This explains why simple global joint modeling cannot solve the fine-grained prediction problem. Therefore, how to overcome the gradient imbalance caused by magnitude differences and ensure the feature learning ability of sparse sequences in multi-level joint training is a problem that must be solved.

[0077] Third, the heterogeneity of covariates and the mismatch of features in multi-level data restrict the effective integration of cross-level information.

[0078] In practical public health surveillance, the types of covariates available at different administrative levels are heterogeneous. As reviewed by Athanasopoulos et al. (2024), one of the advantages of the HTSF framework is that it allows models at each level to use their own covariates. For example, city-level forecasts can benefit from macro-level indicators such as the citywide influenza positivity rate and population mobility index, while street-level forecasts rely on fine-grained features such as community population density. However, this also makes cross-level information sharing difficult: forcibly assigning macro-level indicators to micro-level individuals introduces noise that is mismatched with the micro-scale, thus undermining the forecasting effect. When dealing with such multi-source heterogeneous data, existing technologies either adopt a "global model" strategy, forcibly aligning features from all levels to a unified dimension, which inevitably leads to the discarding of level-specific features or the imputation of missing values ​​without justification; or they abandon cross-level information borrowing and only perform single-level modeling, but this causes the underlying model to lose the constraints of the macro-level situation. Therefore, how to achieve cross-level transfer of valuable information while preserving the original heterogeneous features of each level and avoiding unjustified data alignment operations is a key challenge to ensuring the forecasting performance of the HTSF framework.

[0079] In summary, when existing technologies are applied to the prediction of street-level infectious diseases (a scenario with unique complexities, such as extreme sparsity, large magnitude differences, and heterogeneous covariates), they each face problems such as single-level logical contradictions, fragmented training mechanisms, or failure of sparse data modeling. There is a lack of a technical solution that can adapt to the distribution characteristics of extremely fine-grained data and simultaneously optimize the prediction accuracy and consistency of results at each level in an end-to-end collaborative manner.

[0080] In view of this, the present invention provides a novel solution, namely, the infectious disease prediction method and system based on consistency constraints and zero-expansion perception. The technical concept of the present invention lies in breaking the inherent paradigm of traditional hierarchical prediction of "independent prediction first, then unified coordination" and constructing an end-to-end collaborative prediction architecture of "structural decoupling - distribution adaptation - constraint internalization". Specifically, to address the gradient submersion problem caused by the significant difference in data volume across different levels, the design proposes constructing independent long short-term memory network encoders for cities, districts, and streets respectively. This decouples the learning of high-level macro trends and low-level sparse signals during training, while naturally accommodating the different sets of covariates at each level. For the extremely zero-expansion distribution characteristics of street-level data, a dual-branch gating output mechanism is introduced, deconstructing the prediction task into two parallel subtasks: zero-value probability judgment and positive-value intensity estimation. This aligns with the real data generation process of "first determining whether an illness has occurred, then estimating the number of cases." Regarding hierarchical aggregation constraints, the design abandons traditional post-hoc hard adjustment strategies and instead utilizes orthogonal projection to transform additive consistency into a soft regularization guide term during training, enabling prediction accuracy optimization and consistency constraints to be optimized collaboratively under the same loss function.

[0081] In summary, this invention deeply embeds the prior features of data distribution and the domain structure knowledge of hierarchical aggregation into the training logic of the model. Through end-to-end joint gradient backpropagation, each level can maintain its own prediction accuracy while naturally generating collaborative and consistent prediction results that meet the business logic.

[0082] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings.

[0083] According to a first aspect of the present invention, an infectious disease prediction method based on consistency constraints and zero-inflation perception is provided, comprising the following steps:

[0084] Obtain historical incidence rate sequence data for multiple administrative levels and covariate data for each administrative level;

[0085] Independent long short-term memory network encoders were constructed at the city, district, and street levels, respectively. Covariate data and historical incidence sequence data at each level were input into the corresponding level of the long short-term memory network encoder to extract the temporal features at each level.

[0086] For the city and district levels, the predicted values ​​are generated by linear mapping based on the temporal characteristics of the corresponding levels. For the street level, the predicted values ​​are generated by a zero-inflation gating output structure based on the temporal characteristics of the street level. The zero-inflation gating output structure calculates in parallel the gating probability vector representing the probability of non-zero cases in each street and the intensity vector representing the expected number of non-zero cases. The street-level predicted values ​​are obtained by multiplying the gating probability vector and the intensity vector element by element.

[0087] The city-level, district-level, and street-level predicted values ​​are concatenated into a unified prediction vector.

[0088] Based on the unified prediction vector and the true value vector, the joint loss is calculated. The joint loss includes city-level prediction error loss, district-level prediction error loss, street-level prediction error loss, and soft consistency regularization loss used to penalize prediction values ​​that deviate from the hierarchical summation consistency constraint.

[0089] Based on the joint loss, the parameters of the long short-term memory network encoder and the zero-inflation gated output structure at each level are synchronously updated through backpropagation until training is complete;

[0090] The input data at each level at the time to be predicted are input into the trained long short-term memory network encoders and zero-inflation gating output structures at each level, and the prediction results at the city, district, and street levels are output simultaneously.

[0091] Through the above technical solution, firstly, the method of the present invention can effectively adapt to the zero-inflation distribution of street-level data, improving the accuracy of fine-grained prediction. Specifically, in existing related technologies, the basic predictors of traditional hierarchical time series prediction methods are usually constructed based on assumptions of continuous distribution or general counting distribution (such as Poisson distribution, negative binomial distribution). When applied to street-level infectious disease prediction, the daily incidence rate at this level is zero most of the time, and only occasionally has a non-zero value during the outbreak period. This extreme sparsity and zero-inflation characteristic makes it impossible for conventional distribution assumptions to accurately characterize the dual-process generation mechanism of the data (i.e., "whether there is an illness" and "how many people are infected" are driven by different factors), resulting in a huge bias in the basic prediction results at the street level. Moreover, this bias will be transmitted upward to the district and city levels through aggregation constraints, affecting the prediction quality at all levels.

[0092] The method of this invention, when targeting street-level prediction, generates street-level predicted values ​​through a zero-inflation gated output structure. This structure computes in parallel a gated probability vector representing the probability of non-zero cases in each street and an intensity vector representing the expected number of non-zero cases. The gated probability vector is then multiplied element-wise by the intensity vector to obtain the street-level predicted value. This structure explicitly decouples street-level case prediction into two parallel subtasks at the model architecture level: "case probability judgment" and "case quantity estimation." The former determines whether the predicted value is activated, while the latter determines the magnitude of the activated value. This effectively adapts to the underlying generation mechanism of street-level case data, which consists of "a large number of true zero values ​​+ a small number of positive values." This allows the model to output near-zero predicted values ​​for streets with consistently zero cases, avoiding biased estimations due to distribution assumption mismatches, thereby effectively improving the accuracy of fine-grained hierarchical predictions.

[0093] Second, the method of this invention can avoid gradient submersion across different levels, ensuring effective learning of sparse sequence features. Specifically, in existing related technologies, if a global model with shared parameters is used to fit city-level, district-level, and street-level data simultaneously, the daily incidence rate at the city level can reach hundreds of cases, while at the street level it is mostly zero or single digits, a difference of several orders of magnitude. Scale-sensitive loss functions (such as mean squared error) will cause the loss value and gradient of the city-level sequence to dominate. During backpropagation, the weak gradient signal generated by the street-level sparse sequence is masked by the large numerical gradient of higher-level sequences, i.e., a "gradient submersion" phenomenon occurs. This causes the model parameters to mainly learn the smooth trend of the city-level sequence, while the unique local fluctuation patterns and sparse burst patterns of the street level are almost impossible to learn effectively.

[0094] In the method of this invention, independent long short-term memory (LSTM) network encoders are constructed for the city, district, and street levels, respectively. In subsequent training steps, the parameters of each level's LSM network encoder and the zero-inflation gated output structure are synchronously updated via backpropagation based on the joint loss. This independent design of encoder parameters at each level decouples the feature extraction processes at the city, district, and street levels in the parameter space. Gradients from large numerical sequences at the city level do not flow into the street-level encoder; the gradient signal of the street level is only propagated back within its independent encoder, fundamentally blocking the path of gradient saturation across levels. Simultaneously, "synchronous update" ensures that each level's encoder converges independently under the same optimization objective, fully guaranteeing the feature learning capability of sparse sequences at the street level.

[0095] Third, the method of this invention supports multi-level heterogeneous covariates, avoiding unfounded data alignment operations. Specifically, in existing related technologies, in actual public health monitoring scenarios, city-level models can obtain macro-level covariates such as airport passenger flow, influenza positivity rate, and policy indicators, while street-level models rely on fine-grained features such as community population density and local pedestrian flow. The two are inherently heterogeneous and cannot be converted to each other through simple transformations. Existing global model strategies require forcibly aligning all hierarchical features to a unified dimension, which inevitably leads to the discarding of hierarchical-specific covariates or the artificial expansion of dimensions through interpolation, downfilling, etc., thereby introducing noise that does not match the micro-scale and losing valuable hierarchical-specific information; while abandoning the borrowing of cross-level information would cause the street-level model to lose the trend constraint of the macro-level situation.

[0096] In the data acquisition step of the method of this invention, "covariate data of each administrative level" is acquired, and in the encoding step, "covariate data of each level and historical incidence sequence data are respectively input into the corresponding level's long short-term memory network encoder." Since the encoders at each level are independent and do not share parameters, the city-level input data can include city-specific features such as airport passenger flow and policy indicators, while the street-level input can include street-specific features such as population density. The covariate dimensions of each level do not affect each other and do not require any form of alignment or padding. In this way, while preserving the integrity of the original heterogeneous feature information of each level, the cross-level gradient propagation of the soft consistency regularization term in the joint loss function allows macro-level situational information to indirectly influence street-level predictions in a constrained manner, achieving information sharing that "preserves heterogeneity and indirectly transmits information." This avoids information loss and noise introduction caused by forced alignment while maintaining cross-level information linkage.

[0097] Fourth, the method of this invention enables end-to-end joint optimization, overcoming the fragmented training defects of the two-stage method. Specifically, in the traditional "predict first, coordinate later" two-stage method, the basic predictions at each level are generated independently, isolated from each other during the training phase, and unable to share information. After the predictions at each level are completed, a post-processing coordination step is used to correct them to meet aggregation constraints. This fragmented training mode leads to two serious consequences: First, the training error of the bottom-level street-level model cannot be perceived by higher levels and used to guide parameter adjustment during the training process; cross-level information is only passively superimposed at the final output level. Second, the coordination step can only correct the output value and cannot correct the model parameters; the inherent bias in the basic prediction still exists within the model after coordination, limiting the upper limit of the global prediction accuracy.

[0098] In the method of this invention, city-level, district-level, and street-level predicted values ​​are concatenated into a unified prediction vector. Based on the unified prediction vector and the true value vector, a joint loss is calculated. The joint loss includes city-level prediction error loss, district-level prediction error loss, street-level prediction error loss, and a soft consistency regularization loss used to penalize deviations of predicted values ​​from hierarchical summation consistency constraints. The parameters of the long short-term memory network encoders and zero-inflation gated output structures at each level are synchronously updated through backpropagation based on the joint loss. In other words, this invention unifies the predicted values ​​at each level into the same loss function and uses a backpropagation mechanism to synchronously propagate the gradient of the joint loss back to the encoders and output modules at all levels, thereby enabling cross-level information interaction during the training phase rather than the prediction phase. Parameter updates for each level of the model are no longer performed in isolation but collectively serve the global goal of "optimal accuracy at each level and overall consistency," thus fundamentally overcoming the fundamental defect of fragmented training in two-stage methods.

[0099] Fifth, the method of this invention can output logically consistent multi-granularity prediction results, directly supporting hierarchical collaborative decision-making. Specifically, in existing related technologies, single-level independent prediction methods model cities, districts, and streets separately and output prediction results separately. There is no constraint relationship between the prediction values ​​of each level. When summing them up afterward, logical contradictions often occur, such as "the city-wide prediction value is not equal to the sum of the predictions of each district." This makes it unsuitable for public health decision-making scenarios that require hierarchical linkage (such as material allocation, resource deployment, and risk assessment according to administrative levels). It requires post-event adjustments based on human experience, which increases the uncertainty of the decision-making chain and reduces the timeliness of prediction information.

[0100] The method of this invention introduces a "soft consistency regularization loss" in its loss function to penalize predicted values ​​deviating from hierarchical summation consistency constraints. This loss is then used in parameter updates along with the prediction error losses at each level through backpropagation. This ensures the model is continuously guided during training by the principle that "predicted values ​​should approximate the consistency subspace." Furthermore, in the prediction output step, prediction results at the city, district, and street levels are output simultaneously. Thus, after end-to-end training with soft consistency regularization constraints, the simultaneously output three-level prediction results naturally tend to satisfy an additive relationship (i.e., the city-wide predicted value is approximately equal to the sum of the districts' predictions, and the districts' predicted values ​​are approximately equal to the sum of their subordinate streets' predictions). This allows for direct use in hierarchical collaborative decision-making without any post-processing adjustments, providing logically consistent and readily available quantitative evidence for disease control departments to allocate resources and conduct risk assessments according to administrative levels.

[0101] The technical solution of the present invention will be further described below with reference to an exemplary embodiment.

[0102] 1. Overall framework.

[0103] Please see Figure 2The prediction system constructed in this exemplary implementation can include three levels of prediction sub-models: a city-level prediction sub-model, a district-level prediction sub-model, and a street-level prediction sub-model. All three sub-models use a Long Short-Term Memory (LSTM) network as the core encoder, and the network structures of each sub-model are independent and their parameters are not shared. The outputs of the three-level sub-models are trained end-to-end synchronously through a joint loss function, which consists of four parts: city-level prediction error, district-level prediction error, street-level prediction error, and a soft consistency regularization term.

[0104] The input to the prediction system is three-level hierarchical time series data. This data includes one city-level time series. District-level time series and There are three street-level time series. A strict aggregation relationship exists between the three-level series: the number of city-level cases equals the sum of all district-level cases, and the number of cases in each district equals the sum of all street-level cases under its jurisdiction. Let the historical observation window length of the three-level series be... The predicted step size is After training, for a given historical time window, it can simultaneously output city-level, district-level, and street-level forecasts for the future. The projected values ​​are displayed at each time step. Simultaneous output of city-level, district-level, and street-level forecasts for the future is possible. The single-point predicted value at each time step, i.e. the deterministic point estimate of the number of new cases per day at each level.

[0105] 2. Problem definition and symbol explanation.

[0106] For ease of subsequent description, the relevant symbols are now uniformly defined.

[0107] Set time The city-level observation value The district-level observation vector is The street-level observation vector is , It is the set of natural numbers. For streets... Its covariate vector is denoted as ,in This represents the dimension of street-level features. The corresponding city-level and district-level covariates are denoted as follows: and ,in, and These represent the dimensions of city-level and district-level characteristics, respectively. It should be noted that... , and The values ​​can be different, meaning that each level can use its own set of covariates without the need for forced alignment of feature dimensions.

[0108] Let the aggregation matrix be... Used to characterize hierarchical aggregation relationships. The first row of the aggregation matrix consists entirely of 1s, indicating that the total number of city-level values ​​equals the sum of the values ​​of all street-level values; for the... District , matrix number The elements of the row satisfy If and only if the street If a value belongs to the region, it is 0; otherwise, it is 0. The hierarchical constraints expressed by this matrix form the mathematical basis for subsequent soft consistency regularization.

[0109] The three-level predicted values ​​are stacked into a unified vector in top-down order:

[0110]

[0111] in, This represents the total number of tertiary sequences.

[0112] 3. Three-layer LSTM encoder network structure.

[0113] The city-level, district-level, and street-level prediction sub-models of this invention all employ the same LSTM encoder structure, but the parameters of each sub-model are independent of each other, and are denoted as follows: , ,and This hierarchical, independent modeling design fundamentally avoids the gradient flooding problem caused by the significant difference in data volume between city-level and street-level models. Taking the street-level prediction sub-model as an example, the internal structure of the LSTM encoder is explained. For the street... First, the covariate vector Compared with observed values The street is pieced together in time Input vector:

[0114]

[0115] All The streets in the time window to The input vectors are stacked according to time steps to obtain the street-level prediction sub-model at time step [time]. Input matrix:

[0116]

[0117] The LSTM encoder consists of the following components in sequence: a first LSTM layer with a hidden state dimension of 64, returning the complete time series; a second LSTM layer with a hidden state dimension of 32, returning only the output of the last time step; a Dropout layer with a dropout rate of 0.2; a first fully connected layer with an output dimension of 128, using the ReLU activation function; and a second fully connected layer with an output dimension of 64, also using the ReLU activation function. Let the output vector of the second fully connected layer be... .

[0118] The LSTM encoder structure for the city-level and district-level prediction sub-models is exactly the same as described above, and its input matrix... and The covariates and observed values ​​at the city and district levels were constructed in the same manner.

[0119] For the city-level and district-level prediction sub-models, the encoder output is directly mapped to the predicted value through a linear transformation:

[0120] ,

[0121] in, and As weight, and This is a bias term.

[0122] 4. Zero-inflation gated output mechanism for sparse data.

[0123] To address the extremely sparse and zero-inflation distribution characteristics of street-level incidence data, this exemplary implementation designs a zero-inflation-gated output structure to replace the conventional linear output layer. This structure decomposes the prediction process into two parallel subtasks: zero-value occurrence probability estimation and non-zero-value intensity estimation, enabling the model to more accurately model zero-inflation count data.

[0124] Specifically, the zero-inflation gated output structure acquires the encoder output vector. Then, the gating probability vector is computed in parallel using two independent sets of linear transformations and activation functions. and intensity vector :

[0125] ,

[0126] in, The sigmoid function maps the gate probability to the (0,1) interval; This ensures that the output intensity is non-negative.

[0127] For the street , Indicates the first Each street at any time The probability of non-zero cases, This represents the expected number of non-zero cases under this probability. The final street-level prediction is obtained by element-wise multiplication of the two:

[0128]

[0129] in, This represents the Hadamard product. The mechanism described above ensures that when the model determines that a street is highly likely to have no cases, the predicted value automatically tends towards zero; only when the model determines that there is a possibility of cases will a non-zero intensity prediction be activated and contribute to the final output. This design, from the model structure level rather than simply from the loss function level, naturally adapts to the sparsity of street-level data.

[0130] 5. Hierarchical aggregation constraints and soft consistency regularization.

[0131] As mentioned earlier, the three-level time series should satisfy the additive consistency constraint. This exemplary implementation does not adopt the traditional "two-step method" (i.e., first generating independent base predictions for each level and then adjusting them through post-processing steps to meet the aggregation constraint), but internalizes the hierarchical aggregation constraint as a soft consistency regularization term in the training process to achieve end-to-end joint optimization.

[0132] First, using aggregation matrix Define the constraint matrix for:

[0133]

[0134] in, for An identity matrix of order 1. Then the hierarchical aggregation constraint is equivalent to:

[0135]

[0136] This equation defines the prediction vector space. A linear subspace, denoted as the consistent subspace. , Any vector within this subspace naturally satisfies the summation consistency across all levels. Since each district contains at least one street and the entire city encompasses all districts, the matrix... To complete the full term, therefore Reversible.

[0137] based on Subspace, define orthogonal projection matrix for:

[0138]

[0139] For any prediction vector Its projection for subspace and The point with the closest Euclidean distance is the optimal approximation that satisfies complete summation consistency.

[0140] One of the innovations of this exemplary implementation is that it does not directly... Instead of using the final prediction output, the deviation between the predicted value and the consistency subspace is penalized as part of the loss function during training. This "soft constraint" strategy allows the model to adaptively balance prediction accuracy and hierarchical consistency, avoiding the potential for hard projection to mask the prediction errors of lower layers, while guiding each level of the model to learn the ability to make co-predictions.

[0141] 6. End-to-end joint training and loss function design.

[0142] This exemplary implementation places the city-level prediction sub-model, district-level prediction sub-model, and street-level prediction sub-model (including the zero-inflation gated output module) described above into a unified end-to-end framework for joint training. All trainable parameters, including the parameters of each layer of the LSTM network, are used. , ,and and the weights and biases of the street-level gating module , , , All are updated synchronously through the Adam optimizer.

[0143] The training dataset is divided into a training set (60%), a validation set (20%), and a test set (20%) in chronological order. During training, an early stopping mechanism (patience=10) is applied to the validation set to prevent overfitting.

[0144] The learning rate of the Adam optimizer was set to 0.001, and the remaining hyperparameters used the default values ​​of the TensorFlow / Keras framework. During training, the validation set loss was monitored, and an early stopping mechanism (patience=10) was employed to prevent overfitting.

[0145] Total Loss Function: The total loss function consists of a weighted sum of four components, corresponding to the city-level prediction error, district-level prediction error, street-level prediction error, and soft consistency regularization:

[0146]

[0147] Wherein, the sample size of a single training batch is set to be... .

[0148] Loss functions for city and district levels: Considering the higher density and lower proportion of zero values ​​in city and district-level data, the mean squared error (MSE) is directly used as the loss function.

[0149] ,

[0150] Street-level loss function. To address the high sparsity and right-skewed positive distribution of street-level data, this invention designs a composite loss function consisting of three sub-items:

[0151]

[0152] The first term is the principal loss, calculated using the negative log-likelihood of the Tweedie distribution. Its expression is: The Tweedie distribution is a family of composite Poisson-gamma distributions suitable for modeling count data with zero inflation and variances not equal to the mean. The power parameter... The range of values ​​is (It should be noted that in this invention, the determination was made through preliminary experiments.) At that time, this value allows the model to effectively fit the distribution characteristics of street-level incidence data.

[0153] The second term is the gating loss, which uses binary cross-entropy (BCE). For the first The gating probability vector corresponding to each sample This is the indicator function, taking the value 1 when the true value is greater than zero, and 0 otherwise. This loss term enforces a gating probability. Learning to distinguish between "true zero" and "non-zero" modes is a key monitoring signal for zero-expansion gated output structures.

[0154] The third term is the non-zero sample amplitude loss, calculated only for sample points where the true value is non-zero. This loss term constrains the prediction strength. Approaching the non-zero true value. Among them, For the first In the nth sample The intensity prediction value corresponding to each street It is a very small positive number, used to prevent the denominator from being zero.

[0155] The weighting coefficients of the above-mentioned gating loss and non-zero sample amplitude loss , All values ​​can be set based on preliminary experiments; for example, each value can be 0.1. This setting ensures that the magnitudes of the three sub-loss terms remain coordinated in the early stages of training, avoiding the problem of one term's gradient becoming too large and dominating parameter updates, thus guaranteeing balanced convergence of each subtask in joint optimization.

[0156] Soft consistency regularization loss. The soft consistency regularization loss consists of two parts, both based on the projection matrix. Perform the calculation:

[0157]

[0158] The first item (consistency deviation penalty): measures the original predicted value. Its projection on the consistent subspace The Euclidean distance between them penalizes the degree to which the predicted value deviates from the summation consistency constraint.

[0159] The second term (projected prediction error): measures the projection of the predicted value onto the consistency subspace. Compared with the true value The error between them ensures that the model does not sacrifice prediction accuracy while satisfying the consistency constraint trend.

[0160] The combined effect of these two terms is to guide the model output itself towards summative consistency (through the first term), while ensuring that this tendency is based on maintaining true predictive ability, rather than at the cost of accuracy (through the second term). Compared to the traditional two-step method that directly applies hard projection adjustments after prediction, the soft constraint strategy achieves end-to-end joint optimization of prediction accuracy and hierarchical consistency.

[0161] 7. Predicting the output process.

[0162] After training, the actual usage process of the three-layer joint prediction model is as follows:

[0163] Step 1: Following the method defined above, acquire feature data at all levels up to the current moment and historical case data from the data acquisition module to construct input matrices at the city, district, and street levels. , , ;

[0164] Step two: The above input matrices are fed into the corresponding trained prediction sub-models at each level. The city-level and district-level prediction sub-models directly output the predicted values ​​through a linear output layer; the street-level prediction sub-model calculates the gating probability and intensity vector through a zero-inflation gating output structure and multiplies them element by element to obtain the predicted values.

[0165] Step 3: Calculate the three-level predicted values ​​according to the formula. The data is assembled and output as the system's multi-granularity prediction results.

[0166] In the final output prediction results, the values ​​at each level tend to satisfy an additive relationship due to the training constraints of soft consistency regularization. That is, the city-level prediction value is approximately equal to the sum of the street-level prediction values, and the district-level prediction value is also approximately equal to the sum of the prediction values ​​of its subordinate street-level values. This result provides a logically consistent quantitative basis for subsequent material allocation, resource deployment, and risk assessment according to administrative levels.

[0167] The technical solution of the present invention will be described below with reference to a specific embodiment.

[0168] Assume a city has two districts (District A and District B), and each district has two subdistricts (District A has subdistricts A1 and A2; District B has subdistricts B1 and B2). The hierarchical structure is: 1 city-level node, 2 district-level nodes, and 4 subdistrict-level nodes. For simplicity, assume the input window length L is 7 days and the prediction step size h is 1 day.

[0169] Step 1: Construct the input matrices for each level.

[0170] At prediction time t, historical case counts and covariate data for the past 7 days (from day t-6 to day t) are retrieved from the database and preprocessed using Z-score standardization independently at each level.

[0171] Construct three levels of input matrices according to the formula:

[0172] City-level input matrix (Dimension 7 × City-level Feature Count): Includes two levels of covariates for daily data across the city. The first level comprises demographic and clinical characteristics common to the city, district, and street levels (including average age, average reporting delay days, gender structure, regional distribution of cases, proportion of case categories, clinical severity score, presence of imported cases, and the number and proportion of 12 population categories). The second level comprises city-specific macro-monitoring indicators (including the positivity rate of influenza and a certain coronavirus, fever visits and virus detection indicators at medical institutions, number of critically ill patients in hospitals, airport passenger traffic, and search popularity data for coronavirus-related symptoms and diseases provided by Baidu Index). It also includes standardized historical case counts for the city.

[0173] District-level input matrix (Dimension 7 × Number of District-Level Features × 2): Includes daily level 3 general covariates for each of the two districts (mean age, reporting delay days, gender structure, and other demographic and clinical characteristics, with specific categories as above) and historical incidence rates. District-level data does not include city-specific macroeconomic indicators such as influenza positivity rate and airport passenger flow.

[0174] Street-level input matrix (Dimension 7 × Street-level feature count × 4): Includes daily level 3 general covariates (demographic and clinical characteristics, specific categories as above) and historical incidence rates for each of the four streets. Also excludes city-specific macro-monitoring indicators.

[0175] It should be noted that the types and dimensions of covariates in the three levels of input can be completely different, without the need for downfilling or alignment. For example, the "airport passenger throughput" indicator, which is unique to the city level, is only input into the city-level network and will not be forcibly assigned to each street.

[0176] Step 2: Calculate the predicted value using hierarchical parallel forward propagation.

[0177] like Figure 2 As shown, the three input matrices are fed into the corresponding LSTM prediction sub-models (labeled 1, 2, and 3).

[0178] City-level network (labeled 1): After the LSTM encoder extracts the temporal features, it outputs a non-negative scalar value through a fully connected layer and a softplus activation function, which is the predicted number of cases in the city on day t+1. .

[0179] District-level network (labeled 2): Similarly, a vector containing two elements is output in parallel. After inverse standardization, this vector represents the predicted number of cases in districts A and B on day t+1. .

[0180] Street-level network (label 3): Encoder output feature vector Subsequently, instead of directly outputting the predicted value, it is fed into the zero-inflation gated output module (labeled 8). This module performs parallel calculations:

[0181] The gated probability vector p (left branch of label 8): after passing through the sigmoid function, it outputs four probability values ​​between 0 and 1, such as [0.02, 0.15, 0.03, 0.80], which represent the probabilities of non-zero cases occurring in the four streets respectively.

[0182] Intensity prediction vector (Branch to the right of label 8): The softplus function outputs four non-negative values, such as [1.2, 2.0, 0.5, 4.0], representing the expected number of cases if they occur.

[0183] The final street-level predicted value is obtained by element-wise multiplication of the two values, followed by inverse standardization: =[0.024,0.3,0.015,3.2]. It can be seen that for streets A1 and B1, where the probability is extremely low (zero), the predicted value is close to zero; while for street B2, which may experience an outbreak, a non-zero prediction is given. This naturally aligns with the characteristics of sparse data.

[0184] Step 3: Concatenate the prediction vectors and calculate the loss.

[0185] The three-level predicted values ​​are concatenated into a unified prediction vector in the order of "city-district-street". (Label 4).

[0186] like Figure 2 As shown in the lower right section, the joint loss function begins to be calculated:

[0187] City-level losses (No. 5) and district-level losses (Label 6): Calculate the predicted values ​​respectively (e.g.) The mean square error between the actual number of cases and the number of cases in the city.

[0188] Street-level losses (Label 7): Calculate the composite loss consisting of a weighted sum of three components: the first is the Tweedie loss, which serves as the main loss to fit the overall distribution of the zero-inflated count data; the second is the gating loss (binary cross-entropy), which supervises the gating probability p to learn the zero-value distribution pattern in the real data; the third is the intensity loss (MSE), which is calculated only on samples with true values ​​greater than zero, constraining the accuracy of the intensity prediction μ. The weight coefficients for the last two terms are each 0.1. In this case, the true zero-value distribution pattern supervises the gating probability p, and the true number of non-zero values ​​supervises the intensity μ.

[0189] Soft consistency regularization loss (Label 10): The concatenated predicted vectors are projected onto a "perfectly consistent" subspace using an orthogonal projection matrix M (Label 9), and two soft-constraint losses are calculated: the first is a consistency deviation penalty, calculated by the Euclidean distance between the projected vector and the original predicted vector; the second is the projection prediction error, calculated by the Euclidean distance between the projected vector and the true value vector. The two losses are directly added together and included in the total loss. For example, if the sum of the predicted values ​​for the four streets is 0.024 + 0.3 + 0.015 + 3.2 = 3.539, while the sum of the district-level predictions is 6.0534 (district-level predicted value) + ..., there may be a bias. It will penalize both this bias and the difference between the projected prediction and the true value.

[0190] Step 4: Backpropagation and joint parameter update.

[0191] The four losses are summed at the total loss node (labeled 11) to form the final optimization objective L. Subsequently, through the backpropagation mechanism (labeled 12), the Adam optimizer synchronously updates all parameters of the city-level, district-level, and street-level LSTM networks, as well as the zero-inflation gating module.

[0192] Step 5: Iterative training and prediction.

[0193] Repeat steps one through four to iterate through all training set data. Monitor the total loss on the validation set and apply an early stopping mechanism (terminate training if the validation loss does not decrease for 10 consecutive epochs) to prevent overfitting. After training, the model parameters are fixed. When a new prediction is needed, simply execute steps one and two to simultaneously generate prediction results for the city, district, and street levels, which can be directly used for hierarchical collaborative decision-making.

[0194] Experimental comparison:

[0195] 1. Experimental data and setup.

[0196] The data used for verification in this invention is a three-level time series data of a coronavirus in a certain city, covering the period from January 1, 2023 to December 31, 2025, with a daily observation frequency. The dataset contains 1 city-level unit, 16 administrative districts and 1 economic development zone (a total of 17 district-level units), and 357 street-level units, with a strict top-down aggregation relationship between the series. The core predictive variable is the daily number of new cases. The data is divided into a training set (60%, 658 days), a validation set (20%, 219 days), and a test set (20%, 219 days) in chronological order. In the single-step prediction experiment, the input window length is 7 days, and the prediction step size is 1 day.

[0197] 2. Comparison Model.

[0198] To comprehensively verify the effectiveness of the proposed model (SPAR-Hier), four typical time series forecasting models were selected as the basic predictors: Long Short-Term Memory Network (LSTM), Gated Recurrent Unit (GRU), Temporal Convolutional Network (TCN), and Vector Autoregression (VAR). These four basic predictors were combined with four classic hierarchical two-stage reconciliation strategies to form 20 comparative benchmarks. The four reconciliation strategies include: Bottom-Up (retaining only the bottom-level street-level predictions, with district and city-level predictions obtained through upward aggregation); Top-Down (retaining only city-level predictions, with district and street-level predictions obtained through downward decomposition according to historical proportions); MiddleOut (retaining only district-level predictions, with city-level predictions obtained through upward aggregation and street-level predictions obtained through downward decomposition); and MinT (MinTrace), which estimates the covariance matrix of the basic prediction errors at each level and minimizes the trace of the global prediction error while satisfying aggregation constraints to obtain the optimal reconciliation weights. The above comparative methods cover the current mainstream technical solutions for two-stage hierarchical forecasting.

[0199] The combined models in the first column of Table 3 are named in the format of "Base Predictor / Harmonization Strategy". The part before the slash indicates the model used to generate the base predictions, and the part after the slash indicates the hierarchical harmonization method used. For example, LSTM / BottomUp indicates that a Long Short-Term Memory network is first used for prediction, and then a bottom-up method is used for harmonization. Models labeled only with LSTM, GRU, TCN, or VAR indicate that the predictions of the corresponding models have not been hierarchically harmonized. The above comparison methods cover both unharmonized base prediction methods and typical two-stage hierarchical harmonization methods.

[0200] 3. Evaluation indicators.

[0201] Four indicators were used to evaluate the forecasting performance: mean absolute error (MAE), root mean square error (RMSE), root mean square scaling error (RMSSE), and mean absolute scaling error (MASE). Among them, RMSSE and MASE are scale-independent indicators with weekly seasonal (period m=7) naive forecasts as the benchmark. A value less than 1 indicates that the forecast accuracy is better than the naive benchmark, and the smaller the value, the greater the relative improvement.

[0202] 4. Experiment.

[0203] The prediction errors of each model on the test set are summarized in Tables 1, 2, and 3 below. The results show that this invention (SPAR-Hier) significantly outperforms all baseline models in street-level prediction. The street level is the finest spatially granular level with the most severe zero inflation, and it is also the most challenging prediction task in this problem. In contrast, the best-performing baseline LSTM and its harmonic variants generally have MASE values ​​close to or exceeding 1.0 at this level, meaning their accuracy is inferior to simple weekly seasonal naive predictions. GRU, TCN, VAR, and their harmonic variants also have high errors at the street level, indicating that traditional time series models face fundamental limitations on extremely sparse data.

[0204] At the district and city levels, SPAR-Hier is on par with the optimal baseline, demonstrating that its street-level advantage does not come at the expense of higher-level accuracy. In district-level predictions, LSTM / MinTrace's MAE (5.74) and LSTM / TopDown's RMSSE (0.312) are slightly better than SPAR-Hier. However, both of these reconciliation strategies fail significantly at the street level, with MinTrace achieving a street-level MASE as high as 1.12 and TopDown reaching 1.40. In contrast, this invention (SPAR-Hier) maintains competitive district-level predictions (with a difference of less than 6% from the optimal baseline) while achieving a street-level MASE as low as 0.45. This result reveals that the cost of the two-stage model's pursuit of optimal reconciliation at the street level is the sacrifice of the error covariance structure between the district and street levels; while the soft consistency constraint of this invention (SPAR-Hier) achieves a more balanced cross-level information transmission. Multi-task models face the gradient submersion problem. LSTM performs reasonably well in city-level prediction (RMSSE=0.26), but achieves 0.95 in street-level MASE, close to the naive benchmark. This confirms that when sequences with vastly different magnitudes are forcibly incorporated into a unified parameter space, the large-value city-level sequences dominate gradient updates, making it difficult to effectively learn the patterns of the sparse lower-level sequences. This invention (SPAR-Hier) successfully circumvents this obstacle by using hierarchical independent modeling followed by joint training.

[0205] Table 1. Comparison of average prediction performance of each model across 3 layers

[0206]

[0207] Table 2 Comparison of prediction errors of each model at the city, district, and street levels. (Table 1)

[0208]

[0209] Table 3 Comparison of prediction errors of each model at the city, district, and street levels. (Table 2)

[0210]

[0211] The original VAR and reconciliation model results (without retaining four decimal places) can be found in Table 4 below:

[0212] Table 4. Results of the original VAR and the coordinated model

[0213]

[0214] Please see Figure 1 , Figure 1This paper presents the MASE distribution at the city, district, and street levels using the method of this invention (SPAR-Hier) and four baseline models (LSTM, GRU, TCN, VAR) and their four harmonic variants (BottomUp, TopDown, MiddleOut, MinTrace). The first bin in each group represents SPAR-Hier, and the others represent the variants of the corresponding baselines. Each bin is plotted based on predictions from 10 independent runs, reflecting the median, interquartile range, and outliers of MASE. The vertical axis uses a logarithmic scale to clearly compare the magnitude differences between levels.

[0215] from Figure 1 The results clearly show that SPAR-Hier is lower than or equal to all baseline models across all levels. The difference is most significant at the street level, where the median of SPAR-Hier is far lower than the baseline for each group, while some harmonic variants in the TCN group show extremely high values, indicating severe prediction failure on sparse data. At the district level, the median of SPAR-Hier is roughly equal to the optimal harmonic variant of LSTM, but better than GRU, TCN, and VAR groups. At the city level, SPAR-Hier also maintains its lead. Overall, SPAR-Hier demonstrates higher accuracy compared to the baseline model, more stable performance across levels, and a smaller outlier range.

[0216] Please see Figure 2 ,exist Figure 2 In the text, the meanings of each symbol are as follows:

[0217] Label 1: City-level LSTM prediction network - used to receive city-level input data (including historical incidence rates and city-specific covariates) and output predicted daily incidence rates for the entire city;

[0218] Label 2: District-level LSTM prediction network - used to receive district-level input data (including historical incidence rates and district-specific covariates) and output predicted daily incidence rates for each district;

[0219] Label 3: Street-level LSTM prediction network - used to receive street-level input data (including historical incidence rates and street-level specific covariates) and output predicted daily incidence rates for each street;

[0220] Label 4: Three-level prediction result splicing node—sponging together the city-level, district-level, and street-level prediction values ​​into a unified prediction vector;

[0221] Label 5: City-level forecasting error loss —The deviation between the city-level predicted value and the actual value is calculated using the mean square error.

[0222] Label 6: District-level prediction error loss —The mean square error is used to calculate the deviation between the district-level predicted value and the actual value;

[0223] Label 7: Street-level prediction error loss — It consists of three parts: Tweedie loss, gated binary cross-entropy loss, and intensity mean square error loss, and is used to adapt to zero-inflated counting data;

[0224] Label 8: Zero-inflation gated output module – located at the output of the street-level LSTM prediction network, used to compute the gated probability vector and intensity vector in parallel, and obtain the final street-level prediction value by multiplying them element by element;

[0225] Label 9: Orthogonal Projection Matrix Module – Constructing Projection Matrix Based on Aggregation Matrix , used to project any prediction vector onto a consistency subspace that satisfies hierarchical summation constraints;

[0226] Label 10: Soft consistency regularization loss —It consists of two parts: consistency bias penalty and projection prediction error penalty, which are used to encourage the predicted value to approach the consistent subspace;

[0227] Label 11: Total Loss Summation Node – The above four losses are weighted and summed to form the optimization objective for end-to-end joint training;

[0228] Label 12: Backpropagation and Parameter Update Module – Synchronously updates the parameters of all networks using the Adam optimizer based on the total loss.

[0229] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. An infectious disease prediction method based on consistency constraints and zero-inflation perception, characterized in that, The steps include the following: Obtain historical incidence rate sequence data for multiple administrative levels and covariate data for each administrative level; Independent long short-term memory network encoders were constructed at the city, district, and street levels, respectively. Covariate data and historical incidence sequence data at each level were input into the corresponding level of the long short-term memory network encoder to extract the temporal features at each level. For the city and district levels, the predicted values ​​are generated by linear mapping based on the temporal characteristics of the corresponding levels. For the street level, the predicted values ​​are generated by a zero-inflation gating output structure based on the temporal characteristics of the street level. The zero-inflation gating output structure calculates in parallel the gating probability vector representing the probability of non-zero cases in each street and the intensity vector representing the expected number of non-zero cases. The street-level predicted values ​​are obtained by multiplying the gating probability vector and the intensity vector element by element. The city-level, district-level, and street-level predicted values ​​are concatenated into a unified prediction vector. Based on the unified prediction vector and the true value vector, the joint loss is calculated. The joint loss includes city-level prediction error loss, district-level prediction error loss, street-level prediction error loss, and soft consistency regularization loss used to penalize prediction values ​​that deviate from the hierarchical summation consistency constraint. Based on the joint loss, the parameters of the long short-term memory network encoder and the zero-inflation gated output structure at each level are synchronously updated through backpropagation until training is complete; The input data at each level at the time to be predicted are input into the trained long short-term memory network encoders and zero-inflation gating output structures at each level, and the prediction results at the city, district, and street levels are output simultaneously.

2. The infectious disease prediction method based on consistency constraints and zero-inflation perception according to claim 1, characterized in that, In the step of obtaining historical incidence rate sequence data for multiple administrative levels and covariate data for each administrative level: Set time The city-level observation value The district-level observation vector is The street-level observation vector is ,in, For the set of natural numbers, , The total number of district-level units, The total number of street-level units; For the street Its covariate vector is denoted as ,in, Dimensions for street-level features; City-level covariates are denoted as District-level covariates are denoted as ,in, and These represent the dimensions of city-level and district-level characteristics, respectively.

3. The infectious disease prediction method based on consistency constraints and zero-inflation perception according to claim 2, characterized in that, The step of constructing independent long short-term memory network encoders at the city, district, and street levels, and inputting the covariate data and historical incidence sequence data at each level into the corresponding level of the long short-term memory network encoder: For the street , covariate vector Compared with observed values The street is pieced together in time input vector ; All The streets in the time window to The input vectors are stacked according to time steps to obtain the street-level prediction sub-model at time step [time]. input matrix ,in, This refers to the length of the historical observation window; City-level input matrix and zone-level input matrix The covariates and observed values ​​at the city and district levels were constructed in the same manner.

4. The infectious disease prediction method based on consistency constraints and zero-inflation perception according to claim 3, characterized in that, The Long Short-Term Memory (LSTM) network encoders at each level include the following components: The first layer of the Long Short-Term Memory network has a hidden state dimension of 64 and returns the complete time series. The second layer is a Long Short-Term Memory network with a hidden state dimension of 32, which only returns the output of the last time step. Dropout layer with a dropout rate of 0.2; The first fully connected layer with an output dimension of 128 and using the ReLU activation function; The second fully connected layer has an output dimension of 64 and uses the ReLU activation function. Let the output vector of the second fully connected layer be... .

5. The infectious disease prediction method based on consistency constraints and zero-inflation perception according to claim 4, characterized in that, In the step of generating city-level and district-level predicted values ​​through linear mapping based on the time-series characteristics of the corresponding levels: City-level forecast values ​​are obtained through formulas Calculate, where, As a city-level weight, This is a city-level biased item; District-level predicted values ​​are obtained through the formula Calculate, where, This is a district-level weight matrix. This is the zone-level bias vector.

6. The infectious disease prediction method based on consistency constraints and zero-inflation perception according to claim 4, characterized in that, In the parallel computation of the gated probability vector and intensity vector of the zero-inflation gated output structure: Gating probability vector Through formula Calculate, where, The sigmoid function maps the gate probability to the interval (0,1). and These are the weights and biases of the gated branches, respectively; Intensity vector Through formula Calculate, where, To ensure that the output intensity is non-negative, and These represent the weights and biases of the strength branches, respectively. Street-level predicted values ​​are obtained through a formula Calculate, where, This represents the Hadamard product.

7. The infectious disease prediction method based on consistency constraints and zero-inflation perception according to claim 2, characterized in that, The soft consistency regularization loss is constructed as follows: Let the aggregation matrix be... Its first row contains all 1s, indicating that the total number of city-level values ​​equals the sum of the values ​​of all street-level values; for the... Each district , matrix number The elements of the row satisfy If and only if the street If it belongs to this district, then it is 0; Define constraint matrix ,in, for An identity matrix of order 1; Based on the constraint matrix Define the orthogonal projection matrix ; The soft consistency regularization loss Including the first and second items: The first item is the consistency deviation penalty, calculated as follows: This is used to penalize the degree to which the predicted value deviates from the hierarchical summation consistency constraint, where... This represents the sample size for a single training batch. The second item is the projection prediction error, which is calculated as follows: This is used to ensure that the model does not sacrifice prediction accuracy while satisfying the trend of consistency constraints.

8. The infectious disease prediction method based on consistency constraints and zero-inflation perception according to claim 2, characterized in that, The street-level prediction error loss Including the first, second, and third sub-terms of the weighted summation: The first sub-item is the principal loss, which uses the negative log-likelihood of the Tweedie distribution, and is expressed as: ,in, power parameter The range of values ​​is ; The second sub-item is the gated loss, which uses binary cross-entropy loss, labeled by whether the true value is zero. The calculation method is as follows: ,in, For the first The gating probability vector corresponding to each sample This is an indicator function; it takes the value 1 when the true value is greater than zero, and 0 otherwise. This is the gating loss weighting coefficient; The third sub-item is the non-zero sample amplitude loss, which calculates the mean squared error only for sample points where the true value is non-zero. The calculation method is as follows: ,in, For the first In the nth sample The intensity prediction value corresponding to each street To prevent division by zero constant, The amplitude loss weights are for non-zero samples.

9. The infectious disease prediction method based on consistency constraints and zero-inflation perception according to claim 1, characterized in that, The joint loss Including city-level forecast errors District-level prediction error Street-level prediction error and soft consistency regularization loss : Among them, the city-level forecast error The mean square error is used, and the calculation method is as follows: ; District-level prediction error The mean square error is used, and the calculation method is as follows: .

10. An infectious disease prediction system based on consistency constraints and zero-inflation perception, characterized in that, The system is used to execute the infectious disease prediction method based on consistency constraints and zero-inflation perception as described in any one of claims 1-9, the system comprising: The city-level long short-term memory network prediction module is used to receive city-level input data and generate city-level prediction values; The district-level long short-term memory network prediction module is used to receive district-level input data and generate district-level predicted values; A street-level long short-term memory network prediction module, comprising a zero-inflation gated output submodule, wherein the zero-inflation gated output submodule is used to compute in parallel a gated probability vector representing the probability of non-zero cases occurring in each street and an intensity vector representing the expected number of non-zero cases, and multiplies the gated probability vector and the intensity vector element by element to obtain the street-level prediction value; The hierarchical consistency constraint module is used to transform hierarchical summation constraints into soft consistency regularization loss; The joint training module is used to combine the city-level prediction error loss, district-level prediction error loss, street-level prediction error loss and the soft consistency regularization loss into a total loss, and synchronously update the parameters of all modules through backpropagation. The prediction output module is used to input the input data at each level at the time to be predicted into the trained modules and simultaneously output the prediction results at the city, district, and street levels.