A method for inferring time-varying positive serial interval distribution of infectious diseases
By constructing an infectious disease transmission chain model and inference matrix, the problem of accurate inference of the positive sequence interval distribution of infectious diseases is solved, accurate inference under different growth patterns is achieved, and it adapts to the actual infectious disease situation.
Patent Information
- Application Number
- CN202411854730.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2044-12-17
AI Technical Summary
Existing technologies make it difficult to accurately infer the forward sequence interval distribution of infectious diseases, especially in the early stages of infectious disease outbreaks, when some infected people do not develop the disease, making observation difficult and affecting the accuracy of key parameters.
Construct a basic transmission chain model of infectious diseases, infer the relationship between the forward serial interval distribution and the reverse serial interval distribution by establishing an inequality constraint on the effective reproduction number, construct an inference target, and use the inference matrix for inference.
It achieves accurate inference of the forward sequence interval distribution under different growth modes, especially the sub-exponential growth mode, reduces the forced assumptions on the growth mode, and conforms to the actual time-varying distribution.
Smart Images

Figure CN119694589B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of smart medical technology, and specifically relates to a method for inferring the interval distribution of time-varying positive sequences of infectious diseases. Background Art
[0002] In recent years, more and more infectious diseases have attracted people's attention. For example, after the outbreak of COVID-19, relevant key parameters and indicators of infectious diseases can provide reference information for the formulation of subsequent prevention and control measures, such as basic reproduction number, effective reproduction number, etc. Effective reproduction number, generation time interval, forward serial interval, backward serial interval, and incubation period. After infection, it may take a period of incubation for a patient to develop symptoms, which is also called "onset." In actual observation, it's difficult to directly determine when a patient becomes infected, but it's possible to directly determine when they develop symptoms.
[0003] In previous studies, it was necessary to give a preset growth pattern for the number of cases at each point in time. This could be given explicitly, for example, it is usually assumed to be exponential growth in the early stage; or it could be given implicitly, for example, a transmission mechanism of the infectious disease is assumed (such as SIR, SEIR, etc.). Recent studies have shown that in reality, due to contact patterns, spatial effects, inhomogeneous mixing, reactive behavior changes or other mechanisms, the early growth pattern of a large number of infectious diseases is not exponential growth, but subexponential growth. Different preset growth patterns may have a greater impact on the subsequent inference of related infectious disease parameters, such as the inference of the positive serial interval distribution. Basic reproduction number Key parameters of infectious diseases are the average number of secondary infections resulting from a primary infection, the effective reproduction number, and the rate of increase in the number of cases. These parameters are linked through the generation interval and incubation period. Because the generation interval represents the time delay between the point at which a transmitter becomes infected and the point at which a recipient becomes infected, it is difficult to directly observe. Therefore, the serial interval is often used as a direct proxy or inferred from the serial interval via the incubation period. Therefore, the serial interval is crucial in linking these core parameters or indicators, and using a positive serial interval distribution can more accurately assess relevant parameters of an infectious disease. However, in actual observational statistics, it is easier to obtain a negative serial interval distribution because this only requires searching for information about infected individuals at a point in time. In contrast, using a positive serial interval distribution can miss some transmission pairs because some recipients may not yet have developed symptoms. Therefore, it is crucial to rationally infer the positive serial interval distribution at the relevant points in time. Summary of the Invention
[0004] The problem to be solved by the present invention is to reasonably infer the positive sequence interval distribution at relevant time points, and to propose a method for inferring the time-varying positive sequence interval distribution of infectious diseases.
[0005] To achieve the above object, the present invention is implemented through the following technical solutions:
[0006] A method for inferring the time-varying forward sequence interval distribution of an infectious disease comprises the following steps:
[0007] S1. Construct a basic transmission chain model for infectious diseases;
[0008] S2. Based on the basic transmission chain model of infectious diseases constructed in step S1, establish the time point Inequality constraints on the effective reproduction number of ;
[0009] S3. Based on the time point obtained in step S2 The inequality constraint of the effective reproduction number is used to establish the relationship between the forward sequence interval distribution and the reverse sequence interval distribution;
[0010] S4. Constructing inference targets for the distribution of positive serial intervals;
[0011] S5. Inference Methods for Constructing Forward Serial Interval Distributions.
[0012] Furthermore, the specific implementation method of step S1 includes the following steps:
[0013] S1.1. Set the basic transmission status of the infectious disease to local transmission with no imported infections.
[0014] S1.2. Setting the time point in the transmission chain The number of nodes on , used to represent the number of patients with the disease. In the transmission pair, the initiator of infection is called the infector, and the patient infected by the infector is called the infected person. The infected person is the successor infected node of the infector, and the infector is the predecessor transmission node of the infected person.
[0015] S1.3. Let the maximum interval in the distribution of serial intervals of the transmission chain be , define the time point The nodes on the top have the attribute of complete observation, which means that they have the attribute of observation up to the time point The predecessor propagation node and the successor sensing node of each node;
[0016] S1.4. Assumption represents the effective time point length of the transmission chain, then the time point length of the transmission chain is ,set up It represents the maximum continuous observation time point of the transmission chain, completing the construction of the basic transmission chain model of infectious diseases.
[0017] Furthermore, the specific implementation method of step S2 includes the following steps:
[0018] S2.1. Consider The right time point The effective reproduction number Based on the basic transmission chain model of infectious diseases obtained in step S1, the infection starts at , then when at the time point Time, from the time point The node before must be part of a continuous transmission chain; The node is determined by The nodes are connected and are located at the time point The node is not an isolated node, so The nodes on must all be connected by edges; and The nodes on The out-degree edges on , that is, The edges sent from All nodes on have a unique predecessor propagation node, then the expression is:
[0019] (1)
[0020] in, For the i The number of nodes at a time point, i for t Any one of For the i The effective reproduction number at a point in time;
[0021] S2.2. Based on The outgoing edges do not need to make all the edges in All nodes have a unique predecessor propagation node, then:
[0022] (2)
[0023] Then , the number of nodes at each time point , the maximum interval in the serial interval distribution of the transmission chain as well as , satisfying the following inequality constraints:
[0024] (3)
[0025] S2.3. Further consider the following situation: when considering the time point on Time, time point on is known, ,have:
[0026] (4)
[0027] Therefore, we have:
[0028] (5)
[0029] in, for At the time The lower bound of the upper for At the time upper bound on;
[0030] The expression is:
[0031] (6)
[0032] The expression is:
[0033] (7)
[0034] S2.4. Based on Inequality constraints Make estimates;
[0035] Consider based on exist The above is known, it is estimated on , according to formula (5), when considering the time point When:
[0036] (8)
[0037] Based on the assignment method, take the ratio of the upper and lower bounds , which can be set to change over time or take random values, then:
[0038] (9)
[0039] in, Time point An estimate of the effective reproduction number;
[0040] Next, use calculate and , and then use the same method to obtain Repeat the above process until be estimated;
[0041] Therefore, ,have:
[0042] (10)
[0043] in, The length of the time delay between the onset of the disease in the transmitter and the onset of the disease in the corresponding recipient;
[0044] (11)
[0045] (12).
[0046] Furthermore, the specific implementation method of step S3 includes the following steps:
[0047] S3.1. Construction time The effective reproduction number , time point The total number of outgoing edges from the nodes on and timing The number of nodes on The relationship formula is:
[0048] (13);
[0049] S3.2. Definitions Indicates the starting point is from the time point , the end point is at time The total number of out-degree edges, set , then the construction time point Positive serial interval distribution It is defined by the following formula:
[0050] (14);
[0051] S3.3. The distribution of reverse serial intervals is the time point from the reverse perspective Past-oriented The angle definition of , then the time point is obtained Reverse serial interval distribution The formula is:
[0052] (15);
[0053] S3.4. According to 、 as well as The formula is , the following formula is given:
[0054] (16)
[0055] And because It is easy to derive from observations, At any point in time, it is known Depend on The inequality estimate is given, and the formula for the relationship between the forward serial interval distribution and the reverse serial interval distribution is:
[0056] (17).
[0057] Furthermore, the specific implementation method of step S4 includes the following steps:
[0058] S4.1. Assumption is the maximum continuous observation time point of the transmission chain, indicating the observed time interval The ability to propagate chains from the reverse perspective of all nodes;
[0059] S4.2. Assume that the node is at time The above is a fully observed property, then The relevant links on the above are known quantities, then at the time point on 、 as well as All are known quantities;
[0060] Then for , set in the time interval Get The effective part of for subsequent calculation or inference, so set:
[0061] (18)
[0062] Right now:
[0063] (19)
[0064] Since the considered propagation chain only provides The above has the property of complete observation, so Theoretically, the maximum ,get:
[0065] (20);
[0066] S4.3. Calculation exist The true value of
[0067] S4.4. Constructing a forward serial interval distribution where the inference target is a time interval Positive serial interval distribution on .
[0068] Furthermore, the specific implementation method of step S5 includes the following steps:
[0069] S5.1. Construct the inference matrix Inference Matrix, whose Row, No. The elements of the column represent the time points And the time delay length The forward sequence interval The value of and , set the time point and , so the inference matrix Inference Matrix has Line and Columns, the matrix size is ;
[0070] S5.2. Setting the time point Inference Matrix The time points represented by the rows are:
[0071] (twenty one)
[0072] Then when and When the following inequality is satisfied, the inference matrix Inference Matrix Row, No. The elements of the column have the capabilities inferred from formula (17):
[0073] (twenty two)
[0074] Right now:
[0075] (twenty three)
[0076] According to the above formula, we can get the first column of the first row of the inference matrix to the first column of the inference matrix. The element values of the column are calculated by formula (17);
[0077] like and To satisfy the above inequality, take 、 , then the first Row, No. The elements of the column are given by the following formula:
[0078] (twenty four);
[0079] S5.3. Each row of the inference matrix represents a probability distribution. , the following formula is given:
[0080] (25)
[0081] Then when 、 Time The value is given by formula (24), so the first row of The columns are given by the following formula:
[0082] (26);
[0083] S5.4. Assuming that the change in the distribution of positive serial intervals is continuous under the condition that the intensity of external intervention remains unchanged, then the current time point of Compared with before The differences between the average distributions of time should be as small as possible;
[0084] for Time point and , evenly distributed It is given by:
[0085] (27)
[0086] set up , then the inference matrix The status of the row Calculated at a point in time And the probability value in the forward serial interval distribution to be inferred The status is as follows:
[0087]
[0088] in, in In the above formula, the values are ;
[0089] S5.5. Define the difference between the two distributions obtained in step S5.4 for:
[0090] (28)
[0091] Then define the difference function for:
[0092] (29)
[0093] According to formula (25), there are the following constraints:
[0094] (30)
[0095] Then write the above expression as:
[0096] (31)
[0097] In summary, the inference method for constructing the positive serial interval distribution is to calculate the constraints under Minimize:
[0098] (32);
[0099] S5.6. Constructing Lagrangian Functions as follows:
[0100] (33)
[0101] in, is the Lagrange multiplier;
[0102] Right now:
[0103] (34)
[0104] Then, solve the following equations:
[0105] (35)
[0106] get:
[0107] (36)
[0108] The above formula is correct as well as All are established;
[0109] If it appears , then set the , and then recalculate the other value;
[0110] Then, the inference matrix is calculated by formulas (24) and (26) to satisfy Elements The value of , where the element The value is , and then calculate the element values at other positions using formula (36).
[0111] Beneficial effects of the present invention:
[0112] The method for inferring the time-varying forward sequence interval distribution of infectious diseases described in the present invention is similar to The growth pattern is not mandatory and can be assumed to be exponential, sub-exponential, or other growth patterns.
[0113] The method for inferring the time-varying forward sequence interval distribution of an infectious disease described in the present invention does not require the forward / reverse sequence interval distribution to be fixed, but rather a time-varying distribution that is closer to the actual situation, and does not require the pre-assumption of a certain type of special distribution.
[0114] The present invention provides a method for inferring the time-varying forward sequence interval distribution of infectious diseases. Over time Changing constraints ( Inequality), this constraint condition does not require too many additional assumptions, which can guide subsequent research based on the observed To reasonably presuppose .
[0115] The present invention provides a method for inferring the time-varying forward sequence interval distribution of an infectious disease, and proposes a method for inferring the forward sequence interval distribution, wherein the forward sequence interval distribution varies with time.
[0116] The present invention provides a method for inferring the time-varying forward sequence interval distribution of infectious diseases, focusing on a more reasonable In the growth mode (sub-exponential growth), the positive sequence interval distribution that changes with time is inferred. Through further deduction and exploration, the present invention There is no mandatory assumption on the growth model, which can be either exponential growth or sub-exponential growth, or other forms of growth model; and the positive sequence interval distribution of the present invention allows dynamic changes over time to be more in line with the actual situation. And it can be achieved by When the actual situation is sub-exponential growth, comparing the difference in the positive sequence interval distribution inferred under the assumption of exponential growth and the assumption of sub-exponential growth shows the necessity of reasonably constructing the initial growth model of infectious diseases. BRIEF DESCRIPTION OF THE DRAWINGS
[0117] Figure 1 This is a flow chart of a method for inferring time-varying forward sequence interval distribution of an infectious disease according to the present invention;
[0118] Figure 2 For the present invention Auxiliary schematic diagram of meaning;
[0119] Figure 3 For the present invention Schematic diagram of the impact that the last observation time point may have on the calculation of the forward serial interval distribution;
[0120] Figure 4 For the present invention Schematic diagram illustrating the reverse sequence interval perspective as an example;
[0121] Figure 5 For the present invention Schematic diagram illustrating the forward sequence interval perspective as an example;
[0122] Figure 6 For the present invention Schematic diagram of parameters or distributions that can be calculated when all observations are available at a given time point;
[0123] Figure 7 Schematic diagram of key time points of the present invention;
[0124] Figure 8 This is a schematic diagram of the time points at which the present invention can directly calculate the forward sequence interval distribution. DETAILED DESCRIPTION
[0125] In order to make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present invention and are not intended to limit the present invention. That is, the specific embodiments described herein are only some embodiments of the present invention, not all embodiments. Generally, the components of the specific embodiments of the present invention described and illustrated in the drawings herein can be arranged and designed in various different configurations, and the present invention can also have other embodiments.
[0126] Therefore, the following detailed description of the specific embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention as claimed, but is merely representative of selected specific embodiments of the present invention. All other specific embodiments obtained by those skilled in the art based on the specific embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0127] In order to further understand the content, features and effects of the present invention, the following specific embodiments are given as examples, and the attached Figure 1 -Attached Figure 8 The detailed instructions are as follows:
[0128] A method for inferring the time-varying forward sequence interval distribution of an infectious disease comprises the following steps:
[0129] S1. Construct a basic transmission chain model for infectious diseases;
[0130] Furthermore, the specific implementation method of step S1 includes the following steps:
[0131] S1.1. Set the basic transmission status of the infectious disease to local transmission with no imported infections.
[0132] S1.2. Setting the time point in the transmission chain The number of nodes on , used to represent the number of patients with the disease. In the transmission pair, the initiator of infection is called the infector, and the patient infected by the infector is called the infected person. The infected person is the successor infected node of the infector, and the infector is the predecessor transmission node of the infected person.
[0133] S1.3. Let the maximum interval in the distribution of serial intervals of the transmission chain be , define the time point The nodes on the top have the attribute of complete observation, which means that they have the attribute of observation up to the time point The predecessor propagation node and the successor sensing node of each node;
[0134] S1.4. Assumption represents the effective time point length of the transmission chain, then the time point length of the transmission chain is ,set up It represents the maximum continuous observation time point of the transmission chain, completing the construction of the basic transmission chain model of infectious diseases.
[0135] Furthermore, except for the node at the initial time point, each node at each time point must have a unique predecessor transmission node, but not necessarily a subsequent infection node (this means that the node does not transmit the infectious disease to other nodes / individuals). represents the effective time length of the transmission chain, then the time length involved in the transmission chain is , which means that in the transmission chain under consideration, the time interval The nodes of are "fully observable" and are in The predecessor propagation node of the node on must come from the time interval In addition, in the propagation chain under consideration, the nodes in The nodes on the network must not have subsequent infected nodes, because the focus is on the transmission chain. The infection and infected status of the previous node, and the extended time interval Just to ensure the timing The subsequent infected nodes of the nodes above will not be missed. represents the maximum continuous observation time point of the transmission chain, which means that we can know The accurate and unique predecessor propagation node of all nodes on the network, but the subsequent infection nodes of these nodes are not necessarily known. For example, a Subsequent infected nodes of the node on may appear At this point in time, the link will not be observed in this case.
[0136] S2. Based on the basic transmission chain model of infectious diseases constructed in step S1, establish the time point Inequality constraints on the effective reproduction number of ;
[0137] Furthermore, the specific implementation method of step S2 includes the following steps:
[0138] S2.1. Consider The right time point The effective reproduction number Based on the basic transmission chain model of infectious diseases obtained in step S1, the infection starts at , then when at the time point Time, from the time point The node before must be part of a continuous transmission chain; The node is determined by The nodes are connected and are located at the time point The node is not an isolated node, so The nodes on must all be connected by edges; and The nodes on The out-degree edges on , that is, The edges sent from All nodes on have a unique predecessor propagation node, then the expression is:
[0139] (1)
[0140] in, For the i The number of nodes at a time point, i for t Any one of For the i The effective reproduction number at a point in time;
[0141] Further, The outgoing edges do not need to make all the edges in Each node has a unique predecessor propagation node, this is because The nodes on the same node will also send out some edges to connect Nodes on
[0142] S2.2. Based on The outgoing edges do not need to make all the edges in All nodes have a unique predecessor propagation node, then:
[0143] (2)
[0144] when The outgoing edges from the For all nodes on , the equality sign on the right side of the above inequality holds.
[0145] Then , the number of nodes at each time point , the maximum interval in the serial interval distribution of the transmission chain as well as , satisfying the following inequality constraints:
[0146] (3)
[0147] S2.3. Further consider the following situation: when considering the time point on Time, time point on is known, ,have:
[0148] (4)
[0149] Therefore, we have:
[0150] (5)
[0151] in, for At the time The lower bound of the upper for At the time upper bound on;
[0152] The expression is:
[0153] (6)
[0154] The expression is:
[0155] (7)
[0156] Furthermore, when NPI (Non-Pharmaceutical Intervention) changes, Follow And change, The upper and lower bounds of also change accordingly. Inequality constraints can not only be used to test the current Does the value meet the numerical requirements and can guide Estimates.
[0157] S2.4. Based on Inequality constraints Make estimates;
[0158] Consider based on exist The above is known, it is estimated on , according to formula (5), when considering the time point When:
[0159] (8)
[0160] Based on the assignment method, take the ratio of the upper and lower bounds , then:
[0161] (9)
[0162] in, Time point An estimate of the effective reproduction number;
[0163] Next, use calculate and , and then use the same method to obtain Repeat the above process until be estimated;
[0164] Therefore, ,have:
[0165] (10)
[0166] in, The length of the time delay between the onset of the disease in the transmitter and the onset of the disease in the corresponding recipient;
[0167] (11)
[0168] (12).
[0169] S3. Based on the time point obtained in step S2 The inequality constraint of the effective reproduction number is used to establish the relationship between the forward sequence interval distribution and the reverse sequence interval distribution;
[0170] Furthermore, the specific implementation method of step S3 includes the following steps:
[0171] S3.1. Construction time The effective reproduction number , time point The total number of outgoing edges from the nodes on and timing The number of nodes on The relationship formula is:
[0172] (13);
[0173] S3.2. Definitions Indicates the starting point is from the time point , the end point is at time The total number of out-degree edges, set , then the construction time point Positive serial interval distribution It is defined by the following formula:
[0174] (14);
[0175] Forward Serial Interval Distribution is a distribution that stands at a point in time. Looking to the future The Backward Serial Interval Distribution is defined from the perspective of the reverse point of view. Past-oriented Obviously, since the minimum time point is , so when At the time and time point There is no directed edge between them, that is, Since the transmission chain is continuous, each The node must be an infected node, which means that The node at a certain time has only one predecessor node. The total number of in-degree edges of the previous node must be .
[0176] S3.3. The distribution of reverse serial intervals is the time point from the reverse perspective Past-oriented The angle definition of , then the time point is obtained Reverse serial interval distribution The formula is:
[0177] (15);
[0178] S3.4. According to 、 as well as The formula is , the following formula is given:
[0179] (16)
[0180] And because It is easy to derive from observations, At any point in time, it is known Depend on The inequality estimate is given, and the formula for the relationship between the forward serial interval distribution and the reverse serial interval distribution is:
[0181] (17).
[0182] S4. Constructing inference targets for the distribution of positive serial intervals;
[0183] It is necessary to clarify the specific target to be inferred is the forward serial interval distribution (ForwardSerialIntervalDistribution) over which time period. Since the transmission chain considered has a transmission time of (also referred to as effective time above) and the maximum sequence interval delay is assumed to be , so the maximum duration of the transmission chain is , which makes it possible to obtain superior and The true value of Therefore, we first need to clarify which time intervals The true value can be directly calculated, which time intervals Need to be inferred and have the possibility of inference, and which time intervals cannot be inferred.
[0184] set up is the maximum observation time, which indicates the time interval that can be observed The propagation chain of the reverse view of all nodes on Figure 2 ). This means that each The only predecessor infector node of the node above, but its successor infectee node may not be observed. The subsequent sensing node of the node at the time point may appear At this point in time, the link cannot be observed in this case (see Figure 3 ).
[0185] A node is called at a point in time The meaning of "fully observable" is: (1) from the reverse perspective, we know that the The predecessor propagation node of the node at the time point (see Figure 4 ); (2) Under the positive perspective, knowing that the The subsequent sensing node at the time point (see Figure 5 ).
[0186] Therefore, if the node is at time is fully observable (meaning that The relevant links on the above can be observed), then at time on 、 as well as can be calculated (see Figure 6 ).
[0187] Furthermore, the specific implementation method of step S4 includes the following steps:
[0188] S4.1. Assumption is the maximum continuous observation time point of the transmission chain, indicating the observed time interval The ability to propagate chains from the reverse perspective of all nodes;
[0189] S4.2. Assume that the node is at time The above is a fully observed property, then The relevant links on the above are known quantities, then at the time point on 、 as well as All are known quantities;
[0190] Then for , set in the time interval Get The effective part of for subsequent calculation or inference, so set:
[0191] (18)
[0192] Right now:
[0193] (19)
[0194] Since the considered propagation chain only provides The above has the property of complete observation, so Theoretically, the maximum ,get:
[0195] (20);
[0196] S4.3. Calculation exist The true value of
[0197] Furthermore, from a positive perspective, if we can observe the time point All out-degree edges of the node (considering at most ), you can calculate the time point on Value. ,So , which means that we can calculate exist The true value of Figure 8As shown;
[0198] On the other hand, the timing and time point , formula (17) gives and Since the growth pattern of the number of nodes over time is assumed to be known (for example, the growth of the number of nodes follows a sub-exponential growth), and Known. exist The value of can be calculated accurately and directly from the observed transmission chains, and exist can be estimated using formula (10-12). Therefore, for a specific and , to calculate Need to know The value of is recorded as:
[0199]
[0200] For example, Therefore, when making precise calculations, it is necessary to accurately calculate The value of At the point in time value. The maximum value that can be obtained is on Value. ,but , so it can be accurately calculated at most on value.
[0201] S4.4. Constructing a forward serial interval distribution where the inference target is a time interval Positive serial interval distribution on .
[0202] Furthermore, from a macro perspective The inferential boundary of Initially, no links or infection information between nodes can be observed, only information about and information;
[0203] Obviously, for a given and There are many possible ways to connect nodes, and different connection methods will lead to different Get the value.
[0204] From a positive perspective, the timing and time point The node link between them is the last time point with link or infection information, which cannot be observed. and The link or infection information between them is inferred to be the time interval Positive serial interval distribution on .
[0205] S5. Inference Methods for Constructing Forward Serial Interval Distributions.
[0206] Furthermore, the specific implementation method of step S5 includes the following steps:
[0207] S5.1. Construct the inference matrix Inference Matrix, whose Row, No. The elements of the column represent the time points And the time delay length The forward sequence interval The value of and , set the time point and , so the inference matrix Inference Matrix has Line and Columns, the matrix size is ;
[0208] S5.2. Setting the time point Inference Matrix The time points represented by the rows are:
[0209] (twenty one)
[0210] Then when and When the following inequality is satisfied, the inference matrix Inference Matrix Row, No. The elements of the column have the capabilities inferred from formula (17):
[0211] (twenty two)
[0212] Right now:
[0213] (twenty three)
[0214] According to the above formula, we can get the first column of the first row of the inference matrix to the first column of the inference matrix. The element values of the column are calculated by formula (17);
[0215] like and To satisfy the above inequality, take 、 , then the first Row, No. The elements of the column are given by the following formula:
[0216] (twenty four);
[0217] S5.3. Each row of the inference matrix represents a probability distribution. , the following formula is given:
[0218] (25)
[0219] Then when 、 Time The value is given by formula (24), so the first row of The columns are given by the following formula:
[0220] (26);
[0221] Next, additional assumptions need to be added to move the inference process forward. Assume that the intensity of external intervention remains unchanged (for example, NPI does not change dramatically, vaccine coverage does not change significantly, vaccine efficacy does not change significantly, etc.);
[0222] S5.4. Assuming that the change in the distribution of positive serial intervals is continuous under the condition that the intensity of external intervention remains unchanged, then the current time point of Compared with before The differences between the average distributions of time should be as small as possible;
[0223] for Time point and , evenly distributed It is given by:
[0224] (27)
[0225] set up , then the inference matrix The status of the row Calculated at a point in time And the probability value in the forward serial interval distribution to be inferred The status is as follows:
[0226]
[0227] in, in In the above formula, the values are ;
[0228] S5.5. Define the difference between the two distributions obtained in step S5.4 for:
[0229] (28)
[0230] Then define the difference function for:
[0231] (29)
[0232] According to formula (25), there are the following constraints:
[0233] (30)
[0234] Then write the above expression as:
[0235] (31)
[0236] In summary, the inference method for constructing the positive serial interval distribution is to calculate the constraints under Minimize:
[0237] (32);
[0238] S5.6. Constructing Lagrangian Functions as follows:
[0239] (33)
[0240] in, is the Lagrange multiplier;
[0241] Right now:
[0242] (34)
[0243] Then, solve the following equations:
[0244] (35)
[0245] get:
[0246] (36)
[0247] The above formula is correct as well as All are established;
[0248] If it appears , then set the , and then recalculate the other value;
[0249] Then, the inference matrix is calculated by formulas (24) and (26) to satisfy Elements The value of , where the element The value is , and then calculate the element values at other positions using formula (36).
[0250] It should be noted that relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.
[0251] Although the present application has been described above with reference to specific embodiments, various modifications may be made thereto and components may be substituted with equivalents without departing from the scope of the present application. In particular, as long as there are no structural conflicts, the various features of the embodiments disclosed herein may be combined with each other in any manner, and the omission of an exhaustive description of these combinations in this specification is solely for the sake of space and resource conservation. Therefore, the present application is not limited to the specific embodiments disclosed herein, but includes all technical solutions within the scope of the claims.
Claims
1. A method for inferring the time-varying forward sequence interval distribution of infectious diseases, characterized by: The steps include: S1. Construct a basic transmission chain model for infectious diseases; The specific implementation method of step S1 includes the following steps: S1.
1. Set the basic transmission status of the infectious disease to local transmission with no imported infections; S1.
2. Let P(t) be the number of nodes at time t in the transmission chain, representing the number of patients with the disease. In a transmission pair, the initiator of infection is called the infector, and the patient infected by the infector is called the infected person. The infected person is the successor of the infector, and the infector is the predecessor of the infectee. S1.
3. Let the maximum interval in the serial interval distribution of the transmission chain be τ max , defines that the node at time point t has the property of complete observation, which means that it has the predecessor propagation node and the successor sensing node that observe every node at time point t; S1.
4. Assume t e Represents the effective time point length of the transmission chain, then the time point length of the transmission chain is t e +τ max , let t o Indicates the maximum continuous observation time point of the transmission chain, completing the construction of the basic transmission chain model of infectious diseases; S2. Based on the basic transmission chain model of the infectious disease constructed in step S1, establish an inequality constraint on the effective reproduction number at time point t; The specific implementation method of step S2 includes the following steps: S2.
1. Consider The above inequality constraint on the effective reproduction number R(t) at time point t is based on the basic transmission chain model of infectious diseases obtained in step S1. Assuming that the infection starts at t=1, then when at time point t, the nodes from before time point t-1 must be part of the continuous transmission chain; the node at time point t is definitely connected by the node from [1, t-1] and the node at time point t is not an isolated node, so the nodes on [1, t] must all be connected by edges; and the nodes on t+1 must be connected by the out-degree edges on [1, t], that is, the edges sent from [1, t] must ensure that all nodes on [2, t+1] have a unique predecessor transmission node, then the expression is: Where P(i) is the number of nodes at the i-th time point, i is any one of t, and R(i) is the effective regeneration number at the i-th time point; S2.
2. The outgoing edges from [1, t] do not need to make all the edges in [t+1, t+τ max ] nodes have a unique predecessor propagation node, then: Then The number of nodes at each time point P(t), the maximum interval τ in the serial interval distribution of the propagation chain max and R(t), satisfying the following inequality constraints: S2.
3. Further consider the following situation: when considering R(t) at time point t, R(·) at time points 1, 2, ..., t-1 is known, have: Therefore, we have: l Rt (t)≤R(t)≤u Rt (t) (5) Among them, l Rt (t) is the lower bound of R(t) at time point t, u Rt (t) is the upper bound of R(t) at time point t; l Rt The expression of (t) is: u Rt The expression of (t) is: S2.
4. Estimate R(t) based on the R(t) inequality constraints; Considering that R(t) is known on [1,t], estimate [t+1,t+τ max ], according to formula (5), when considering the time point t+1, we have: l Rt (t+1)≤R(t+1)≤u Rt (t+1) (8) Based on the assignment method, take the ratio of the upper and lower bounds θ∈[0,1], then: in, is the estimated value of the effective reproduction number at time point t+1; Next, using R(1), R(2), ..., R(t), Calculate l Rt (t+2) and u Rt (t+2), and then use the same method to obtain Repeat the above process until be estimated; Therefore, have: Where τ is the time delay between the onset time of the transmitter and the onset time of the corresponding recipient; S3. Based on the inequality constraint of the effective reproduction number at time point t obtained in step S2, establish the relationship between the forward sequence interval distribution and the reverse sequence interval distribution; The specific implementation method of step S3 includes the following steps: S3.
1. The relationship between the effective reproduction number R(t) at time point t, the total number of outgoing edges e(t) from the nodes at time point t, and the number of nodes P(t) at time point t is: S3.
2. Define e(t,t+τ) to represent the total number of out-degree edges starting from time point t and ending at time point t+τ. Set τ>0, and construct the forward sequence interval distribution f at time point t. t (τ) is defined by the following formula: S3.
3. The reverse serial interval distribution is defined as the angle from the reverse perspective time point s to the past time point s-τ, and the reverse serial interval distribution b at time point t is obtained. s The formula for (τ) is: S3.
4. According to R(t), f t (τ) and b s The formula for (τ) is s = t + τ, which gives the following formula: f t (τ)×R(t)×P(t)=b t+τ (τ)×P(t+τ) (16) And because b t+τ (τ) is easy to obtain from observations, P(t) is known at any time point, R(t) is estimated by the R(t) inequality, and the formula for the relationship between the forward serial interval distribution and the reverse serial interval distribution is: S4. Construct an inference target for the positive serial interval distribution; The specific implementation method of step S4 includes the following steps: S4.
1. Assume t o is the maximum continuous observation time point of the transmission chain, indicating that the observation time interval t∈[1,t o ]The ability to propagate chains from the reverse perspective of all nodes on the network; S4.
2. Assuming that the node has the fully observed attributes at time point t, then in [t-τ max ,t+τ max ] is a known quantity, then R(t) and f at time point t t (τ) and b t (τ) are all known quantities; Then for t o , set in the time interval [1,t o ] get b t The effective part of (τ) is used for subsequent calculation or inference, so it is set as follows: t o -t max ≥1 (18) Right now: t o ≥τ max +1 (19) Since the considered propagation chain only provides e ] has the property of complete observation, so t o Theoretically, the maximum value is t o =t e ,get: t o ∈[1+τ max ,t e ](20); S4.
3. Calculation of f t (τ) in [1,t o -τ max ] on the true value; S4.
4. The inference target for constructing a forward serial interval distribution is the time interval [t o -τ max +1,t o -1] on the positive serial interval distribution f t (τ); S5. Constructing an inference method for the distribution of positive serial intervals; The specific implementation method of step S5 includes the following steps: S5.
1. Construct an inference matrix, whose elements in the mth row and nth column represent the time point t=t o -τ max +m and time delay length τ = n t (τ), where t = t o -τ max +m and τ=n, set the time point t∈[t o -τ max +1,t o -1] and τ∈{1,2,...,τ max }, so the inference matrix InferenceMatrix has τ max -1 row and τ max Column, the matrix size is (τ max -1)×τ max ; S5.
2. Assume time t m Indicates the time point represented by the mth row of the inference matrix InferenceMatrix, that is: t m =t o -t max +m, m=1,2,...,τ max -1 (21) Then, when m and n satisfy the following inequality, the elements in the mth row and nth column of the inference matrix have the ability inferred by formula (17): t m +n=t o -τ max +m+n≤t o (22) Right now: m+n≤τ max (23) According to the above formula, we can get the first column of the first row of the inference matrix to the τth column. max The element values of the -1 column are calculated by formula (17); If m and n satisfy the above inequality, take t = t m =t o -τ max +m, τ=n, then the element in the mth row and nth column of the inference matrix that meets the constraints is given by the following formula: S5.
3. Each row of the inference matrix represents a probability distribution. For m∈{1,2,...,τ max -1}, the following formula: Then when m=1、n∈{1,2,...,τ max -1} The value is given by formula (24), so the τth in the first row is max The columns are given by the following formula: S5.
4. Assuming that the change in the distribution of positive sequence intervals is continuous under the condition that the intensity of external intervention remains unchanged, then at the current time point t m of The difference from the average distribution of the previous T time should be as small as possible; For t m Time point and n∈{1,2,...,τ max }, average distribution a(t m ,n) is given by: Set m∈{2,3,...,τ max -1}, then the state of the mth row of the inferred matrix is the same as t m a(t m ,n) and the probability value x(t m ,n) is as follows: a(t m ,1),...,a(t m ,τ max -m),a(t m ,τ max -m+1),...,a(t m ,τ max ); Among them, x(t m ,n) in the above formulas are respectively taken as n=τ max -m+1,τ max -m+2,...,τ max ; S5.
5. Define the difference Δ between the two distributions obtained in step S5.4 as: Then define the difference function for: According to formula (25), there are the following constraints: Then write the above expression as: In summary, the inference method for constructing the positive serial interval distribution is to calculate the constraints under Minimize: S5.
6. Constructing Lagrangian Functions as follows: in, is the Lagrange multiplier; Right now: Then, solve the following equations: get: The above formula is correct as well as All are established; If x(t m ,n)<0, then set the x(t m ,n)=0, and then recalculate other x(t m ,n) value; Then, the inference matrix Inference Matrix is calculated by formulas (24) and (26) to satisfy m+n≤τ max The value of the element (m,n) of , and then calculate the element values at other positions using formula (36).
Citation Information
Patent Citations
Hand-foot-and-mouth disease prediction method using fine-grained data, electronic equipment, and medium
CN111863276A
SIS model availability evaluation method and system based on discrete data
CN115376703A