Prediction method and prediction device of the number of infection-positive persons

The method uses moving and weighted geometric averaging to filter noise from sewage pathogen data, ensuring accurate prediction of infectious disease case trends by aligning with data distribution characteristics.

JP2025159745APending Publication Date: 2025-10-22KUBOTA CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024062455
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-09
Publication Date
2025-10-22

AI Technical Summary

Technical Problem

Conventional methods for predicting infectious disease cases based on sewage pathogen data are inaccurate due to noise and other factors, leading to unreliable trend predictions.

Method used

A method involving moving geometric averaging and weighted geometric averaging of sewage pathogen data, combined with inflow water volume weighting, to filter noise and accurately predict case trends.

Benefits of technology

The method effectively filters noise and other factors, enabling accurate prediction of infectious disease case trends by aligning with data distribution characteristics, thus improving prediction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025159745000001_ABST
    Figure 2025159745000001_ABST
Patent Text Reader

Abstract

To provide a prediction method of the number of infection-positive persons which can accurately predict a tendency to increase or decrease the number of infection-positive persons based on information on infection disease pathogens present in sewage.SOLUTION: A prediction method 1 of the number of infection-positive persons includes data acquisition step S1, moving geometric averaging step S4, and prediction step S5. In the data acquisition step S1, observation data, which is information on infection disease pathogens in water of a water treatment plant, is acquired as a time-series data string. In the moving geometric averaging step S4, a data string is accepted as an input data string, and the input data string is subjected to moving geometric averaging processing to calculate a moving average data string. In the prediction step S5, a tendency to increase or decrease the number of positive persons is predicted using the number of positive persons infected by the infection disease pathogen and the moving average data string. In the moving geometric averaging processing, an input data string is geometrically averaged while an average period is moved during a predetermined average period in the input data string.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method for predicting the number of infected people and a device for predicting the number of infected people. [Background technology]

[0002] To prevent the spread of infectious diseases such as the novel coronavirus, it is necessary to understand the increasing or decreasing trend in the number of people who test positive for the infection, monitor the infection situation, and then take appropriate measures. However, conventional methods of understanding the number of people who test positive using PCR tests, antigen tests, etc., have issues such as delays in reporting test results, making it time-consuming to understand the infection situation. In response to this issue, for example, the infectious disease testing method described in Patent Document 1 acquires information regarding the presence of infectious pathogens in sewage discharged from a facility. From the acquired information, the possibility that there are people who test positive for the infection among those staying at the facility is indicated, and the increasing or decreasing trend in the number of people who test positive is predicted. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent No. 7031957 Summary of the Invention [Problem to be solved by the invention]

[0004] However, the testing method described in Patent Document 1 may not accurately indicate the possibility of the presence of a person who is positive for an infectious disease if unnecessary information such as noise is included in the acquired information regarding the presence of an infectious pathogen. As a result, the testing method described in Patent Document 1 may not accurately predict the increase or decrease in the number of positive cases.

[0005] The present invention has been made in consideration of the above-mentioned problems, and aims to provide a method for predicting the number of infectious disease-positive cases that can accurately predict the increase or decrease in the number of infectious disease-positive cases based on information about infectious disease pathogens present in sewage. [Means for solving the problem]

[0006] According to one aspect of the present invention, a method for predicting the number of infected people includes a data acquisition step of acquiring observation data, which is information about infectious disease pathogens in water at a water treatment plant, as a time-series data string; a moving geometric averaging step of receiving the data string as an input data string and performing a moving geometric averaging process on the input data string to calculate a moving average data string; a prediction step of predicting an increase / decrease trend in the number of positive cases of an infection caused by the infectious disease pathogen using the number of positive cases and the moving average data string; Equipped with In the moving geometric mean process, the input data string is subjected to the geometric mean process while the averaging period is moved within a predetermined averaging period for the input data string.

[0007] According to this, in the method for predicting the number of infected people, a data string is subjected to a moving geometric mean process, thereby performing an averaging process suited to the distribution characteristics of the information in the data string (log-normal distribution). In other words, in the method for predicting the number of infected people, the effects of noise and other factors contained in the data string are appropriately eliminated, and the data string is averaged. Therefore, the method for predicting the number of infected people can appropriately grasp the relationship between the averaged data string and the number of infected people.

[0008] The method for predicting the number of infected positive cases according to the second invention further comprises a weighted geometric mean step in which the data string is subjected to a weighted geometric mean process, In the data acquisition step, the data strings for each of the plurality of water treatment plants and weighted values ​​associated with the catchment population for each of the water treatment plants are acquired; In the weighted geometric mean step, the data strings weighted by the respective weight values ​​are subjected to a weighted geometric mean process to calculate a weighted average data string; In the moving geometric mean step, the weighted average data string is received as the input data string and the moving geometric mean process is performed.

[0009] According to this, in the method for predicting the number of positive cases of infectious diseases, a weighted geometric mean is applied to a data sequence using a weight value related to the basin population, thereby performing a weighted mean process that is appropriate for the distribution characteristics of the information in the data sequence (log-normal distribution). In other words, in the method for predicting the number of positive cases of infectious diseases, the influence of noise and other factors contained in the data sequence can be appropriately eliminated, and the data sequence can be appropriately weighted using a weight value. Therefore, the method for predicting the number of positive cases of infectious diseases can appropriately grasp the relationship between the weighted average data sequence and the number of positive cases.

[0010] The method for predicting the number of infected positive cases according to the third invention further comprises an averaging period determination step for determining the averaging period, In the averaging period determination step, the moving average data sequence corresponding to each of the averaging periods is calculated using the averaging period as a variable, In the averaging period determination step, a correlation coefficient indicating the degree of correlation between the processed data subjected to the moving geometric average process and the number of positive cases, which is included in each of the moving average data strings, is calculated, and The averaging period is determined based on the value of the correlation coefficient.

[0011] According to this, the method for predicting the number of infected people optimizes the averaging period so that the correlation coefficient between the processed data and the number of infected people becomes large.

[0012] In the method for predicting the number of infectious disease positive cases according to the fourth invention, the weighted value is the amount of inflow water at each of the water treatment plants.

[0013] According to this, the method for predicting the number of positive cases of infectious diseases assigns appropriate weighting to data sequences with large basin populations by taking a weighted average using the inflow water volume as a weighted value. Therefore, the method for predicting the number of positive cases of infectious diseases can appropriately grasp the trends in the values ​​of information that appear frequently in the data sequence.

[0014] In the method for predicting the number of infectious disease positive cases according to the fifth invention, in the prediction step, a regression line showing the relationship between the number of positive cases and the observation data processed by the moving geometric average in the moving average data sequence is calculated, and the number of positive cases for specific observation data is predicted as the increase / decrease trend using the regression line.

[0015] According to this, the method for predicting the number of infected people appropriately grasps the relationship between the number of infected people and the observed data using the regression line. Therefore, the method for predicting the number of infected people can predict the number of infected people as an increase or decrease trend for specific observed data using the regression line.

[0016] The sixth aspect of the present invention provides a predicting device for the number of infected people who test positive for infectious diseases, comprising: a data acquiring unit that acquires observation data, which is information about infectious disease pathogens in water at a water treatment plant, as a time-series data string; a moving geometric average unit that receives the data string as an input data string and performs a moving geometric average process on the input data string to calculate a moving average data string; a prediction unit that predicts an increase or decrease trend in the number of positive cases of an infection caused by the infectious disease pathogen using the number of positive cases and the moving average data string; Equipped with In the moving geometric mean process, the input data string is subjected to the geometric mean process while the averaging period is moved within a predetermined averaging period for the input data string.

[0017] According to this, the device for predicting the number of infected positive cases performs an averaging process suited to the distribution characteristics of information in the data sequence (log-normal distribution) by performing a moving geometric mean process on the data sequence. In other words, the device for predicting the number of infected positive cases averages the data sequence by appropriately eliminating the influence of noise and other factors contained in the data sequence. Therefore, the device for predicting the number of infected positive cases can appropriately grasp the relationship between the averaged data sequence and the number of infected cases. [Effects of the Invention]

[0018] According to the method for predicting the number of infectious disease positive cases of the present invention, it is possible to accurately predict the increase or decrease in the number of infectious disease positive cases based on information about infectious disease pathogens present in sewage. [Brief explanation of the drawings]

[0019] [Figure 1] 1 is a flow chart showing the processing steps in a method for predicting the number of infected people according to an embodiment of the present invention. [Figure 2] This is a diagram showing an example of a distribution curve of virus concentration in observation data acquired in the data acquisition process of the method for predicting the number of infected people. [Figure 3] FIG. 1 is a flow chart showing an example of processing in each step of the method for predicting the number of infected people. [Figure 4] FIG. 10 is a flow chart showing an example of processing in the average period determination step of the method for predicting the number of infected person positive cases. [Figure 5] This is a flow chart showing a generalized process in each step of the method for predicting the number of infected people. [Figure 6] FIG. 10 is a diagram showing an example of the processing results obtained by the method for predicting the number of infected people. [Figure 7] FIG. 10 is a diagram showing an example of the processing results obtained by the method for predicting the number of infected people. [Figure 8] FIG. 10 is a diagram showing an example of the processing results obtained by the method for predicting the number of infected people. [Figure 9] FIG. 1 is a block diagram of a predicting device for the number of infected people that executes the method for predicting the number of infected people. DETAILED DESCRIPTION OF THE INVENTION

[0020] Hereinafter, a method 1 for predicting the number of infected people who have tested positive for an infectious disease according to an embodiment of the present invention will be described with reference to the drawings. In the drawings, the same or corresponding parts are denoted by the same reference symbols, and description thereof will not be repeated.

[0021] [Overview of each process] The processing steps in a method 1 for predicting the number of infected people according to an embodiment of the present invention will be described with reference to Figure 1. Figure 1 is a flow diagram showing the processing steps in a method 1 for predicting the number of infected people. As will be described in detail later, as shown in Figure 1, a method 1 for predicting the number of infected people includes a data acquisition step S1, a weighted geometric mean step S2, an averaging period determination step S3, a moving geometric mean step S4, and a prediction step S5.

[0022] In the data acquisition process S1, observation data of virus concentration (information about infectious disease pathogens, an example of observation data) in sewage from a sewage treatment plant P (an example of a water treatment plant) is acquired as a time-series data string x. In this embodiment, the observation data are virus concentration values, and multiple pieces of data are acquired at regular time intervals (sampling times). The observation data are arranged in time series to form the data string x. Viruses to be observed include, for example, the novel coronavirus, influenza virus, norovirus, etc. The virus concentration is, for example, the number of copies per unit volume. The copy number is the number of genes of the target virus contained in the sewage. The observation data may be information about the presence of viruses other than virus concentration, such as the virus detection rate.

[0023] [Regarding virus concentration observation data] Next, the observation data of virus concentration acquired in the data acquisition step S1 will be described with reference to Fig. 2. Fig. 2 is an example of a distribution curve of virus concentration in the acquired observation data. Specifically, Fig. 2 shows a probability density function as a distribution curve of virus concentration estimated from the acquired observation data. In Fig. 2, the horizontal axis represents a random variable, which is the value of virus concentration in the observation data. The vertical axis represents the probability (probability density) that the random variable (virus concentration) takes on that value.

[0024] A distribution curve is found by estimating the parameters of a probability density function based on observed data. The parameters are estimated, for example, using maximum likelihood estimation. In maximum likelihood estimation, a likelihood function is defined as a function that represents the probability that observed data on virus concentration will result from a specific distribution curve. The parameter value that maximizes the likelihood function is then found as the estimated value. The distribution curve in Figure 2 follows a lognormal distribution. A lognormal distribution is a distribution curve in which the logarithm of a random variable, i.e., the probability corresponding to the logarithm of the virus concentration on the horizontal axis, follows a normal distribution.

[0025] The log-normal distribution has an asymmetric shape with respect to the random variable, and is characterized by the fact that the distribution remains with a certain probability even when the random variable takes a large value (has a long tail on the right). The reason why the virus concentration in sewage follows a log-normal distribution is thought to be because, for example, unexpected sewage inflow at sewage treatment plant P can cause the virus concentration to suddenly increase. In other words, it is thought that the observation data with high virus concentrations is included as noise due to unexpected sewage inflow, etc.

[0026] The observed data on virus concentrations in sewage, which is the subject of processing in Method 1 for predicting the number of positive cases of infectious diseases, is sometimes called sewage surveillance data (sewage epidemiological survey data) that tests and monitors viruses in sewage. Sewage surveillance data is expected to be used as an objective indicator that can reflect the infection status, including asymptomatic infected individuals, without being influenced by factors such as the medical visit behavior of test subjects or the number of tests. In order to accurately grasp the infection status using sewage surveillance data, it is effective to use it while minimizing the influence of noise, as explained above.

[0027] [About the processing at each stage] Next, a specific example of processing in each step of the method 1 for predicting the number of infected people will be described with reference to Fig. 3. Fig. 3 is a flow chart showing an example of processing in each step of the method 1 for predicting the number of infected people.

[0028] As shown in Figure 3, method 1 for predicting the number of infectious disease positive cases predicts the number of positive cases by acquiring observation data from two sewage treatment plants P: a first sewage treatment plant P1 and a second sewage treatment plant P2. To repeat, in the data acquisition step S1, observation data on the virus concentration in the sewage from the sewage treatment plant P is acquired as a time-series data string x.

[0029] If there are multiple sewage treatment plants P, in the data acquisition step S1, data strings x are acquired from each of the multiple sewage treatment plants P. Multiple pieces of observation data included in each data string x are acquired at regular time intervals (sampling times). In this embodiment, the data strings x acquired from the multiple sewage treatment plants P are acquired at equal sampling times.

[0030] 3, first, in the data acquisition step S1, a first data sequence x1(t) is acquired from the first sewage treatment plant P1, and a second data sequence x2(t) is acquired from the second sewage treatment plant P2. The first data sequence x1(t) is composed of observation data of virus concentration at the first sewage treatment plant P1 at time t, arranged in chronological order. The second data sequence x2(t) is composed of observation data of virus concentration at the second sewage treatment plant P2 at time t, arranged in chronological order.

[0031] In the data acquisition step S1, the inflow water volume w is acquired as an example of a weighted value associated with the catchment population of each sewage treatment plant P1, P2. The inflow water volume w is, for example, the inflow water volume per day, [m 3 3, in the data acquisition step S1, a first influent flow rate w1 at the first sewage treatment plant P1 and a second influent flow rate w2 at the second sewage treatment plant P2 are acquired.

[0032] Next, in the weighted geometric mean step S2, the data strings x weighted by the inflow water volumes w1 and w2 are weighted to obtain a weighted mean.

[0033] The weighted average takes into account the importance of data string x and grasps the tendency of virus concentrations that are frequently observed in data string x (hereinafter referred to as "central tendency"). Data string x at sewage treatment plant P with a large basin population has a large contribution to the overall number of positive cases, and therefore has a high importance. Sewage treatment plant P with a large basin population has a large inflow water volume w. Therefore, data string x at sewage treatment plant P with a large inflow water volume w has a high importance. By performing a weighted average using the inflow water volume w as a weighting value, method 1 for predicting the number of positive infection cases assigns an appropriate weight to data string x with a large basin population. Therefore, method 1 for predicting the number of positive infection cases can appropriately grasp the central tendency in data string x.

[0034] In the weighted geometric mean step S2, each data sequence x is subjected to weighted geometric mean processing to calculate a weighted average data sequence xaw(t). Specifically, in the weighted geometric mean step S2, the weighted average data sequence xaw(t) is calculated by performing weighted geometric mean processing as shown in equation F1 in Figure 3. That is, as shown in equation F1, the weighted average data sequence xaw(t) is obtained by multiplying the first data sequence x1(t) raised to the power of the ratio of the first inflow water volume w1 to the integrated inflow water volume W by the second data sequence x2(t) raised to the power of the ratio of the second inflow water volume w2 to the integrated inflow water volume W.

[0035] Here, the cumulative inflow water volume W is the sum of the first inflow water volume w1 and the second inflow water volume w2, as shown in formula F2. Specifically, the weighted geometric mean process is performed on each of the observation data included in each of the first data sequence x1(t) and the second data sequence x2(t), and the weighted average processed observation data (hereinafter simply referred to as "processed data") is calculated. In other words, the weighted average data sequence xaw(t) is composed of the processed data arranged in chronological order.

[0036] Next, in a moving geometric averaging step S4, the weighted average data sequence xaw(t) is accepted as an input data sequence xi(t), and the input data sequence xi(t) is averaged using a moving average. In the moving averaging step, the input data sequence xi(t) is subjected to a moving geometric averaging process to calculate a moving average data sequence xam(t). In the moving geometric averaging process, the input data sequence xi(t) is geometrically averaged while the averaging period is shifted within a predetermined averaging period that includes a plurality of processed data.

[0037] In the averaging period determination step S3, the number of data to be averaged by moving average is determined as the number of averaged data N (details will be described later). In Fig. 3, the number of averaged data N included in the averaging period is, for example, 6.

[0038] Specifically, in the moving geometric mean step S4, the input data sequence xi(t) is subjected to moving geometric mean processing as shown in equations F3 and F4. As shown in equation F3, for example, in the moving geometric mean processing for processed data x(1) at time t=1 in the input data sequence xi(t), the processed data x(1) at time t=1 through the processed data x(6) at time t=6 are multiplied. The result of the multiplication is the radical root of the number of data to be averaged N (N=6). Similarly, in the moving geometric mean processing for processed data x(2) at time t=2 in the input data sequence xi(t), the processed data x(2) at time t=2 through the processed data x(7) at time t=7 are multiplied. The result of the multiplication is the radical root of the number of data to be averaged N (N=6). The same process is performed on all processed data that can be subjected to moving geometric mean processing among the processed data included in the input data sequence xi(t), and the moving average data sequence xam(t) is calculated.

[0039] Finally, in the prediction step S5, a prediction process is performed in which the number of positive cases for a specific virus concentration (an example of an increase / decrease trend in the number of positive cases) is predicted using the number of positive cases of the infectious disease and the moving average data sequence xam(t).

[0040] Specifically, in the prediction step S5, a data sequence y(t) of the number of positive cases (e.g., the number of positive cases in the treatment area of ​​a sewage treatment plant P) corresponding to the calculated moving average data sequence xam(t) is acquired. The data sequence y(t) of the number of positive cases is time-series data of the number of positive cases, and is acquired from a database or the like (e.g., a database managed by an administrative agency). The data of the number of positive cases included in the data sequence y(t) of the number of positive cases is acquired, for example, at a sampling time corresponding to the moving average data sequence xam(t). In the prediction step S5, a correlation graph G1 such as that shown in FIG. 3 is created. In the correlation graph G1, for example, the X-axis (horizontal axis) represents the virus concentration, and the Y-axis (vertical axis) represents the number of virus-positive cases.

[0041] Although not shown in the figure, correlation graph G1 plots processed data (hereinafter simply referred to as "processed data") that has undergone moving average processing and the data on the number of positive cases in the moving average data sequence xam(t). In correlation graph G1, the processed data and the data on the number of positive cases at the same time are plotted as point data. The point data is plotted, for example, with the value of the processed data as the X value and the value of the data on the number of positive cases as the Y value.

[0042] In the prediction step S5, a regression line L is obtained for the set of point data for each time plotted on the correlation graph G1. The regression line L is obtained, for example, by the least squares method. The regression line L may also be obtained, for example, by various regression methods using machine learning. In the regression line L, the number of positive cases, which is the Y-axis, is expressed as a function of the virus concentration, which is the X-axis. Then, in the prediction step S5, the number of positive cases y1 for a specific virus concentration Xa1 is predicted based on the regression line L. Also in the prediction step S5, a correlation coefficient C is obtained for the set of point data for each time. The correlation coefficient C is obtained, for example, as Pearson's correlation coefficient.

[0043] The closer the correlation coefficient C is to 1, the more linear the relationship between the virus concentration on the X-axis and the number of positive cases on the Y-axis. In other words, the closer the correlation coefficient C is to 1, the stronger the relationship between virus concentration and the number of positive cases, and the more accurate the prediction of the number of positive cases from virus concentration.

[0044] In this embodiment, the method 1 for predicting the number of infectious disease-positive cases is configured to acquire data sequences x from two sewage treatment plants P, but it may also be configured to acquire data sequences x from three or more sewage treatment plants P. The method 1 for predicting the number of infectious disease-positive cases may also be configured to acquire data sequences x from one sewage treatment plant P, in which case the weighted geometric mean step S2 may be omitted. When the weighted geometric mean step S2 is omitted, the moving geometric mean step S4 accepts the data sequence x in which observation data is arranged in time series as the input data sequence xi(t), and performs moving geometric mean processing.

[0045] In this embodiment, the moving average at time t=1 is calculated by averaging the observation data at times t=1, 2, 3, 4, 5, and 6. That is, in this embodiment, when the processing target time t=t1, the averaging period is configured to be the period from time t=t1 onwards, but the averaging period may be configured to be a different period. For example, when the processing target time t=t1 is the averaging period, the averaging period may be configured to be the period before time t=t1. Furthermore, the averaging period may be configured to be a period that includes both processing data before time t=t1 and processing data after time t=t1.

[0046] Next, an example of the processing in the averaging period determination step S3 described above will be described with reference to Fig. 4. Fig. 4 is a flow diagram showing an example of the processing in the averaging period determination step S3. In the averaging period determination step S3, an averaging period in the moving geometric average step S4 is determined. The averaging period is determined as the number N of averaged data, which is the number of processed data items included in the averaging period and averaged. If the data acquisition interval for processed data (and observed data) is set to sampling time Δt, the relationship between the averaging period and the number N of averaged data items can be expressed by the following formula: averaging period = [sampling time Δt] × [number N-1 of averaged data items].

[0047] As shown in Fig. 4, in the averaging period determination step S3, the input data sequence x i (t) described above is subjected to a moving geometric mean process and a prediction process, and the correlation coefficient C calculated in the prediction process is evaluated. In this embodiment, in the averaging period determination step S3, processing is performed using a preliminary data sequence (not shown) acquired in advance in the data acquisition step S1 (see Fig. 1) as a data sequence x different from the data sequence x described above. The preliminary data sequence is subjected to weighted geometric mean processing in the weighted geometric mean step S2, similar to the processing performed on the data sequence x described above. Then, a weighted average data sequence x aw (t) is calculated from the preliminary data sequence.

[0048] In step S31, the averaging period determination step S3 accepts the calculated weighted average data sequence xaw(t) as the input data sequence xi(t). The input data sequence xi(t) is then subjected to moving geometric averaging to calculate the moving average data sequence xam(t). When the input data sequence xi(t) is subjected to moving geometric averaging, the number of data to be averaged, N, is used as the variable n. In step S32, the input data sequence xi(t) is subjected to moving geometric averaging to calculate the moving average data sequence xam(t) for each n as the nth moving average data sequence xan(t) (n=1, 2, ..., k).

[0049] Next, in the prediction process, a data sequence y(t) of the number of positive cases corresponding to the nth moving average data sequence xan(t) is acquired from a database, etc. In the prediction process, in step S34, a correlation coefficient C between the number of positive cases data included in the acquired data sequence y(t) of the number of positive cases and the processed data included in the nth moving average data sequence xan(t) is calculated as the nth correlation coefficient Cn (n=1, 2, ..., k).

[0050] In the averaging period determination step S3, the calculated value of the nth correlation coefficient Cn is evaluated. Specifically, as shown in graph G2 in the figure, the calculation results are plotted, for example, with the horizontal axis representing the variable n and the vertical axis representing the nth correlation coefficient Cn (step S35). Then, in the averaging period determination step S3, the averaging period is determined as the number N of averaged data based on the value n1 corresponding to the maximum correlation coefficient Cmax, which is the largest value among the nth correlation coefficients Cn, as shown in graph G2 (step S36).

[0051] However, the number of data items N to be averaged may be determined based on the value of n1, and need not be equal to the value of n1. For example, the number of data items N to be averaged may be configured to be determined based on one week (7 days). For example, when one piece of data is acquired per day, the number of data items N to be averaged may be configured to be determined as the closest multiple of 7 to the value of n1.

[0052] That is, in the averaging period determination step S3, the number of processed data included in the averaging period is used as a variable, and each moving average data sequence xam(t) corresponding to each number of processed data (i.e., averaging period) is calculated. In the averaging period determination step S3, a correlation coefficient C between the processed data included in each moving average data sequence xam(t) and the number of positive cases is calculated. Then, the number of averaged data N as the averaging period is determined based on the value of the correlation coefficient C.

[0053] [Processing at each stage (generalized)] Next, the processing of each step of the method 1 for predicting the number of infectious disease positive cases will be generalized and explained with reference to Fig. 5. Fig. 5 is a flow chart showing the processing of each step explained above in a generalized manner.

[0054] 5, in method 1 for predicting the number of positive cases of infectious diseases, a data sequence x is acquired from m systems of sewage treatment plants P as multiple sewage treatment plants P. In the data acquisition step S1, a j-th data sequence xj(t) and a j-th inflow water volume wj are acquired from the j-th sewage treatment plant Pj (j=1, 2, ..., m).

[0055] In the weighted geometric mean step S2, a weighted average data sequence xaw(t) is calculated from each data sequence x using the formula indicated by the symbol F5. The cumulative inflow water volume W in formula F5 is calculated using the formula indicated by the symbol F6. In formula F5, Π is a multiplication symbol, indicating that the multiple values ​​written after Π are multiplied. Also, in formula F6, Σ is an accumulation symbol, indicating that the multiple values ​​written after Σ are accumulated.

[0056] That is, as shown in equation F5 in Figure 5, the weighted average data sequence xaw(t) is calculated by multiplying the jth data sequence xj(t) (j=1, 2, ..., m) by the ratio of the jth inflow water volume wj to the accumulated inflow water volume W. Also, as shown in equation F6 in Figure 5, the accumulated inflow water volume W is calculated by accumulating the jth inflow water volume wj (j=1, 2, ..., m).

[0057] The moving geometric average step S4 receives the weighted average data sequence xaw(t) calculated above as the input data sequence xi(t). In the moving geometric average step S4, the moving average data sequence xam(t) is calculated using the formula indicated by the symbol F7. In formula F7, the averaging period is from time t1 to time t2. Furthermore, the number of consecutive processed data from the processed data x(t1) at time t1 to the processed data x(t2) at time t2 (the number of averaged data N) is N.

[0058] That is, as shown in formula F7, the moving average data string xam(t) is calculated as the Nth root of the product of the processed data x(t1) at time t1 to the processed data x(t2) at time t2. The subsequent processing is the same as the contents of the prediction step S5 described above with reference to Figure 3, so please refer to the above and a detailed description will be omitted.

[0059] [About the action] Next, also with reference to FIG. 5, the operation based on the configuration of the above-described method 1 for predicting the number of infected people will be described below. As described above, the observed data processed by the method 1 for predicting the number of infected people is virus concentration data. The virus concentration in the observed data follows a log-normal distribution. Therefore, the j-th data sequence xj(t) (j = 1, 2, ..., m) composed of multiple observed data and the input data sequence xi(t) obtained by taking a weighted average of the j-th data sequence xj(t) follow a log-normal distribution.

[0060] Method 1 for predicting the number of infected positive cases performs the moving geometric mean process shown in formula F7 in Figure 5 as the averaging process for the input data sequence xi(t). Taking the logarithm of both sides of formula F7, formula F7 becomes formula F9 below. log[xam(t)]=1 / N×Σ[log[xi(k)]](k=t1,…,t2)…(Formula F9)

[0061] Here, for example, if log[xi(t)]=Xi(t) and log[xam(t)]=Xam(t), then formula F9 becomes formula F10 below. Xam(t)=1 / N×Σ[Xi(k)](k=t1,…,t2)…(Formula F10)

[0062] As explained above, the input data sequence xi(t) follows a log-normal distribution. In a log-normal distribution, the logarithm of a random variable follows a normal distribution. Therefore, Xi(t) follows a normal distribution. As shown in equation F10, Xam(t) is the result of taking the moving arithmetic average of Xi(t), which follows a normal distribution. In a normal distribution, the arithmetic mean is equal to the median. Therefore, Xam(t) can appropriately capture the central tendency in the distribution of data.

[0063] On the other hand, when calculating the moving average data sequence xam(t) by performing a moving arithmetic average process, the moving average data sequence xam(t) is calculated by equation F11. xam(t)=1 / N×Σ[xi(k)](k=t1,…,t2)…(Formula F11)

[0064] Since the input data sequence xi(t) follows a log-normal distribution, the distribution curve of the input data sequence xi(t) has a long tail on the right and contains large data values ​​with a certain probability. Therefore, the moving average data sequence xam(t) in formula F11 will show larger values ​​than in formula F10.

[0065] An example of the processing results by the moving geometric mean step S4 of the method 1 for predicting the number of infected positive cases will be described with reference to Fig. 6. Fig. 6 is a diagram showing an example of the processing results by the moving geometric mean step S4 of the method 1 for predicting the number of infected positive cases. In addition to the result of performing the moving geometric mean process on the acquired data string x (time series data of virus concentration observation data) (an example of the result according to the present invention), Fig. 6 also shows the result of performing the moving arithmetic mean process on the data string x as a comparative example.

[0066] In Figure 6, reference numeral 6 denotes a data string of observed data, line 7 is the result of performing a moving geometric average on data string 6, and dotted line 8 is the result of performing a moving arithmetic average on data string 6. The number of data points to be averaged, N, will be described in detail below, but processing was performed with N=14. As explained above, dotted line 8, which is the result of the moving arithmetic average processing, fluctuates at values ​​larger than line 7, which is the result of the moving geometric average processing according to the present invention. Dotted line 8, which is the result of the moving arithmetic average processing, does not adequately capture the central tendency of the observed data due to the influence of large-value observation data such as observation data 6b and 6d, which are included as noise in data string 6.

[0067] Next, we attempted to apply method 1 of the present invention for predicting the number of positive infectious disease cases using sewage surveillance data and open data on the number of positive COVID-19 cases (open data from 59 locations in 18 cities made public through a Cabinet Secretariat demonstration project from 2022 to 2023). That is, we used the sewage surveillance data in the above open data as the jth data sequence xj(t) (j = 1 to 59) and the data on the number of positive COVID-19 cases as the data sequence y(t), and applied method 1 for predicting the number of positive infectious disease cases. As a result, the correlation coefficient Cg between the processed data of the sewage surveillance data obtained by method 1 of the present invention for predicting the number of positive infectious disease cases and the number of positive cases was 0.817 (average value for 59 locations), a high value.

[0068] When the moving geometric mean processing of the present invention was replaced with moving arithmetic mean processing, the correlation coefficient Ca between the processed sewage surveillance data and the number of positive cases was 0.789 (average value of 59 locations). The difference in average value between the correlation coefficient Cg when using the moving geometric mean and the correlation coefficient Ca when using the moving arithmetic mean was 0.027, and the standard deviation of the difference was 0.058. A significance test was performed on the difference between the correlation coefficient Cg when using the moving geometric mean and the correlation coefficient Ca when using the moving arithmetic mean for the above 59 locations, and the t-value was 3.59, indicating that the difference was significant (95% significant). In other words, by performing moving geometric mean processing in method 1 for predicting the number of positive cases of infectious diseases, the correlation coefficient C was improved compared to when the commonly used moving arithmetic mean processing was performed.

[0069] As described above, in method 1 for predicting the number of infected people, the input data sequence xi(t) is subjected to a moving geometric mean process, thereby performing an averaging process suited to the distribution characteristics (log-normal distribution) of the virus concentration in the input data sequence xi(t). In other words, in method 1 for predicting the number of infected people, the effects of noise and other factors contained in the input data sequence xi(t) are appropriately eliminated, and the input data sequence xi(t) is averaged. Therefore, method 1 for predicting the number of infected people can appropriately grasp the relationship between the averaged processed data and the number of infected people. As a result, method 1 for predicting the number of infected people can accurately predict the increase or decrease trend in the number of infected people.

[0070] Furthermore, the method 1 for predicting the number of infectious disease positive cases according to the present invention performs the weighted geometric mean process shown in formula F5 as the weighted mean process of the data string x observed at each of multiple observation locations. In the method 1 for predicting the number of infectious disease positive cases according to the present invention, for example, the inflow water volume w into the sewage treatment plant P is set as a weighted value as a value related to the catchment population of the sewage treatment plant P.

[0071] The method 1 for predicting the number of infected positive cases according to the present invention performs the weighted average of the data string x as a weighted geometric mean process, as shown in formula F5. When both sides of formula F5 are logarithmized, formula F5 becomes formula F12 below. log[xaw(t)]=1 / W×Σ{wj×log[xj(t)]}(j=1,…,m)…(Formula F12)

[0072] Here, similarly to the above, if log[xaw(t)]=Xaw(t) and log[xj(t)]=Xj(t), for example, then formula F12 becomes formula F13 below. Xaw(t)=1 / W×Σ[wj×Xj(t)]…(Formula F13)

[0073] As explained above, the jth data sequence xj(t) follows a log-normal distribution, and therefore Xj(t) follows a normal distribution. From equation F13, Xaw(t) is the result of the weighted arithmetic mean of Xj(t), which follows a normal distribution. Therefore, Xaw(t) is weighted as intended and can appropriately capture the central tendency of the observed data.

[0074] On the other hand, when the weighted average data sequence xaw(t) is calculated by weighted arithmetic averaging, for example, the weighted average data sequence xaw(t) is calculated by the following formula F14. xaw(t)=1 / W×Σ[wj×x(j)](j=1,…,m)…(Formula F14)

[0075] Because the j-th data sequence xj(t) follows a log-normal distribution, the distribution curve for the j-th data sequence xj(t) has a long tail on the right, and large-value data is included with a certain probability. Therefore, the weighted average data sequence xaw(t) shown in formula F14 is strongly influenced by the weighting given to the j-th data sequence xj(t), which has large values, and therefore may not achieve the intended weighting compared to the case shown in formula F13.

[0076] Next, an example of the results of processing by the weighted geometric mean step S2 of the present invention will be described with reference to Figure 7. Figure 7 shows the results of processing the acquired data string x by the weighted geometric mean processing method 1 of the present invention for predicting the number of infected positive cases, as well as the results of processing the data string x by the weighted arithmetic mean processing as a comparative example. Figure 7 shows an example of the processing results for two data strings x acquired from two sewage treatment plants P in Komatsu City, which are part of the open data from 59 locations in 18 cities made public through the Cabinet Secretariat demonstration project described above.

[0077] In Figure 7, line 9 shows the results of weighted geometric mean processing, using the inflow water volumes w of each of the two sewage treatment plants P as weights. Similarly, dotted line 10 shows the results of weighted arithmetic mean processing, instead of the weighted geometric mean processing of the present invention, using the inflow water volumes w of each of the two sewage treatment plants P as weights. Reference numeral 11 indicates the trend in the number of positive cases. As shown particularly in the period indicated by reference numeral 12 in Figure 7, the results of the weighted arithmetic mean processing (dotted line 10) tend to be partially larger than the results of the weighted geometric mean processing (line 9). Since no sudden increase in the number of positive cases was observed during period 12, the results of dotted line 10 may not adequately capture the central tendency of data string x due to the intended weighting not being applied, as explained above.

[0078] When the correlation coefficient between the processed data and the number of positive cases was checked for the data shown in Figure 7, the correlation coefficient was 0.739 when the weighted geometric mean processing was performed, and the correlation coefficient was 0.723 when the weighted arithmetic mean processing was performed. In other words, it was confirmed that the correlation coefficient when the weighted geometric mean processing was performed was larger than the correlation coefficient when the weighted arithmetic mean processing was performed.

[0079] As described above, in method 1 for predicting the number of infected people, the jth data sequence xj(t) (j = 1, ..., m) is subjected to a weighted geometric mean process using the jth inflow water volume wj as a weight related to the basin population. By performing the weighted geometric mean process, method 1 for predicting the number of infected people performs a weighted mean process appropriate for the distribution characteristics (log-normal distribution) of the virus concentration in each jth data sequence xj(t). In other words, method 1 for predicting the number of infected people can appropriately eliminate the influence of noise and other factors contained in each jth data sequence xj(t) and appropriately weight each jth data sequence xj(t) by the jth inflow water volume wj. Therefore, method 1 for predicting the number of infected people can appropriately grasp the relationship between the weighted-averaged processed data and the number of infected people. As a result, method 1 for predicting the number of infected people can accurately predict the increase or decrease in the number of infected people.

[0080] Furthermore, in the method 1 for predicting the number of infected people according to the present invention, the averaging period in the moving geometric average process is determined in the averaging period determination step S3. Specifically, the number of data to be averaged N is determined so that the correlation coefficient C between the processed data of the moving average data sequence xam(t) and the number of infected people becomes large.

[0081] An example of processing in the averaging period determination step S3 of the method 1 for predicting the number of infected positive cases according to the present invention will be described with reference to FIG. 8. FIG. 8 is a diagram showing the results of processing in the averaging period determination step S3. In FIG. 8, observation data acquired once a day was used. The horizontal axis of FIG. 8 indicates the number of days for which the moving average was calculated (i.e., the number of data points for which the moving average was calculated). The vertical axis of FIG. 8 is the correlation coefficient C corresponding to the number of days for which the moving average was calculated on the horizontal axis. From FIG. 8, the averaging period corresponding to the maximum correlation coefficient Cmax was 11 days (i.e., the number of data points was 11). In this embodiment, taking into consideration the cyclical nature of society, the averaging period was set to, for example, 14 days (i.e., the number of days to be averaged N=14), which is the number of days closest to 11 days among multiples of 7 days that make up a week.

[0082] Therefore, in method 1 for predicting the number of positive cases of infectious disease, the averaging period is optimized so that the correlation coefficient C between the processed data and the number of positive cases of infectious disease becomes large. As a result, method 1 for predicting the number of positive cases of infectious disease can accurately predict the increase or decrease trend in the number of positive cases of infectious disease.

[0083] The averaging period is set appropriately depending on the interval (sampling time) at which the observation data is acquired. The averaging period may be configured as the number of days to be averaged, as shown in the example of Fig. 8, or may be configured as the number of observation data to be averaged.

[0084] [Predictor of the number of positive cases of infectious diseases 2] Next, referring to FIG. 9, a description will be given of a prediction device 2 for the number of infected positive cases that executes a prediction method 1 for the number of infected positive cases according to the present invention. FIG. 9 is a block diagram of the prediction device 2 for the number of infected positive cases. As shown in FIG. 9, the prediction device 2 for the number of infected positive cases includes a control unit 21, an input / output unit 22, a data processing unit 30, a memory unit 23, and a communication unit 24. The control unit 21 includes a processor such as a CPU (Central Processing Unit). The memory unit 23 includes a storage device (not shown) and stores data and computer programs. The input / output unit 22 is a user interface that has an input function for accepting various information input from a user to the prediction device 2 for the number of infected positive cases and an output function for outputting various information from the prediction device 2 for the number of infected positive cases. The communication unit 24 communicates with the water treatment plant P, database, and the like described above. The control unit 21 controls the input / output unit 22, the memory unit 23, the communication unit 24, and the data processing unit 30.

[0085] As shown in FIG. 9, the data processing unit 30 includes a data acquisition unit 31, a weighted geometric mean unit 32, an averaging period determination unit 33, a moving geometric mean unit 34, and a prediction unit 35. The data acquisition unit 31 executes a data acquisition step S1 (see FIG. 1). The weighted geometric mean unit 32 executes a weighted geometric mean step S2. The averaging period determination unit 33 executes an averaging period determination step S3. The moving geometric mean unit 34 executes a moving period averaging step S4. The prediction unit 35 executes a prediction step S5. The data processing unit 30 is a computer program stored in the storage device of the storage unit 23. The processor of the control unit 21 executes the computer program stored in the storage device of the storage unit 23 to control the data processing unit 30. The control unit 21 outputs the predicted result of the number of infected people, which is the result of processing by the data processing unit 30, to the input / output unit 22. In other words, the device 2 for predicting the number of infected people executes the method 1 for predicting the number of infected people.

[0086] The device 2 for predicting the number of positive infection cases executes the method 1 for predicting the number of positive infection cases, thereby performing an averaging process suited to the distribution characteristics (log-normal distribution) of information in the data sequence x, as explained above. In other words, the device 2 for predicting the number of positive infection cases averages the data sequence x by appropriately eliminating the influence of noise and the like contained in the data sequence x. Therefore, the device 2 for predicting the number of positive infection cases can appropriately grasp the relationship between the averaged data sequence x and the number of positive infection cases.

[0087] The embodiments of the present invention have been described above with reference to the drawings. However, the present invention is not limited to the above embodiments and can be embodied in various forms without departing from the spirit and scope of the present invention. The drawings mainly show each component in a schematic manner for ease of understanding, and the thickness, length, number, spacing, etc. of each component shown in the drawings may differ from the actual components due to the convenience of creating the drawings. Furthermore, the materials, shapes, dimensions, etc. of each component shown in the above embodiments are merely examples and are not particularly limited, and various modifications are possible within a scope that does not substantially deviate from the configuration of the present invention. [Explanation of symbols]

[0088] 1. Methods for predicting the number of infected people S1 Data acquisition process S2 weighted geometric averaging process S3 Average period determination process S4 Moving geometric mean process S5 Prediction process P1 No. 1 Sewage Treatment Plant P2 Second Sewage Treatment Plant x data column x1 First data column x2 Second data column xi Input data sequence xaw Weighted average data sequence xam Moving Average Data Column y Positive case count data column N Number of averaged data w Inflow water volume w1 1st inflow water volume w2 2nd inflow water volume W Accumulated inflow water volume L regression line C correlation coefficient

Claims

1. a data acquisition step in which observation data, which is information about infectious disease pathogens in water at the water treatment plant, is acquired as a time-series data string; a moving geometric averaging step of receiving the data string as an input data string and performing a moving geometric averaging process on the input data string to calculate a moving average data string; a prediction step of predicting an increase / decrease trend in the number of positive cases of an infection caused by the infectious disease pathogen using the number of positive cases and the moving average data string; Equipped with A method for predicting the number of infected positive cases, wherein the moving geometric mean processing involves performing geometric mean processing on the input data string while moving the averaging period over a predetermined averaging period in the input data string.

2. The method further comprises a weighted geometric mean step in which the data string is subjected to a weighted geometric mean process; In the data acquisition step, the data strings for each of the plurality of water treatment plants and weighted values ​​associated with the catchment population for each of the water treatment plants are acquired; In the weighted geometric mean step, the data strings weighted by the respective weight values ​​are subjected to a weighted geometric mean process to calculate a weighted average data string; 2. The method for predicting the number of infected positive cases according to claim 1, wherein the moving geometric mean step receives the weighted average data string as the input data string and performs the moving geometric mean processing.

3. further comprising an averaging period determination step of determining the averaging period, In the averaging period determination step, the moving average data sequence corresponding to each of the averaging periods is calculated using the averaging period as a variable, In the averaging period determination step, a correlation coefficient indicating the degree of correlation between the processed data subjected to the moving geometric average process and the number of positive cases, which is included in each of the moving average data strings, is calculated, and The method for predicting the number of infected people according to claim 1 or claim 2, wherein the average period is determined based on the value of the correlation coefficient.

4. The method for predicting the number of infected people according to claim 2 , wherein the weighted value is the amount of inflow water at each of the water treatment plants.

5. A method for predicting the number of infectious disease positive cases described in claim 1 or claim 2, wherein in the prediction process, a regression line showing the relationship between the number of positive cases and the observation data processed by the moving geometric average in the moving average data sequence is calculated, and the number of positive cases for specific observation data is predicted as the increase / decrease trend using the regression line.

6. a data acquisition unit that acquires observation data, which is information about infectious disease pathogens in water at the water treatment plant, as a time-series data string; a moving geometric average unit that receives the data string as an input data string and performs a moving geometric average process on the input data string to calculate a moving average data string; a prediction unit that predicts an increase or decrease trend in the number of positive cases of an infection caused by the infectious disease pathogen using the number of positive cases and the moving average data string; Equipped with The moving geometric mean unit performs geometric mean processing on the input data string while moving over a predetermined averaging period in the input data string, in a prediction device for the number of infected positive cases.

Citation Information

Patent Citations

  • Infectious disease testing methods

    JP7031957B1