Model generation device, flood probability prediction device, model generation method, model generation program, flood probability prediction method, flood probability prediction program, flood probability prediction system, and trained model

The model generation device and flood probability prediction system using water detectors and machine learning address inefficiencies in conventional methods by predicting flooding probability accurately and efficiently in areas where water is not normally present, reducing power consumption and installation complexity.

JP2026075423APending Publication Date: 2026-05-08TOKYO DENKI UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
TOKYO DENKI UNIVERSITY
Filing Date
2024-10-22
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Conventional flood prediction methods, such as those using water level gauges and water detectors, are inefficient and power-intensive, and cannot accurately predict flooding in areas where water does not normally exist, like roads and residential areas, due to the need for continuous power and installation space, and lack of time-series data during non-flooding periods.

Method used

A model generation device and flood probability prediction system using a water detector to detect changes in water levels, combined with machine learning, to create a trained model that predicts flooding probability based on rainfall time series data, allowing for accurate predictions with reduced power consumption and simpler installation.

Benefits of technology

Enables highly accurate flooding probability predictions in areas where water detectors are installed, reducing power consumption and installation complexity while providing timely flood risk assessments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026075423000001_ABST
    Figure 2026075423000001_ABST
Patent Text Reader

Abstract

This system enables highly accurate water level prediction using a water detector with a simpler configuration compared to a water level gauge. [Solution] The model generation device 3 generates a model that predicts the probability of new flooding occurring at specific locations P0 to Pm after a predetermined time. The model generation device 3 includes a water detector 2 that detects whether or not flooding occurs at a specific location, a training data creation unit 35 that creates a training dataset including training rainfall time series information RL for the specific location based on the detection results of the water detector 2, and training flooding probability information PL for the specific location after the latest time point of the training rainfall time series information RL, and a model learning unit 36 ​​that takes the training rainfall time series information RL from the training dataset as input to the model 4 and the training flooding probability information PL as output to the model 4, performs machine learning so that the input / output relationship of the model 4 approaches the input / output relationship of the training dataset, and generates a trained model 4A.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure relates to a model generation device, a flood probability prediction device, a model generation method, a model generation program, a flood probability prediction method, a flood probability prediction program, a flood probability prediction system, and a trained model. [Background technology]

[0002] Climate change is causing an increase in the frequency of heavy rainfall, raising concerns that more rainwater will flow onto roads in a shorter amount of time than before. This is exacerbated in areas where urbanization has increased the amount of surface soil that makes it difficult for rainwater to penetrate the ground. Road sections that cannot adequately drain the incoming rainwater will become flooded, requiring traffic restrictions. For road users, residents near flooded areas, and road administrators, predicting the occurrence of traffic restrictions due to road flooding is useful for ensuring safety and efficiency in daily life and travel. For example, based on predictive information, road users can determine safe routes, and road administrators can manage roads more efficiently.

[0003] As a technology related to flood prediction, for example, Patent Documents 1 and 2 disclose a method for measuring time-series data of water levels using water level gauges at multiple points along a river to be predicted, and then using the measured water level data to predict the river's water level.

[0004] On the other hand, Patent Document 3 discloses a configuration in which a water detector is installed at the target location to detect when a predetermined water level has been reached, for purposes such as detecting flooding of roads. [Prior art documents] [Patent Documents]

[0005] [Patent Document 1] Japanese Patent Application Publication No. 09-256338 [Patent Document 2] Japanese Patent Publication No. 2020-134300 [Patent Document 3] Japanese Patent Publication No. 2023-66488 [Overview of the project] [Problems that the invention aims to solve]

[0006] Conventional water level prediction methods, such as those described in Patent Documents 1 and 2, target rivers, so the water level gauges are installed in measurement environments where water is constantly flowing. Therefore, continuously measuring water levels using water level gauges and acquiring time-series data is effective in improving the accuracy of water level predictions. However, water level gauges used in conventional water level prediction methods measure and transmit water levels at regular time intervals, requiring constant power, expense, and installation space. For this reason, water level gauges are generally not installed in places where water does not normally exist, such as roads or residential areas. Furthermore, even if a water level gauge is installed in a place where water does not normally exist, it is unlikely that water levels can be measured for most of the time, making it highly probable that most of the time-series data measured will be wasted. In addition, because water level gauges must be operated continuously, wasted power consumption may occur during periods when valid data cannot be measured.

[0007] On the other hand, conventional flood detection methods, such as those described in Patent Document 3, recognize the occurrence of flooding at a target location only after detecting that a predetermined water level has been reached at that location using a water detector. Water detectors do not transmit data to external devices during periods when there is no change in the presence or absence of water, thus reducing power consumption, and their structure is simpler compared to water level gauges, which is an advantage. However, conventional flood detection methods using water detectors cannot predict in advance when flooding will occur at a target location.

[0008] Furthermore, the water level prediction methods described in Patent Documents 1 and 2 acquire time-series data at the same time intervals regarding the water level of the river to be predicted and use it for water level prediction. On the other hand, the water detector in Patent Document 3 is a device for detecting changes in the presence or absence of water at a predetermined water level at the installation site. Therefore, even if multiple water detectors are installed at the same installation site to detect at different water levels, while information on the time when the presence or absence of water at the predetermined water level changed can be obtained from each water detector, it is not possible to continuously measure the water level before and after that time at regular time intervals and obtain time-series data. Consequently, the water detector in Patent Document 3 cannot be applied instead of a water level meter in the water level prediction methods described in Patent Documents 1 and 2.

[0009] This disclosure aims to enable highly accurate flood probability prediction using a water detector with a simpler configuration compared to a water level gauge. [Means for solving the problem]

[0010] A model generation device according to one aspect of an embodiment of the present invention is a model generation device for generating a model for predicting the probability of new flooding occurring at a specific location after a predetermined time, comprising: a water detector installed at the specific location that detects the presence or absence of flooding at the specific location by detecting a change in the presence or absence of water at a predetermined water level; a training data creation unit that creates a training dataset including training rainfall time series information relating to time series information of rainfall at the specific location based on the detection result of the water detector, and training flooding probability information relating to the probability of new flooding occurring at the specific location after the latest time point of the training rainfall time series information; and a model learning unit that uses the training dataset created by the training data creation unit to perform machine learning so that the input / output relationship of the model approaches the input / output relationship of the training dataset, using the training rainfall time series information as input to the model and the training flooding probability information as output to the model, thereby generating a trained model.

[0011] Similarly, a flooding probability prediction device according to an aspect of an embodiment of the present invention is a flooding probability prediction device for predicting the probability that new flooding will occur at a specific location after a predetermined time point, comprising: a learned model obtained by machine learning of the correspondence between a learning explanatory variable including learning rainfall time series information regarding the time series information of the rainfall at the specific location, and a learning objective variable including learning flooding probability information regarding the probability that new flooding will occur after the latest time point of the learning rainfall time series information at the specific location; and a prediction unit configured to input an explanatory variable including prediction rainfall time series information regarding the time series information of the rainfall at the specific location before the implementation time including the implementation time for carrying out the prediction into the learned model, output information of an objective variable related to the information of the explanatory variable, and predict the probability after the implementation time based on the objective variable.

Advantages of the Invention

[0012] According to the present disclosure, it is possible to perform highly accurate flooding probability prediction using a water detector with a simple configuration compared to a water level gauge.

Brief Description of the Drawings

[0013] [Figure 1] Diagram showing the overall configuration of a flooding probability prediction system according to the first embodiment [Figure 2] Functional block diagram of a model generation device according to the first embodiment [Figure 3] Diagram showing an example of a time series pattern of the time interval of learning rainfall time series information set by an explanatory variable setting unit [Figure 4] Diagram showing an example of rainfall data at the time of flooding detection [Figure 5] Functional block diagram of a flooding probability prediction device according to the first embodiment [Figure 6] Hardware configuration diagram of a model generation device and a flooding probability prediction device [Figure 7] Flowchart of model learning control according to the first embodiment [Figure 8] Flowchart of flooding probability prediction control according to the first embodiment [Figure 9]This figure shows an example of the prediction results for the probability of flooding occurring according to the first embodiment. [Figure 10] This figure shows an example of flood prediction and evaluation according to the first embodiment. [Figure 11] A diagram showing an example of the occurrence rate of flooding probability of 0.5 or higher. [Figure 12] Functional block diagram of the model generation device according to the second embodiment [Figure 13] This figure shows an example of a time-series pattern of multiple step sizes set by the explanatory variable setting unit of the second embodiment. [Figure 14] This figure illustrates a modified example of the explanatory variable setting unit. [Figure 15] Functional block diagram of the flood probability prediction device according to the second embodiment. [Modes for carrying out the invention]

[0014] The embodiments will be described below with reference to the attached drawings. To facilitate understanding of the explanation, the same reference numerals are used for identical components in each drawing whenever possible, and redundant explanations are omitted.

[0015] [First Embodiment] The first embodiment will be described with reference to Figures 1 to 12. Figure 1 is a diagram showing the overall configuration of the flooding probability prediction system 1 according to the first embodiment.

[0016] The flood probability prediction system 1 is a system for predicting the probability (flood occurrence probability) Pe of new flooding occurring after a predetermined time in specific locations P1 to Pm where water detectors 2 are installed. In this system, the specific locations P1 to Pm that are the targets for predicting the flood occurrence probability Pe include places where there is no water under normal circumstances, such as roads and residential areas. In particular, it includes places where there is a relatively high possibility of flooding during rainfall, such as lowlands, depressions, and underpasses, as well as places in urban areas where there is a relatively high possibility of flooding, such as basements, subway stations, underground shopping malls, underground parking lots, and electrical rooms. Furthermore, specific locations P1 to Pm can also include places where there is water under normal circumstances and where there is a relatively high possibility of overflow during rainfall, such as rivers, lakes, reservoirs, and their banks.

[0017] In this embodiment, the term "flooding" is used as a broader concept encompassing the aforementioned events such as inundation, flooding, and overflowing. Furthermore, although the specific locations targeted for prediction by the flooding probability prediction system 1 are shown as m locations from P1 to Pm in Figure 1, they may be a single location or any multiple locations.

[0018] The flood probability prediction system 1 acquires various information when a predetermined water level is reached using water detectors 2 installed at specific locations P1 to Pm, and uses machine learning to obtain the correspondence between rainfall and the probability of flooding at these specific locations P1 to Pm based on this acquired information. The flood probability prediction system 1 can use these learning results to predict the probability of flooding Pe at these specific locations P1 to Pm.

[0019] As shown in Figure 1, the flood probability prediction system 1 comprises a water detector 2, a model generation device 3, a prediction model 4, and a flood probability prediction device 5.

[0020] The water detector 2 is installed at specific locations P1 to Pm that are the target of the flood probability Pe prediction, and is a device that detects when the presence or absence of water changes at a predetermined water level. In the example shown in Figure 1, the water detector 2 has a detection unit 21 that detects the presence or absence of water at a predetermined height h vertically above a reference surface G such as the ground, and by detecting the presence or absence of water with the detection unit 21, it is possible to detect when the presence or absence of water at the predetermined water level h at the specific location has changed. As shown in Figure 1, in this embodiment, the detection unit 21 is installed so as to be able to detect the presence or absence of water at a height h (i.e., flood height h) from the reference surface G where the position of the water surface W can be judged as flooding has occurred. As a result, when water is detected by the detection unit 21, the water detector 2 can detect that the water surface W has risen to the flood height h and that flooding has occurred.

[0021] The flood height h at which the detection unit 21 is installed may be set individually according to various conditions (for example, the elevation and uneven shape of the reference surface G) of the specific locations P1 to Pm that are the target of the flood probability Pe prediction. Furthermore, the flood probability prediction system 1 may be configured to install multiple water detectors 2 at each of the specific locations P1 to Pm and calculate the flood probability for multiple heights h.

[0022] Here, water detector 2 can be defined as a device that does not transmit information during periods when there is no change in the presence or absence of water, and only transmits information about the occurrence of a change in the presence or absence of water at a certain height to an external device via wireless or wired communication. On the other hand, a comparative example of water detector 2 is a device generally called a water level meter. A water level meter is a device that acquires water levels at regular intervals. In the aforementioned Patent Documents 1 and 2, such water level meters are used to measure time-series data of river water levels at the same time intervals and transmit it to an external device. For the operation of a water level meter, a commercial power supply or solar power generation and a large-capacity storage battery are required, so construction work including the power supply and a certain amount of space are required. In contrast, water detector 2 can limit the value of the water level to be detected and does not need to transmit information related to the water level to an external device at regular intervals, so power consumption can be reduced compared to a water level meter, and for example, it can operate for several years on dry cell batteries. Also, because water detectors are smaller and cheaper than water level meters, they are more likely to be installed in locations where road flooding is a concern and where the installation of water level meters is difficult. Furthermore, if there are surveillance cameras at specific locations P1 to Pm that are the target of the flood probability Pe prediction, it is conceivable to use them as water detectors by having the surveillance cameras determine changes in water relative to a threshold. In other words, the water detector 2 of this embodiment also includes devices that can detect changes in the presence or absence of water using a mechanism different from the configuration in which the presence or absence of water is directly detected by the detection unit 21 as in the example in Figure 1.

[0023] Model generation device 3 is a device that generates a model (trained model 4A) for predicting the probability of flooding occurrence Pe at specific locations P1 to Pm, using prediction model 4.

[0024] Predictive model 4 is a model that undergoes machine learning using the model generation device 3, and after machine learning, is used by the flood probability prediction device 5 to predict the flood probability Pe. In the following explanation, predictive model 4, after machine learning is complete and specific input / output relationships have been acquired, will also be referred to as "trained model 4A".

[0025] The flood probability prediction device 5 is a device for predicting the probability of flooding Pe for each specific location P1 to Pm using the trained model 4A generated by the model generation device 3.

[0026] The model generation device 3 will be described with reference to Figures 2 to 4. Figure 2 is a functional block diagram of the model generation device 3 according to the first embodiment.

[0027] As shown in Figure 2, the model generation device 3 includes a rainfall data acquisition unit 31, an explanatory variable setting unit 32, a detection information storage unit 33, a rainfall data storage unit 34, a training data creation unit 35, and a model learning unit 36.

[0028] The rainfall data acquisition unit 31 acquires time-series data (rainfall data) DA of rainfall at specific locations P1 to Pm for which the probability of flooding is predicted. The rainfall data acquisition unit 31 can be configured to acquire rainfall data DA from outside the flooding probability prediction system 1, for example, by collecting rainfall data DA for specific locations P1 to Pm via an internet connection from observation data of a real-time rainfall observation system (XRAIN) using high-performance weather radar operated by the Ministry of Land, Infrastructure, Transport and Tourism. Alternatively, the flooding probability prediction system 1 can be equipped with rain gauges installed at specific locations P1 to Pm, and the rainfall data acquisition unit 31 can acquire rainfall data DA measured by the rain gauges, thus acquiring rainfall data DA using elements within the flooding probability prediction system 1.

[0029] Furthermore, in this embodiment, in order to reduce the number of explanatory variables described later and to suppress the variability of rainfall data, rainfall data DA is created by aggregating rainfall temporally and spatially. For example, when using precipitation nowcasting (10-minute precipitation, 1km mesh) for predicted rainfall, the number of explanatory variables and the variability of data can be reduced by averaging XRAIN (1-minute rainfall intensity, 250m mesh) every 10 minutes from the correct time and every 1km mesh to create rainfall data DA. In addition, when predicted rainfall is used at a certain time, as soon as the observed rainfall data arrives, the predicted rainfall at that time is replaced with the more accurate observed rainfall data for flood probability prediction processing, thus enabling seamless use of observed and predicted rainfall.

[0030] The explanatory variable setting unit 32 sets the time series pattern of the step size for the rainfall time series information to be used for learning, which will be described later. Based on command values ​​set in advance, for example by user input, the explanatory variable setting unit 32 determines how many data points to extract from the rainfall data DA and with what step size to create explanatory variables, and outputs this information to the training data creation unit 35.

[0031] The detection information storage unit 33 stores and holds detection information including when water is observed by the water detector 2, that is, when the water surface W exceeds the flood height h (state H), or when the observed water disappears, that is, when the water surface W falls below the flood height h (state L), and the time T at which these states H and L were detected.

[0032] The rainfall data storage unit 34 stores and holds the rainfall data DA acquired by the rainfall data acquisition unit 31.

[0033] The training data creation unit 35 creates a training dataset that includes training rainfall time series information RL and training flood probability information PL. The training rainfall time series information RL is time series information of rainfall at specific locations P1 to Pm based on the detection results of the water detector 2. The training flood probability information PL is information about the probability that new flooding will occur at specific locations P1 to Pm after the latest time point in the training rainfall time series information RL. Of the training dataset, the training rainfall time series information RL is an explanatory variable in machine learning by the model learning unit 36 ​​described later, and is used as input information for the prediction model 4. Also, of the training dataset, the training flood probability information PL is an objective variable in machine learning by the model learning unit 36 ​​described later, and is used as output information for the prediction model 4.

[0034] More specifically, the training dataset includes a "positive dataset" and a "negative dataset." The positive dataset represents data from when flooding actually occurred. The negative dataset, on the other hand, represents data from when the trends, such as the time course of rainfall, are similar to those of the positive dataset, but flooding did not actually occur.

[0035] In the positive dataset, the RLP (Rainfall Time Series) information is extracted from past rainfall data DA, going back to a predetermined time before the point in time when flooding was detected by water detector 2. 1:n These are the explanatory variables and are used as input information for prediction model 4. The sign n in the above rainfall time series information is an arbitrary positive number and indicates the number of explanatory variables. In this embodiment, as will be described later (see Figure 3, etc.), the number of explanatory variables is 6 (n=6). In the case of a positive dataset, water is detected by the water detector 2 and flooding actually occurs, so the flooding probability PLP is "1", and this flooding probability PLP is the target variable and is used as output information for prediction model 4.

[0036] In the negative dataset, the RLN (Rainfall Time Series) data is selected from past rainfall data DA (Digital Amount Data) for sections where flooding was not detected by water detector 2, but whose characteristics are similar to the rainfall time series information in the positive dataset.1:n These are the explanatory variables and are used as input information for prediction model 4. In the case of a negative dataset, no water is detected by water detector 2, and no flooding actually occurs, so the flooding probability PLN becomes "0". This flooding probability PLN is the dependent variable and is used as output information for prediction model 4.

[0037] In this embodiment, the water detector 2 can be said to be an element used to set the training flood probability information PL (0 or 1) of the training dataset. Furthermore, in this embodiment, the detection information from the water detector 2 is not used for flood prediction using the flood probability prediction device 5.

[0038] Figure 3 shows an example of a time series pattern of the step size for the learning rainfall time series information RL set by the explanatory variable setting unit 32. As shown in Figure 3, a predetermined period (10 minutes in the example of Figure 3) including the time when flooding was detected by the detection unit 21 of the water detector 2 in the positive dataset is defined as the "0 minute" time period. The rainfall during this "0 minute" time period is extracted as the first learning rainfall time series information RLP1. At this time, the learning rainfall time series information RLP1 may be calculated and used as the average value of the rainfall over the entire flooding detection time, i.e., the "0 minute" time period, or a representative value of the rainfall during this time period may be used after filtering to remove outliers. The extraction method for other learning rainfall time series information thereafter is the same.

[0039] In the example in Figure 3, the interval for the 40 minutes immediately preceding the detection time is set to 10 minutes. That is, the predetermined period for dividing the time into time zones is set to 10 minutes. At this time, the 10-minute time zone immediately preceding the "0 minute" time zone is defined as the "-10 minute" time zone, the 10-minute time zone immediately preceding the "-10 minute" time zone is defined as the "-20 minute" time zone, and the 10-minute time zone immediately preceding the "-20 minute" time zone is defined as the "-30 minute" time zone. The rainfall for each of these "-10 minute," "-20 minute," and "-30 minute" time zones is then extracted as the second, third, and fourth learning rainfall time series information RLP2, RLP3, and RLP4. The rainfall during the 40 minutes from 30 minutes before the flood detection is called "immediate rainfall."

[0040] For time periods prior to 40 minutes before the most recent rainfall, the time increment is set to 3 hours. In other words, the predetermined period for dividing the time period is set to 3 hours. In this case, the time period immediately preceding the "-30 minutes" time period is defined as the "3-hour" time period. The rainfall amount during the "3-hour" time period is then extracted as the fifth learning rainfall time series information, RLP5.

[0041] Furthermore, for time periods prior to the "3-hour" time period, the interval is set to 9 hours. That is, the predetermined period for dividing the time period is set to 9 hours. In this case, the time period immediately preceding the "3-hour" time period is defined as the "9-hour" time period. Then, the rainfall during the "9-hour" time period is extracted as the sixth learning rainfall time series information, RLP6.

[0042] Figure 4 shows an example of rainfall data at the time of flood detection. The horizontal axis in Figure 4 represents time (minutes), and the vertical axis represents rainfall intensity (mm / h). Rainfall intensity is an indicator of the strength of rainfall, and rainfall intensity (mm / h) is the strength of the rain that corresponds to the amount of rainfall if it rained continuously for one hour. The vertical axis in Figure 4 represents the rainfall intensity for a 1km mesh that includes the specific location where flooding was detected.

[0043] Figure 4 illustrates four types of graphs: solid lines, dotted lines, dashed lines, and double-dash lines. These four graphs show time-series data of rainfall immediately before flooding is detected under different circumstances. In the example in Figure 4, flooding is detected at time 0, and the time-series data of rainfall up to 50 minutes prior to the flooding detection time is shown.

[0044] In Figure 4, the plot of each graph at time 0 minutes corresponds to the first training rainfall time series information RLP1 for the "0 minute" time period mentioned above. Similarly, the plot of each graph at time -10 minutes corresponds to the second training rainfall time series information RLP2 for the "-10 minute" time period mentioned above. The plot of each graph at time -20 minutes corresponds to the third training rainfall time series information RLP3 for the "-20 minute" time period mentioned above. The plot of each graph at time -30 minutes corresponds to the fourth training rainfall time series information RLP4 for the "-30 minute" time period mentioned above. These training rainfall time series information RLP1 to RLP4 are rainfall data at the time of crown detection, and therefore correspond to the positive dataset mentioned above.

[0045] Figure 4 illustrates the time-series information of the preceding rainfall. Referring to each graph in Figure 4, it can be seen that the preceding rainfall has a significant impact on the occurrence of flooding. For this reason, in the 40 minutes prior to the preceding rainfall, the time step size for the learning rainfall time-series information RLP1 to RLP4 is set to 10 minutes, which is smaller than the time step size for the learning rainfall time-series information RLP5 and RLP6 for the time period prior to the preceding rainfall.

[0046] Incidentally, water levels may rise due to rainfall prior to the most recent rainfall. For this reason, rainfall prior to the most recent rainfall is incorporated into the prediction model. However, if explanatory variables are assigned in 10-minute intervals, similar to the most recent rainfall, the number of variables increases, leading to a decrease in generalization performance due to surface concentration phenomena. Furthermore, if the number of variables exceeds the number of training data, the prediction model will not converge. Therefore, in this embodiment, longer time periods are aggregated, and the average rainfall intensity for the 3 hours prior to the most recent rainfall and the 9 hours prior to that are each used as explanatory variables (i.e., training rainfall time series information RLP5 and RLP6). This results in a total of six explanatory variables.

[0047] Furthermore, regarding the spherical concentration phenomenon due to an increase in the number of explanatory variables, or the so-called curse of dimensionality which leads to a decrease in generalization performance, simulations using epidemiological data have been conventionally performed regarding the number of explanatory variables per data point (EPV: events per variable) in logistic regression, for example. These studies have reported that (a) there were accuracy problems when the EPV was small, (b) there were fewer problems when the EPV was 10 or more, and (c) the correlation between explanatory variables needs to be considered when setting the EPV. Since it is unlikely that there is a linear correlation (multicollinearity) between hourly rainfall, it is considered possible to set the EPV to 10 or less in this embodiment. The number of experimental data points (see Figure 10), which will be described later, is 83 or more. However, the positive and negative datasets are imbalanced, and the positive dataset is small. The small number of classes is important for ensuring generalization performance. Taking this into consideration, in this embodiment, the number of explanatory variables is set to 6.

[0048] The number of explanatory variables, i.e., the number of learning rainfall time series information RLs, can be appropriately changed and set according to, for example, the environmental conditions of the installation location of the water detector 2, and is not limited to six as shown in the example in Figure 3. For this reason, in Figure 2, the output of the training data creation unit 35 uses an arbitrary positive number n to obtain the learning rainfall time series information of the positive dataset as RLP. 1:n This notation is used to represent the rainfall time series information used for training in the negative dataset as RLN. 1:n This notation has become commonplace.

[0049] The training data creation unit 35 can, for example, create a positive dataset first and then create a negative dataset.

[0050] First, the time T at which the water surface W transitioned to a state H where the water level W exceeded the flood height h is extracted from the rainfall data DA stored in the rainfall data storage unit 34.

[0051] Next, based on any one of the plurality of extracted times T, using the time series pattern of the step widths (i.e., information such as the plurality of step widths and the number of data of the learning rainfall time series information set by the explanatory variable setting unit 32) described with reference to FIG. 3, a plurality (in this embodiment, n = 6) of rainfall data in the time period before the time T are extracted. A set of these rainfall data is used as a set of learning rainfall time series information RLP of the positive data set. 1:6 Let it be so. Also, in association with the corresponding learning flooding probability information PLP (= 1), a positive data set (RLP 1:6 , PLP) is created at the time T. This process is repeated at the plurality of extracted times T, and finally a plurality of sets of positive data sets are created.

[0052] Next, using the plurality of sets of learning rainfall time series information RLP of the created positive data set 1:6 , in the rainfall data DA, intervals that are similar to the characteristics of these learning rainfall time series information RLP 1:6 (for example, the rainfall in each time period, the change amount of the rainfall between time periods, etc.) and in which the state L where the water surface W is below the flooding height h is maintained are extracted. In this embodiment, the "interval" refers to a continuous period of time from one time to another time. A set of rainfall data extracted in this way is used as a set of learning rainfall time series information RLN of the negative data set 1:6 . Note that the time series pattern of the step widths between the rainfall data of the learning rainfall time series information RLN 1:6 is the same as the time series pattern of the step widths described with reference to FIG. 3, similar to the positive data set. Note that the time from after flood detection to flood elimination cannot detect a new flood occurrence, and the data cannot be determined as positive or negative. Therefore, it is desirable to exclude rainfall data with these times set to "0 minutes" from the creation of the data set.

[0053] A set of learning rainfall time series information RLN 1:6 that is extracted is associated with the corresponding learning flooding probability information PLN (= 0), and a set of negative data sets (RLN 1:6Create a PLN (Patent Linear Negative). Repeat this process to ultimately create multiple sets of negative datasets.

[0054] Returning to Figure 2, the model learning unit 36 ​​uses the training dataset created by the training data creation unit 35 to train the prediction model 4 and generate a trained model 4A. The model learning unit 36 ​​uses the training rainfall time series information RL from the training data as input (explanatory variable) to the prediction model 4 and the training flood probability information PL as output (dependent variable) to the prediction model 4, and performs machine learning so that the input / output relationship of the prediction model 4 approaches the input / output relationship of the training dataset, thereby generating a trained model 4A. In this embodiment, it is preferable to use logistic regression analysis as the machine learning method, and it is preferable that the prediction model 4 and the trained model 4A are regression models obtained by logistic regression analysis.

[0055] In this embodiment, logistic regression analysis is applied to the machine learning of the model learning unit 36 ​​for the following reasons. (a) The prediction result is not a binary result of whether or not flooding will occur, but rather the "probability" of flooding occurring. Therefore, flexible operation is possible, for example, by changing the probability threshold (cutoff value) for determining whether road flooding will occur depending on factors such as traffic volume and time of day. (b) Interpretation of the basis for the prediction results, specifically, the influence of rainfall on the probability of flooding can be directly confirmed.

[0056] The logistic regression analysis model in this embodiment is shown in equation (1).

[0057]

number

[0058] Here, p(Y=1) represents the probability of flooding occurring. p(Y=1) corresponds to the PLP (Planetary Flooding Probability Information) used for training in the positive dataset mentioned above. x1, x2, ..., x n x1, x2, ..., x n This refers to the above-mentioned learning rainfall time series information RLN1:n Corresponds to β1, β2, ..., β n This represents the partial regression coefficient.

[0059] Let P be the probability of flooding occurring (Y=1). Then, P can be expressed as the log odds log. e The transformation to (P / (1-P)) is called a logit transformation. By performing a logit transformation on P, the probability between 0 and 1 can be transformed into a continuous value from -∞ to +∞. Equation (2) is obtained from equation (1) by the inverse logit transformation (logistic transformation).

[0060]

number

[0061] The maximum likelihood method is used to estimate the partial regression coefficients. The partial regression coefficients (maximum likelihood estimators) that maximize the log-likelihood function shown in equation (3) can be obtained by numerical calculations such as the Newton-Raphson method.

[0062]

number

[0063] Here, L represents the likelihood function.

[0064] By applying the partial regression coefficients and explanatory variables obtained by equation (3) to equation (2), the probability of road flooding occurring, P, can be calculated.

[0065] The reason for adopting logistic regression analysis is that the basis for the prediction results can be interpreted as follows. The right-hand side of equation (1) is a linear combination of the explanatory variables, which are the average rainfall intensities at different times. In this case, as the right-hand side increases, equation (2) shows that the probability of occurrence increases. More precisely, x i When increases by 1, P / ((1-P)) (called the odds) becomes e βi It doubles. Thus, the explanatory variable β i This allows us to interpret the impact that has on the odds.

[0066] Furthermore, in this embodiment, by including the aforementioned positive and negative datasets in the training dataset, it is possible to reflect rainfall time series data similar to that during flooding, but in situations where flooding did not actually occur, in the training of the prediction model 4. This improves the training accuracy of the prediction model 4 and the prediction accuracy of the flooding probability prediction device 5 using the trained model 4A.

[0067] Next, the flood probability prediction device 5 will be described with reference to Figure 5. Figure 5 is a functional block diagram of the flood probability prediction device 5 according to the first embodiment.

[0068] As shown in Figure 5, the flood probability prediction device 5 includes a rainfall data acquisition unit 51 for prediction, a rainfall data storage unit 52 for prediction, a rainfall time-series information creation unit 53 for prediction, and a prediction unit 54. The flood probability prediction device 5 can also be described as having a trained model 4A.

[0069] The trained model 4A uses the aforementioned training rainfall time series information RL. 1:n This model uses machine learning to obtain the correspondence between training explanatory variables, which include the training inundation probability information PL, and training target variables, which include the training inundation probability information PL. The trained model 4A uses the rainfall time series information Re for prediction. 1:n Based on the input of explanatory variables including the input variables, the dependent variable (flood occurrence probability Pe) related to the input explanatory variables is output.

[0070] The rainfall data acquisition unit 51 acquires time-series data (rainfall data for prediction) of rainfall at specific locations P1 to Pm for which the probability of flooding is predicted, prior to the time of prediction. If the rainfall data for prediction is for a time when rainfall observed before the time of prediction is performed, the rainfall data acquisition unit 51 can be configured to acquire the rainfall data for prediction from outside the flood probability prediction system 1, for example, by collecting the rainfall data for prediction for specific locations P1 to Pm via an internet connection from observation data of the real-time rainfall observation system (XRAIN) using high-performance weather radar operated by the Ministry of Land, Infrastructure, Transport and Tourism, similar to the rainfall data acquisition unit 31. Alternatively, the flood probability prediction system 1 can be equipped with rain gauges installed at specific locations P1 to Pm, and the rainfall data acquisition unit 51 can acquire the rainfall data for prediction measured by the rain gauges, thus acquiring the rainfall data for prediction using internal elements of the flood probability prediction system 1. For periods between the time the forecast is to be executed and the desired forecast time (which may also be referred to as "the time the forecast will be performed" or "the time the forecast will be performed"), when no rainfall data has been observed, predicted rainfall data such as high-resolution precipitation nowcasts or short-term precipitation forecasts can be obtained and used. For example, if the current time is 10:00 and the desired forecast time is 10:30, observed rainfall data can be obtained up to 10:00, but predicted rainfall data will be needed from then until 10:30.

[0071] The rainfall forecast data storage unit 52 stores and holds the rainfall forecast data DB acquired by the rainfall forecast data acquisition unit 51.

[0072] The rainfall time series information generation unit 53 generates the rainfall time series information Re, which is an explanatory variable to be input into the trained model 4A. 1:nThe rainfall time series information creation unit 53 creates a predictive rainfall time series information database. Based on information such as the time series pattern of the step size and the number of data points of the explanatory variables (predictive rainfall time series information) set by the explanatory variable setting unit 32 of the model generation device 3, the predictive rainfall data storage unit 52 extracts multiple (six in this embodiment) rainfall data for a time period prior to the time the prediction is made from the predictive rainfall data DB. The predictive rainfall time series information creation unit 53 then combines the extracted rainfall data into a set of predictive rainfall time series information Re 1:n This is output to the prediction unit 54.

[0073] The prediction unit 54 receives the rainfall time series information Re for prediction from the trained model 4A. 1:n The system takes explanatory variables, including the input variables, as input and outputs the target variable related to the input explanatory variables. The prediction unit 54 predicts the probability of flooding occurrence Pe at the time the prediction is performed, based on the target variable output by the trained model 4A.

[0074] Figure 6 is a hardware configuration diagram of the model generation device 3 and the flood probability prediction device 5.

[0075] As shown in Figure 6, the model generation device 3 and flood probability prediction device 5 according to the first embodiment can be physically configured as a computer system including a CPU (Central Processing Unit) 101, a GPU (Graphics Processing Unit) 108, main memory such as RAM (Random Access Memory) 102 and ROM (Read Only Memory) 103, input devices such as a keyboard and mouse 104, output devices such as a display 105, a communication module 106 which is a data transmission and reception device such as a network card, and auxiliary storage devices such as a hard disk 107.

[0076] Each function of the model generation device 3 shown in Figure 2 is realized by loading predetermined computer software (model generation program) onto hardware such as the CPU 101 and RAM 102, thereby operating the communication module 106, input device 104, and output device 105 under the control of the CPU 101, and reading and writing data to the RAM 102 and auxiliary storage device 107. In other words, by executing the model generation program of the first embodiment on a computer, the model generation device 3 functions as the rainfall data acquisition unit 31, explanatory variable setting unit 32, detection information storage unit 33, rainfall data storage unit 34, training data creation unit 35, model learning unit 36, and prediction model 4 shown in Figure 2.

[0077] Similarly, each function of the flood probability prediction device 5 shown in Figure 5 is realized by loading predetermined computer software (flood probability prediction program) onto hardware such as the CPU 101 and RAM 102, thereby operating the communication module 106, input device 104, and output device 105 under the control of the CPU 101, and reading and writing data to the RAM 102 and auxiliary storage device 107. In other words, by executing the flood probability prediction program of the first embodiment on a computer, the flood probability prediction device 5 functions as the rainfall data acquisition unit 51 for prediction, the rainfall data storage unit 52 for prediction, the rainfall time series information creation unit 53 for prediction, the prediction unit 54, and the trained model 4A shown in Figure 5.

[0078] The model generation program and flood probability prediction program according to the first embodiment are stored, for example, in a storage device provided by a computer. Alternatively, part or all of the model generation program and flood probability prediction program may be transmitted via a transmission medium such as a communication line and received and recorded (including installation) by a communication module 106 provided by the computer. Furthermore, part or all of the model generation program and flood probability prediction program may be stored on a portable storage medium such as a CD-ROM, DVD-ROM, or flash memory, and then recorded (including installation) into the computer.

[0079] Similarly, the trained model 4A generated by the model generation device 3 may be stored on a storage medium and carried around, transmitted via a transmission medium, or recorded in a computer.

[0080] Next, with reference to Figure 7, the model generation method according to the first embodiment will be described. Figure 7 is a flowchart of the model learning control according to the first embodiment. Each process in the flowchart shown in Figure 7 is performed by the model generation device 3 according to the first embodiment.

[0081] In step S101, the rainfall data DA acquired by the rainfall data acquisition unit 31 is stored in the rainfall data storage unit 34. The rainfall data acquisition unit 31 creates rainfall data DA by averaging XRAIN observation data every 10 minutes and for every 1km mesh, and then stores the created rainfall data DA in the rainfall data storage unit 34. In this embodiment, the rainfall data DA is stored in the rainfall data storage unit 34 for a 1km mesh area that includes specific locations P1 to Pm where the water detector 2 is set.

[0082] In step S102, the explanatory variable setting unit 32 sets the time series pattern of the step size for the explanatory variables of the prediction model 4 (i.e., the learning rainfall time series information RL). The explanatory variable setting unit 32 can set the time series pattern of the step size based on command values ​​that are set in advance, for example by user input. In this embodiment, as explained with reference to Figure 3, for example, in the time period before a reference time such as the flood detection time, step sizes of 10 minutes x 4, 3 hours, and 9 hours are set in order from the closest to the reference time. As a result, the number of explanatory variables is set to 6.

[0083] In step S103, the process of acquiring data to be used to create training data is started, and the processes in steps S104 to S106 are repeated until it is determined in step S107 that the acquisition of a predetermined number of data has been completed. The processes in steps S103 to S107 are carried out by the detection information storage unit 33, and the acquired data corresponds to the "detection information" described above.

[0084] In the data acquisition process, when the water detector 2 detects a water surface W at flood height h (YES in step S103), the information of the state H in which the water surface W exceeds the flood height h and the time T when the transition to state H occurs are linked and recorded in the detection information storage unit 33.

[0085] On the other hand, when the water detector 2 does not detect the water surface W at flood height h (NO in step S103), the detection information storage unit 33 records the information of the state L in which the water surface W is lower than the flood height h and the time T when the transition to state L occurred, linking them together.

[0086] The process in steps S104 to S106 is repeated until the number of datasets stored in the detection information storage unit 33 reaches a predetermined number, at which point the process proceeds to step S108 and beyond. In this embodiment, the "determined number" of datasets stored in the detection information storage unit 33 is not an arbitrary fixed number, but rather a number that can be arbitrarily changed to allow the creation of any desired number of positive datasets. For example, the condition for determining that the "determined number has been reached" can be the number of occurrences of an event in the acquired data that transitions from state L to state H (i.e., an event in which the water level W rises and the flood height exceeds h, resulting in flooding). In other words, if the frequency of flooding occurrences is relatively low, the number of datasets stored in the detection information storage unit 33 will be relatively large, while if the frequency of flooding occurrences is relatively high, the number of datasets stored in the detection information storage unit 33 will be relatively small. By setting a predetermined number of datasets in this way, it becomes possible to create any desired number of positive datasets for machine learning, thereby stabilizing the learning process.

[0087] In step S108, the training data creation unit 35 acquires the detection information recorded in steps S105 and S106, and the rainfall data DA saved in step S101.

[0088] In step S109, the training data creation unit 35 extracts time-series data with the step size set in step S102 from the rainfall data DA acquired in step S108 for the section including the flood detection time, and uses this extracted rainfall time-series data to create the learning rainfall time-series information RLP for positive data. 1:n This is created.

[0089] In step S110, the training data creation unit 35 generates the training rainfall time series information RLP created in step S109. 1:n The corresponding learning flood probability information (flood occurrence probability) PLP is set to 1.

[0090] In step S111, the training data creation unit 35 uses the section of rainfall data DA acquired in step S108 where no flooding was detected to create the training rainfall time series information RLP of the positive data created in step S109. 1:n Extract similar time-series data and use this extracted rainfall time-series data to create a training rainfall time-series information RLN for negative data. 1:T This is created.

[0091] Negative data is extracted from rainfall data DA based on a time series pattern with the same step size as the positive data. One example of a method for extracting negative data is as follows: First, among the observation times accumulated in the rainfall data DA, the time of flooding (the "0 minutes") is excluded because it is positive data. Furthermore, the time from the detection of flooding until the flooding is resolved is also excluded because it is not possible to detect the occurrence of new flooding during this time, so it is not judged as either positive or negative data. The remaining data in the rainfall data DA after these exclusions become candidates for the "0 minutes" (reference time) for negative data. Next, using the rainfall immediately preceding the positive data, the value of rainfall that is likely to cause flooding (threshold) is calculated from the rainfall in the tens of minutes immediately preceding the occurrence of flooding. For example, this could be the average value of the rainfall immediately preceding the flooding or the smallest of the maximum values ​​for each of the multiple positive data. Time series data that include rainfall above the threshold obtained in this way in the rainfall immediately preceding the "0 minutes" (reference time) are designated as negative data.

[0092] In step S112, the training data creation unit 35 generates the training rainfall time series information RLN created in step S111. 1:n The corresponding training flood probability information (flood occurrence probability) PLN is set to 0.

[0093] In step S113, the training data creation unit 35 generates the training rainfall time series information RLP created in step S109. 1:n Then, a positive dataset is created by associating it with the training flood occurrence probability PLP set in step S110.

[0094] In step S114, the training data creation unit 35 generates the RLN created in step S111. 1:n Then, a negative dataset is created by associating it with the training flood probability information PLN set in step S112.

[0095] In step S115, the training data creation unit 35 combines the positive dataset created in step S113 and the negative dataset created in step S114 to create a training dataset.

[0096] In step S116, the model learning unit 36 ​​trains the prediction model 4 using supervised machine learning with the training dataset created in step S115 (model learning step). The model learning unit 36 ​​can apply various algorithms as supervised machine learning, such as regression, decision trees, support vector machines, and neural networks. The model learning unit 36 ​​only needs to be able to train the prediction model 4 using supervised learning, and there are no restrictions on the algorithm that can be applied, but in this embodiment, logistic regression analysis is applied as explained with reference to equations (1) to (3) above.

[0097] Next, with reference to Figure 8, the flood probability prediction method according to the first embodiment will be described. Figure 8 is a flowchart of the flood probability prediction control according to the first embodiment. Each process in the flowchart shown in Figure 8 is performed by the flood probability prediction device 5 according to the first embodiment.

[0098] In step S201, the rainfall prediction data DB acquired by the rainfall prediction data acquisition unit 51 is stored in the rainfall prediction data storage unit 52. Similar to step S101 in Figure 7, the rainfall prediction data acquisition unit 51 creates a rainfall prediction data DB by averaging XRAIN observation data every 10 minutes and for every 1km mesh, and then stores the created rainfall prediction data DB in the rainfall prediction data storage unit 52. In this embodiment, the rainfall prediction data DB is stored in the rainfall prediction data storage unit 52 for a 1km mesh area that includes specific locations P1 to Pm where the water detector 2 is set.

[0099] Furthermore, if the flooding probability prediction device 5 is configured to predict the probability of flooding occurring at a time later than the current time, in step S201, the rainfall data acquisition unit 51 acquires predicted rainfall data, such as high-resolution precipitation nowcasts or short-term precipitation forecasts, for at least the interval up to the time the prediction is made, and stores the predicted rainfall data in the rainfall data database in the rainfall data storage unit 52.

[0100] In step S202, the trained model 4A is set (model setting step). The trained model 4A can be one that has been pre-generated by the model generation device 3, or one that has been obtained by other methods.

[0101] In step S203, the rainfall time series information creation unit 53 obtains information on the time series pattern of the step size of the explanatory variables (i.e., the rainfall time series information Re for prediction) used in the model learning control described with reference to the flowchart in Figure 7, from the explanatory variable setting unit 32 of the model generation device 3. In other words, in this embodiment, the step size and number of the two types of explanatory variables, the learning rainfall time series information RL in the model generation control and the rainfall time series information Re for prediction in the flood probability prediction control, are made identical.

[0102] In step S204, the rainfall forecast time-series information creation unit 53 retrieves the rainfall forecast data stored in the rainfall forecast data storage unit 52 in step S201.

[0103] In step S205, the rainfall time series information creation unit 53 extracts time series data with a step size set in step S203 from the rainfall data DB for prediction acquired in step S204 for the section including the current time, and uses the extracted time series data to create the rainfall time series information Re 1:n This is created.

[0104] In step S206, the prediction unit 54 uses the explanatory variables created in step S205 (i.e., the rainfall time series information Re for prediction) to process the explanatory variables. 1:nThe data is input to the trained model 4A, and the target variable is output.

[0105] In step S207, the prediction unit 54 calculates the flooding probability Pe based on the target variable of the trained model 4A acquired in step S206.

[0106] The flood probability prediction device 5 can perform the processing shown in the flowchart of Figure 8 at any time, for example, when new observation data is input from an external source such as XRAIN and the rainfall data database for prediction is updated. In addition, in the case of a configuration that predicts the probability of flooding occurring at a time later than the current time, the processing shown in the flowchart of Figure 8 may be triggered by, for example, the acquisition of information on the "time at which the prediction should be performed" through user input.

[0107] Figure 9 shows an example of the flooding probability Pe prediction results according to the first embodiment. In Figure 9, the horizontal axis represents the date and time, and the vertical axis represents the water level (m), flooding probability, and rainfall intensity (mm / h). In Figure 9, the rainfall intensity for 10 minutes at each date and time is shown as a bar graph, the water level is shown as a dotted line graph, and the flooding probability prediction results are shown as a solid line graph. The flooding probability and rainfall intensity in Figure 9 correspond to the flooding probability Pe and prediction rainfall data DB shown in Figure 5, etc.

[0108] In Figure 9, the water level is the height of the water surface W at the flood height h, with the road surface G shown in Figure 1 being 0 (m). Therefore, in the example in Figure 9, as shown by the dotted line graph, for the period prior to "February 1, 2021, 4:50," the water level is 0 (m) or less, meaning that the water detector 2 outputs a state L where the water surface W is below the flood height h (0.2m), and no flooding has occurred. On the other hand, for the period from "February 1, 2021, 5:00" to "February 1, 2021, 5:10," the water level increases above 0 (m), and the water detector 2 transitions to outputting a state H where the water surface W exceeds the flood height h (0.2m), thus detecting the occurrence of flooding. In the subsequent period, the water level increases further, so the water detector 2 continues to output a state H where the water surface W exceeds the flood height h, indicating that flooding is continuing.

[0109] Furthermore, in the example shown in Figure 9, as indicated by the bar graph, the rainfall intensity (mm / h) increased sharply after the date and time of "February 1, 2021, 4:00," and it can be said that the rainfall during this period corresponds to the "immediate rainfall" explained with reference to Figure 3.

[0110] In the example shown in Figure 9, as indicated by the solid line graph, the probability of flooding is below 0.5 during the period up to "February 1, 2021, 4:20," when no rise in water level is observed. However, it reaches 0.5 at "February 1, 2021, 4:30" and continues to exceed 0.5 until "February 1, 2021, 7:30." Furthermore, as mentioned above, the probability of rainfall rises to a value close to 1 during the period from "February 1, 2021, 5:00" to "February 1, 2021, 5:10," when flooding is detected. Then, prior to "February 1, 2021, 6:50," when rainfall intensity begins to decrease, the probability of flooding remains at approximately 1.

[0111] In the example shown in Figure 9, it can be said that the increasing trend in the probability of flooding predicted by this embodiment matches the increasing trend in the water level. Furthermore, it can be said that the probability of flooding predicted by this embodiment begins to increase earlier than the timing at which flooding is actually detected. In other words, this embodiment demonstrates that the probability of flooding can be predicted with high accuracy.

[0112] Figure 10 shows an example of flood prediction and evaluation according to the first embodiment. In Figure 10, three locations A, B, and C were selected as specific locations for predicting the probability of flooding. Also, the subscript "40" is used for each location name. ※ " means that if the pre-rainfall (immediate rainfall) is set to 40 minutes, that is, as in the example in Figure 3, it is a configuration that creates four predictive rainfall time series information at 10-minute intervals with respect to a reference time such as the flood detection time. The subscript "30" for each location name. ※ This means that if the preceding rainfall (immediate rainfall) is set to occur within 30 minutes, that is, the system will create three time-series data points for predicting rainfall at 10-minute intervals relative to a reference time such as the time of flood detection.

[0113] Figure 10 shows the values ​​for each location as items: "Rainfall," "True," "False," "True Positive," "False Positive," "New Negative," "False Negative," "True Positive Rate," "False Positive Rate," "Precision Rate," and "Accuracy Rate." "Rainfall" is the time-series rainfall information used for prediction at each location. 1:n This corresponds to the number of occurrences. "True" and "False" are observation results regarding the presence or absence of flooding. Here, flooding is determined to have occurred (positive) when the flooding probability Pe is 0.5 or higher, and to not have flooding (negative) when it is less than 0.5. In this case, the value 0.5 is the probability value (cutoff value) that distinguishes between the presence or absence of flooding.

[0114] "True positive" refers to the number of locations where flooding was observed and predicted, out of the total rainfall. "False positive" refers to the number of locations where flooding was predicted but did not actually occur, out of the total rainfall. "True negative" refers to the number of locations where no flooding was predicted and no flooding actually occurred, out of the total rainfall. "False negative" refers to the number of locations where no flooding was predicted but flooding actually occurred, out of the total rainfall.

[0115] The "true positive rate" is the ratio of true positives to true positives (true positives / true), indicating the proportion of times actual flooding was correctly predicted as flooding. The "false positive rate" is the ratio of false positives to false positives (false positives / false), indicating the proportion of times non-flooding data was predicted as flooding. The "precision rate" is the proportion of true positives among the positives that were determined to be flooding (true positives / (true positives + false positives)). The "accuracy rate" is the ratio of the number of correct answers (sum of true positives and true negatives) to the number of data points (number of rainfall events).

[0116] As shown in Figure 10, at all locations A, B, and C, the true positive rate was 0.9 to 1.0 regardless of whether the preceding rainfall was 30 minutes or 40 minutes, indicating a very high rate of correct prediction of actual flooding.

[0117] Figure 11 shows an example of the occurrence rate of cases with a flood probability of 0.5 or higher. The horizontal axis of Figure 11 shows the flood probability, and the vertical axis shows the occurrence rate. In addition, Figure 11 shows two bar graphs representing the occurrence rate of true positives and false positives for each range of flood probability in 0.1 increments: "0.5~0.6", "0.6~0.7", "0.7~0.8", "0.8~0.9", and "0.9~1.0".

[0118] As shown in Figure 11, in the range of "0.9 to 1.0" for the probability of flooding, true positives (actual flooding occurrences) were significantly higher than false positives (predictions of flooding occurrences even though no flooding occurred).

[0119] [Second Embodiment] A second embodiment will be described with reference to Figures 12 to 15.

[0120] Figure 12 is a functional block diagram of the model generation device 3A according to the second embodiment. The model generation device 3A according to the second embodiment differs from the explanatory variable setting unit 32 of the first embodiment in the configuration of the explanatory variable setting unit 32A.

[0121] The explanatory variable setting unit 32A creates rainfall time series information RL using multiple time series patterns with different step sizes, performs "preliminary training" of the prediction model 4 for each time series pattern of the created rainfall time series information RL, and sets the time series pattern of the step size based on the results of the preliminary training. In the second embodiment, "multiple time series patterns with different step sizes" means that in addition to multiple combinations of multiple types of step sizes, multiple types of number of step sizes can also be set.

[0122] Figure 13 shows examples of time-series patterns with multiple step sizes set by the explanatory variable setting unit 32A of the second embodiment. Figure 12 illustrates four pattern examples, (A) to (D). Figure 13(A) is the same as the time-series pattern of the first embodiment described with reference to Figure 3. That is, four data points are extracted by dividing the 40 minutes of rainfall immediately preceding the flood detection time, including the flood detection time, into 10-minute intervals, and two data points are extracted by dividing the time before the immediate rainfall into 3-hour and 9-hour intervals, for a total of six data points.

[0123] In the example pattern shown in Figure 13(B), the time interval for the preceding rainfall is set to be longer than in (A). Specifically, five data points are extracted by dividing the 50 minutes of preceding rainfall before the flood detection time into 10-minute intervals, and two data points are extracted for the time before the preceding rainfall into 3-hour and 9-hour intervals, for a total of seven data points.

[0124] In the example pattern in Figure 13(C), the time interval for the preceding rainfall is set to be even longer than in (B). Specifically, the 60 minutes of rainfall immediately preceding the time of flooding detection, including the time of flooding detection, is divided into 10-minute intervals, and 6 data points are extracted. For time periods prior to the preceding rainfall, 2 data points are extracted in 3-hour and 9-hour intervals, for a total of 8 data points.

[0125] In the example pattern in Figure 13(D), the time interval for the immediate rainfall is the same as in (A), but the time intervals prior to that are set to be longer than in (A) to (C). Specifically, four data points are extracted by dividing the 40 minutes of the immediate rainfall preceding the flood detection time, including the flood detection time, into 10-minute intervals, and two data points are extracted by dividing the time prior to the immediate rainfall into 4-hour and 10-hour intervals, for a total of six data points.

[0126] In addition, in the example pattern in Figure 13(D), the time interval for the preceding rainfall may be changed, similar to (B) and (C). Furthermore, the time interval prior to the preceding rainfall may be set to a length other than 4 hours or 10 hours.

[0127] In the "preliminary learning" described above, the explanatory variable setting unit 32A uses multiple patterns of rainfall time series information RL, created as shown in the example pattern in Figure 13, to perform machine learning of the prediction model 4 using the training data creation unit 35 and the model learning unit 36. In the preliminary learning, for example, the training data creation unit 35 prepares training data and evaluation data for each of the multiple patterns of rainfall time series information RL. Then, the model learning unit 36 ​​performs machine learning of the prediction model 4 individually for each of the multiple patterns of rainfall time series information RL using the training data, under the same conditions as the model learning unit 36 ​​in the first embodiment. As a result, the same number of trained models 4A as the number of patterns of rainfall time series information RL are generated. Subsequently, the prediction performance of the multiple trained models 4A is evaluated using the evaluation data. As an evaluation method, for example, for each of the multiple patterns of rainfall time series information RL, the explanatory variables of the evaluation data are input to the trained model 4A corresponding to each pattern, and the output of the model is compared with the target variable of the evaluation data. For example, the "true positive rate" explained with reference to Figure 10 is calculated for each of the multiple patterns of rainfall time series information RL. Then, the one with the highest true positive rate is selected as the optimal rainfall time series information RL. Note that the parameters related to the multiple patterns of rainfall time series information RL used for preliminary training may include information such as cutoff values ​​for distinguishing between positive and negative results, in addition to the various combinations of step sizes and the number of extracted data points exemplified in Figure 13.

[0128] The time-series pattern of the step size set by the explanatory variable setting unit 32A for preliminary learning may be determined based on a command value that has been set in advance, for example, by user input. Alternatively, a minimum condition may be set, such as keeping the number of explanatory variables less than the number of data points at the time of flooding, extracting multiple rainfall information from the most recent rainfall, and extracting rainfall up to a maximum of a certain number of hours prior to the most recent rainfall, and the step size and number of data points may be automatically determined to satisfy this minimum condition.

[0129] The explanatory variable setting unit 32A can receive feedback from the model learning unit 36 ​​regarding the results of preliminary learning, as shown in Figure 12, for example, and determine the optimal time series pattern for the step size of the rainfall time series information for learning based on the feedbacked information. The preliminary learning results information fed back from the model learning unit 36 ​​may include combinations of multiple trained models 4A corresponding to the multiple patterns of rainfall time series information RL mentioned above, and evaluation data, and the explanatory variable setting unit 32A may perform the evaluation. Alternatively, the model learning unit 36 ​​may perform the evaluation of multiple patterns of rainfall time series information RL, and the preliminary learning results information may include the information for the optimal rainfall time series information RL.

[0130] The explanatory variable setting unit 32A outputs information on the time series pattern with the determined optimal step size to the training data creation unit 35. Based on the time series pattern information with the step size input from the explanatory variable setting unit 32A, the training data creation unit 35 creates training data in the same manner as in the first embodiment.

[0131] Figure 14 illustrates a modified example of the explanatory variable setting unit 32A. In the second embodiment, the explanatory variable setting unit 32A is shown as a configuration in which the time series pattern of the step size of the learning rainfall time series information is arbitrarily set in the region where the water detector 2 is installed (region R0 in Figure 14). However, it is also possible to set other elements other than the step size arbitrarily, as long as these elements can influence the learning of the prediction model 4.

[0132] In the example shown in Figure 14, the explanatory variable setting unit 32A can arbitrarily set the rainfall measurement locations for the learning rainfall time series information RL. For example, the system may be configured to surround area R0 where the water detector 2 is installed from all sides, and set up eight other areas R1 to R8 that have approximately the same area as area R0, and then select the rainfall measurement location from among these eight areas R1 to R8. In the example shown in Figure 14, multiple areas R1 to R8 are shown as having approximately the same area, but the areas of each area R1 to R8 may be different. Furthermore, instead of selecting one measurement location from multiple areas R1 to R8, the system may use a mix of the measurement results from multiple areas R1 to R8 (for example, the average or weighted average of rainfall).

[0133] In the example shown in Figure 14, the water detector 2 is installed on road B adjacent to river A. River A has its upstream side in the upper region R2 in the figure relative to region R0, and its downstream side in the left region R8→R7 and the lower region R6→R5 in the figure. In this example, if the amount of rainfall increases in region R2, where the upstream side of river A is located, and the river overflows, it is likely that flooding will also occur in the downstream region R0. Therefore, in the example shown in Figure 14, the explanatory variable setting unit 32A can set region R2, where the upstream side of river A is located, as the rainfall measurement location for the learning rainfall time series information RL, instead of region R0 where the water detector 2 is set.

[0134] In the modified configuration, the explanatory variable setting unit 32A creates rainfall time-series information for a section (region R0) containing a specific location P0 to Pm, and for multiple sections (regions R1 to R8) surrounding that section. Based on the created rainfall time-series information, the model 4 described above is pre-trained for each of the multiple sections, and based on the results of the pre-training, one of the multiple sections is set as the measurement location. Here, "section" refers to each individual part of a plot of land or similar area (in this case, regions R1 to R8).

[0135] In the example shown in Figure 14, the training data creation unit 35 can create training rainfall time series information RL based on rainfall information for measurement locations (i.e., any of the regions R1 to R8) set by the explanatory variable setting unit 32A.

[0136] As explained with reference to Figures 12 to 14, the explanatory variable setting unit 32A is configured to select and determine time-series patterns and measurement locations with arbitrary rainfall step sizes, allowing the prediction model 4 to be trained using information that is influenced by whether or not flooding occurs at specific locations P0 to Pm. This further improves the training accuracy of the prediction model 4 and the prediction accuracy of the flood probability prediction device 5A that uses the trained model 4A.

[0137] In addition, similar to the first embodiment, the explanatory variable setting unit 32A may be configured to set a predetermined rainfall measurement location (one of the regions R0 to R8 in Figure 14) without performing preliminary learning.

[0138] Figure 15 is a functional block diagram of the flood probability prediction device 5A according to the second embodiment. The flood probability prediction device 5A according to the second embodiment differs from the first embodiment in that it has a flood occurrence determination unit 55. It also differs from the first embodiment in that the rainfall time series information creation unit 53 for prediction creates rainfall time series information Re for prediction based on information such as the time series pattern of the step size of the rainfall information selected by the explanatory variable setting unit 32A according to the second embodiment, as explained with reference to Figures 12 to 14, and the measurement location.

[0139] The flood occurrence determination unit 55 determines whether or not new flooding will occur at specific locations P0 to Pm after the time of prediction, based on the flood occurrence probability Pe predicted by the prediction unit 54. For example, the flood occurrence determination unit 55 can determine that flooding will occur when the flood occurrence probability Pe predicted by the prediction unit 54 exceeds a predetermined threshold. This "predetermined threshold" can be, for example, the cutoff value (e.g., 0.5) explained with reference to Figure 10.

[0140] Furthermore, the flood occurrence determination condition by the flood occurrence determination unit 55 may be that the cutoff value is maintained for a predetermined time. The numerical values ​​of the cutoff value and predetermined time can be changed. For example, determination conditions can be set such as maintaining a cutoff value of 0.5 or higher for 30 seconds or more, or maintaining a cutoff value of 0.7 or higher for 20 seconds or more. The numerical values ​​of the cutoff value and predetermined time can be determined, for example, for each specific location P0 to Pm targeted for flood prediction, according to the time of day, road traffic volume, land drainage capacity, etc.

[0141] By providing the flood occurrence determination unit 55 in this way, it becomes possible to allow users to intuitively recognize whether or not there is a possibility of flooding occurring, thereby improving convenience.

[0142] The embodiments have been described above with reference to specific examples. However, this disclosure is not limited to these specific examples. Modifications made to these specific examples by those skilled in the art are also included within the scope of this disclosure, as long as they retain the features of this disclosure. The elements, their arrangement, conditions, shapes, etc., of each of the aforementioned specific examples are not limited to those illustrated and can be modified as appropriate. The elements of each of the aforementioned specific examples can be combined in different ways as appropriate, as long as no technical inconsistencies arise. [Explanation of Symbols]

[0143] 1. Flooding Probability Prediction System 2 Water detectors 3.3A Model Generation Device 32, 32A Explanatory variable setting section 35 Training Data Creation Department 36 Model Learning Department 4. Predictive Models 4A Pre-trained model 5.5A Flooding Probability Prediction Device 54 Prediction Section 55 Flood occurrence determination unit

Claims

1. A model generation device for generating a model to predict the probability of new flooding occurring at a specific location after a predetermined time, A water detector installed at the aforementioned specific location detects whether or not flooding has occurred at the aforementioned specific location by detecting a change in the presence or absence of water at a predetermined water level, A training data creation unit creates a training dataset that includes training rainfall time series information relating to the rainfall at a specific location based on the detection results of the water detector, and training flood probability information relating to the probability that new flooding will occur at the specific location after the latest point in time of the training rainfall time series information. A model learning unit generates a trained model by using the training dataset created by the training data creation unit, taking the training rainfall time series information as input to the model and the training flood probability information as output to the model, and performing machine learning so that the input / output relationship of the model approaches the input / output relationship of the training dataset. Equipped with, Model generation device.

2. The aforementioned Teacher dataset includes a positive dataset and a negative dataset. In the aforementioned positive dataset, the input is rainfall time-series information extracted from past rainfall data, going back from the time when flooding was detected by the water detector to a predetermined time prior, and the probability of flooding occurring in the output is 1. In the negative dataset, the input is time-series rainfall information from past rainfall data for sections where flooding has not been detected by the water detector, and the probability of flooding occurring in the output is 0. The model generation apparatus according to claim 1.

3. The unit includes an explanatory variable setting unit for setting the time series pattern of the step size of the learning rainfall time series information, The training data creation unit creates the learning rainfall time series information based on the time series pattern of the step size set by the explanatory variable setting unit. A model generation apparatus according to claim 1 or 2.

4. The explanatory variable setting unit creates rainfall time series information using multiple time series patterns with different step sizes, performs preliminary training on the model for each time series pattern using the created rainfall time series information, and sets the time series pattern with the specified step size based on the results of the preliminary training. The model generation apparatus according to claim 3.

5. The system includes an explanatory variable setting unit for setting the measurement locations of the rainfall in the aforementioned learning rainfall time-series information, The training data creation unit creates the learning rainfall time series information based on the rainfall information of the measurement location set by the explanatory variable setting unit. A model generation apparatus according to claim 1 or 2.

6. The explanatory variable setting unit creates rainfall time series information for a section including the specific location and for multiple sections surrounding the section, performs preliminary training of the model for each of the multiple sections using the created rainfall time series information, and sets one of the multiple sections as the measurement location based on the results of the preliminary training. The model generation apparatus according to claim 5.

7. The machine learning of the aforementioned model learning unit uses logistic regression analysis. The model generation apparatus according to claim 1.

8. A flood probability prediction device for predicting the probability of new flooding occurring at a specific location after a predetermined time, A trained model obtained by machine learning the correspondence between a training explanatory variable, which includes training rainfall time series information relating to the time series information of rainfall at the specified location, and a training target variable, which includes training flood probability information relating to the probability of new flooding occurring at the specified location after the latest point in time of the training rainfall time series information. A prediction unit inputs explanatory variables into the trained model, including the time at which the prediction will be made and time-series information on rainfall at the specific location prior to the time at which the prediction will be made, and outputs information on the target variable related to the information of the explanatory variables, and predicts the probability after the time at which the prediction will be made based on the target variable. Equipped with, Flooding probability prediction device.

9. The system includes a flood occurrence determination unit that determines whether or not new flooding will occur at the specific location after the specified time, based on the probability predicted by the prediction unit. The flooding probability prediction device according to claim 8.

10. The flood occurrence determination unit determines that flooding will occur when the probability predicted by the prediction unit exceeds a predetermined threshold. The flooding probability prediction device according to claim 9.

11. A model generation method for generating a model to predict the probability of new flooding occurring at a specific location after a predetermined time, A detection step in which a water detector installed at the specified location detects whether or not flooding has occurred at the specified location by detecting a change in the presence or absence of water at a predetermined water level, A training data creation step to create a training dataset that includes training rainfall time series information relating to the rainfall at a specific location based on the detection results of the water detector, and training flood probability information relating to the probability that new flooding will occur at the specific location after the latest point in time of the training rainfall time series information, A model training step involves using the training dataset created in the training data creation step, taking the training rainfall time series information as input to the model and the training flood probability information as output to the model, performing machine learning so that the input-output relationship of the model approaches the input-output relationship of the training dataset, and generating a trained model; including, Model generation method.

12. A model generation program for generating a model to predict the probability of new flooding occurring at a specific location after a predetermined time, A detection function that detects whether or not flooding has occurred at the specified location by detecting a change in the presence or absence of water at a predetermined water level using a water detector installed at the specified location, A training data creation function that creates a training dataset including training rainfall time series information relating to the rainfall at a specific location based on the detection results of the water detector, and training flood probability information relating to the probability that new flooding will occur at the specific location after the latest point in time of the training rainfall time series information. A model learning function generates a trained model by using the training dataset created by the aforementioned training data creation function, taking the training rainfall time series information as input to the model and the training flood probability information as output to the model, and performing machine learning so that the input / output relationship of the model approaches the input / output relationship of the training dataset. including, Make the computer execute it. Model generation program.

13. A flood probability prediction method for predicting the probability of new flooding occurring at a specific location after a predetermined time, A model setting step involves setting up a trained model in which a correspondence relationship is obtained by machine learning between a training explanatory variable, which includes training rainfall time series information relating to the time series information of rainfall at the specified location, and a training target variable, which includes training flood probability information relating to the probability of new flooding occurring at the specified location after the latest point in time of the training rainfall time series information; A prediction step in which the trained model set in the model setting step is input to the trained model, which includes the time at which the prediction will be made and includes forecast rainfall time series information relating to the rainfall at the specific location before the time at which the prediction will be made, and outputs information of the target variable relating to the information of the explanatory variables, and predicts the probability after the time at which the prediction will be made based on the target variable, including, Methods for predicting the probability of flooding.

14. A flood probability prediction program for predicting the probability of new flooding occurring at a specific location after a predetermined time, A model setting function that sets up a trained model in which the correspondence between a training explanatory variable, which includes training rainfall time series information relating to the time series information of rainfall at the specified location, and a training target variable, which includes training flood probability information relating to the probability of new flooding occurring at the specified location after the latest point in time of the training rainfall time series information, is obtained by machine learning. A prediction function that inputs explanatory variables, including the implementation time for performing the prediction and time-series information on rainfall at a specific location prior to the implementation time, into the trained model set by the model setting function, outputs information on the target variable related to the information of the explanatory variables, and predicts the probability after the implementation time based on the target variable, Make the computer execute it. A program for predicting the probability of flooding.

15. A flood probability prediction system for predicting the probability of new flooding occurring at a specific location after a predetermined time, A model generation device for generating a model for predicting the probability of new flooding occurring at a specific location after a predetermined time point, A flood probability prediction device for predicting the probability that new flooding will occur at the specified location after a predetermined time, It is equipped with, The aforementioned model generation device is A water detector installed at the aforementioned specific location detects whether or not flooding has occurred at the aforementioned specific location by detecting a change in the presence or absence of water at a predetermined water level, A training data creation unit creates a training dataset that includes training rainfall time series information relating to the rainfall at a specific location based on the detection results of the water detector, and training flood probability information relating to the probability of new flooding occurring at the specific location after the latest point in time of the training rainfall time series information. A model learning unit generates a trained model by using the training dataset created by the training data creation unit, taking the training rainfall time series information as input to the model and the training flood probability information as output to the model, and performing machine learning so that the input / output relationship of the model approaches the input / output relationship of the training dataset. Equipped with, The flooding probability prediction device is A trained model obtained by machine learning the correspondence between a training explanatory variable, which includes training rainfall time series information relating to the time series information of rainfall at the specified location, and a training target variable, which includes training flood probability information relating to the probability of new flooding occurring at the specified location after the latest point in time of the training rainfall time series information. A prediction unit inputs explanatory variables into the trained model, including the time at which the prediction will be made and time-series information on rainfall at the specific location prior to the time at which the prediction will be made, and outputs information on the target variable related to the information of the explanatory variables, and predicts the probability after the time at which the prediction will be made based on the target variable. Equipped with, Flooding probability prediction system.

16. A trained model obtained by machine learning that shows the correspondence between a training explanatory variable, which includes training rainfall time series information for predicting the probability of new flooding occurring after a predetermined time point, for a specific location, and a training target variable, which includes training flooding probability information for predicting the probability of new flooding occurring after the latest time point of the training rainfall time series information at the specific location.

Citation Information

Patent Citations

  • Water level forecasting device for river

    JP1997256338A

  • Prediction method, prediction program and information processing apparatus

    JP2020134300A

  • Water detection sensor and water cell used in the same, and flooding detection method

    JP2023066488A