Pipe failure prediction device, pipe failure prediction method, and pipe failure prediction program
By clustering underground pipes and calculating failure likelihoods at the cluster level, the method addresses the challenge of understanding pipe failure impacts, optimizing repair planning, and predicting long-term maintenance needs.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-30
- Publication Date
- 2026-04-16
Smart Images

Figure 2026066217000001_ABST
Abstract
Description
Technical Field
[0001] This application claims the priority of each of the following provisional applications and incorporates them herein by reference. U.S. Provisional Patent Application No. 63 / 703,587 (filed on October 4, 2024), and U.S. Provisional Patent Application No. 63 / 703,685 (filed on October 4, 2024).
[0002] This disclosure relates to an apparatus, a method, and a program for predicting pipe failures.
Background Art
[0003] There is a strong demand for methods and systems for accurately predicting the probability of underground pipe failures. According to the degradation diagnosis technology of pipes (pipe segments) that utilizes artificial intelligence (AI) and environmental big data, by using the buried environment data, pipeline information, and leakage information related to the physical and chemical degradation of pipes, it is possible to accurately grasp the future degradation risk without directly checking the pipe body by excavation. The future degradation risk is provided on the user screen, enabling risk assessment.
[0004] However, in the prior art, since the likelihood of leakage (failure) is determined for each individual underground pipe with the end as the joint, it may not always be easy to understand the impact when leakage occurs.
Summary of the Invention
Means for Solving the Problems
[0005] This disclosure provides a method for clustering one or more underground pipes and predicting the leakage risk at the cluster level.
[0006] An apparatus for predicting the occurrence of failures in an underground piping network according to one aspect of the present disclosure includes: a reading module configured to read connection data, wherein the connection data includes data indicating one or more pipe segments and data indicating one or more valves, and the end of each pipe segment is connected to the end of another pipe segment or a valve; a clustering module configured to automatically generate data indicating one or more clusters from the connection data, wherein each cluster includes pipe segments that form a connection network between valves; a cluster failure likelihood calculation module configured to calculate the failure likelihood of each cluster by summing the failure likelihoods of the pipe segments included in each cluster; and a display module configured to display the data indicating the one or more clusters and the failure likelihood of each cluster on a display.
[0007] This makes it possible to display the likelihood of failure for each cluster separated by valves. Overall, it becomes possible to understand the impact of failures at a more practical level, and it becomes easier to apply repair plans to the cluster level that can be shut off by valves.
[0008] The loading module may be further configured to load population forecast data for each region, and the display module may be further configured to display the population forecast data in addition to data indicating the one or more clusters and the failure likelihood of each cluster. This makes it possible to visualize both the likelihood of failure and the population that is predicted to be affected if a failure occurs in the future.
[0009] The system may further include an individual failure likelihood calculation module configured to calculate the failure likelihood of the pipe segment based on machine learning of correlations between previously collected data.
[0010] The individual failure likelihood calculation module may be configured to calculate an estimator of the baseline cumulative hazard function by Breslow estimation of the baseline hazard function. The individual failure likelihood calculation module may be configured to calculate the baseline survival curve estimator using an augmented Breslow estimation. This makes it possible to predict the likelihood of failure in the far future, and can be applied to the formulation of long-term maintenance plans.
[0011] A method for predicting the occurrence of failures in an underground piping network by computer, according to one aspect of the present disclosure, includes the steps of: reading connection data, the connection data including data indicating one or more pipe segments and data indicating one or more valves, wherein the end of each pipe segment is connected to the end of another pipe segment or a valve; automatically generating data indicating one or more clusters from the connection data, wherein each cluster includes pipe segments that form a connection network between valves; and calculating the failure likelihood of each cluster by summing the failure likelihoods of the pipe segments included in each cluster. The method includes the step of displaying the one or more clusters and the failure likelihood of each cluster on a display.
[0012] The process may further include a step of loading population forecast data for each region, and the display may show the population forecast data in addition to data indicating the one or more clusters and the failure likelihood of each cluster.
[0013] The failure likelihood of the pipe segment may be calculated based on machine learning of the correlations between previously collected data. The step of calculating the failure likelihood may include calculating an estimator of the baseline cumulative hazard function by Breslow estimation of the baseline hazard function. The step of calculating the failure likelihood may include calculating an estimator of the baseline survival curve using an augmented Breslow estimate.
[0014] An apparatus according to one aspect of the present disclosure includes one or more processors and one or more memories communicating with the one or more processors, the one or more memories containing computer-executable instructions, which, when executed by the one or more processors, cause the one or more processors to perform a method for predicting the occurrence of a failure in an underground conduit network.
[0015] A computer program according to one aspect of the present disclosure causes a computer to perform a method for predicting the occurrence of failures in an underground conduit network when it is executed on a computer.
[0016] In this disclosure, a pipe segment failure refers to a condition in which the distribution of water through a pipe segment becomes insufficient or impossible due to damage to that pipe segment, such as a break. For example, if the pipe segment is a water pipe, a leak (water leakage) caused by damage would be considered a failure. In this disclosure, failure, damage, and even leaks / water leakage are often used interchangeably. [Brief explanation of the drawing]
[0017] [Figure 1] This is a schematic diagram showing the configuration of a pipe failure prediction device according to one embodiment of the present disclosure. [Figure 2] This is a schematic diagram of a pipe failure prediction system that may include a pipe failure prediction device according to one embodiment of the present disclosure. [Figure 3] An example of an underground piping network including pipe segments and gate valves is shown. [Figure 4] This is a simplified version of Figure 3. [Figure 5] This diagram illustrates the intersection of two pipe segments. [Figure 6] This is an example of displaying cluster-level LOFs on a map. [Figure 7]An example of a display obtained by superimposing population prediction on the LOF of underground piping. [Figure 8] An example of a display obtained by superimposing population prediction on the LOF of underground piping. [Figure 9] An example of a display obtained by superimposing population prediction on the LOF of underground piping. [Figure 10] An example of a display obtained by superimposing population prediction on the LOF of underground piping. [Figure 11] An example showing the position of a pipe segment, the predicted population density, and the predicted population change rate. [Figure 12] An example of a GUI for filtering the display of FIG. 11 is shown. [Figure 13] A display showing the filtering input by the GUI of FIG. 12 for the display of FIG. 11 is shown. [Figure 14] An example showing the Breslow estimate of the survival curve. [Figure 15] An example showing the linear regression of FIG. 14 is shown. [Figure 16] An example of the average failure probability obtained by applying the long-term LOF prediction procedure for each material of the pipe segment. [Figure 17] An example of the average failure probability obtained by applying the long-term LOF prediction procedure for each type of joint. [Figure 18] An example of an area-specific risk heat map obtained by the long-term LOF prediction procedure. [Figure 19] An example of an area-specific risk heat map obtained by the long-term LOF prediction procedure. [Figure 20] An example of an area-specific risk heat map obtained by the long-term LOF prediction procedure. [Figure 21] A flowchart showing an example of a computer-based piping failure prediction method according to an embodiment of the present disclosure. [Figure 22] A diagram showing an example of the hardware configuration of a device that executes the piping failure prediction method according to the present disclosure.
Mode for Carrying Out the Invention
[0018] The embodiments of the present invention will be described in detail below.
[0019] Figure 1 is a schematic diagram showing the configuration of a pipe failure prediction device 100 according to one embodiment of this disclosure. The pipe failure prediction device is a device that predicts the occurrence of failures in an underground piping network. The pipe failure prediction device 100 in Figure 1 includes a reading module 110, a clustering module 120, a cluster failure likelihood calculation module 130, and a display module 140. The pipe failure prediction device 100 may include an individual failure likelihood calculation module 135 in addition to, or instead of, the cluster failure likelihood calculation module 130. These modules 110-140 may be contained within the same enclosure, or some or all of them may be in separate enclosures. They may also be hardware-based, or their functions may be implemented in software. In particular, the functions of each module may be implemented within a cloud server system. Furthermore, depending on the application to which the pipe failure prediction device 100 is used, modules or components other than modules 110 to 140 may be added as appropriate.
[0020] The pipe failure prediction device 100 in Figure 1 can constitute part of the pipe failure prediction system in Figure 2. The pipe failure prediction system in Figure 2 includes a front-end interface, a management system, and a machine learning system. The front-end interface includes pages for uploading pipe segment data, damage data, and any auxiliary data; pages for viewing the results of machine learning analysis in both map view and auxiliary statistics; pages for downloading cleaned data and machine learning result maps, as well as statistics; and an interface that allows small businesses to access the company's pipe failure prediction solution even without the correct data or software. The management system includes a management server for creating instances and processes, a database containing customer information, and a file server for storing files. The machine learning system (instance) includes scripts for spatial joining, geoprocessing, and machine learning, and a temporary database for storing files.
[0021] In the pipe failure prediction system shown in Figure 2, in step 11, the customer logs in and uploads data. In step 12, the customer's data is uploaded, and in step 13, a request is made to the process manager of the management server. In step 14, the operator (i.e., the operator of the company predicting pipe failures) logs in and issues a request to the instance manager of the management server. In step 15, the instance manager issues a request to the machine learning instance of the machine learning system. In step 16, raw files from the file server are loaded into the data process of the machine learning instance. In step 17, the data process inserts pipe segment and damage data into the Geographic Information System (GIS) database. The geoprocess receives information from the GIS and national databases. In step 18, the geoprocessed information is supplied to the predictor. In step 19, the predicted results are transferred to the Likelihood of Failure (LOF) results on the file server. In step 20, the LOF results are uploaded to the front-end viewer. In step 21, the customer logs in and views / downloads the LOF data.
[0022] Next, the Valve-Isolated Segment (VIS) evaluation function will be described. This function clusters underground pipe segments that form an underground piping network and predicts the failure risk for each cluster. In step 11 of Figure 2, the VIS evaluation function provides the VIS LOF when the LOF results of the pipe segments are uploaded to the front-end viewer, and also provides a function to display the VIS LOF in the front-end viewer. The piping failure prediction device 100 in Figure 1 includes the following procedure.
[0023] The reading module 110 is configured to read connection data. The connection data includes data indicating one or more pipe segments (intermediate piping). The connection data also includes data indicating one or more valves. The end of each pipe segment is connected to the end of another pipe segment or a valve. This allows for the identification of the location of gate valves in the underground piping network. Connection data can be read from data held by the operating company (e.g., water utilities, water companies, etc.). Alternatively, it can be identified by other means, such as reading data from other databases or reading user input. Location information indicating the location of the identified gate valve may be displayed on the map screen of the piping failure prediction system, for example, as a plot.
[0024] Each valve is associated with the nearest pipe segment, establishing a direct link. If two or more pipe segments share a valve, they are considered part of the same cluster. However, they are only directly connected if the network specifies such an association. In other words, the connection of two or more pipe segments by a valve must be specified in the piping network data.
[0025] Figure 3 shows an example of an underground piping network including pipe segments 1-8 and gate valves A-E. The numbers 1-8 assigned to each pipe segment are pipe segment identifiers (IDs). The symbols A-E for the gate valves are gate valve IDs. Figure 4 is a simplified version of Figure 3. For example, the two ends of pipe segment 4 are connected to pipe segment 6 and valve A, respectively.
[0026] The clustering module 120 is configured to automatically generate data indicating one or more clusters from the connection data. Each cluster includes a pipe segment that forms a connection network between valves. In this way, the sections (clusters) separated by gate valves are defined as a single VIS.
[0027] Referring to Figure 3 or Figure 4, it can be seen that there is VIS1 consisting of pipe segments 1 to 3 and VIS2 consisting of pipe segments 4 to 8. VIS1 and VIS2 are separated from each other by gate valve A. Therefore, even if a leak occurs in pipe segment 2, for example, closing gate valves A and B will prevent the leak from affecting VIS2.
[0028] Based on the pipe segments 1-8, gate valves A-E, and connection data indicating their connections included in Figure 3, one or more VISs included in Figure 3 can be identified. For any underground piping network, it is possible to automatically obtain multiple VISs by clustering the underground piping and applying an algorithm known as Depth-First Search to a graph where gate valves are vertices and one or more pipe segments between gate valves are edges. The method used in the clustering module 120 to automatically generate one or more clusters from the connection data may be any other algorithm and is not particularly limited.
[0029] Furthermore, the pipe failure prediction device 100 of this disclosure can perform clustering even for three-dimensional underground pipe networks, which are not planar, by using connection data. For example, Figure 5 shows the intersection of pipe segment 1 and pipe segment 2 projected onto a plane. If the soil or other coverings are removed and pipe segment 1 and pipe segment 2 are exposed, and photographed from above, a photograph like Figure 5(a) would be obtained. In this case, even if pipe segment 1 and pipe segment 2 intersect in the projection plane, they may actually be connected (i.e., the medium in pipe segment 1 can flow through the intersection to pipe segment 2, Figure 5(b)) or they may not actually be connected (there is no intersection between pipe segment 1 and pipe segment 2, and the medium in pipe segment 1 cannot flow to pipe segment 2, Figure 5(c)). Thus, when only the projection plane is used, that is, when the underground piping network is simplified and represented in a planar manner, clustering may not be accurate. In contrast, by using connection data, it becomes possible to perform accurate clustering even for underground piping networks that are not planar.
[0030] The cluster failure likelihood calculation module 130 is configured to calculate the failure likelihood of each cluster by summing the failure likelihoods of the pipe segments included in each cluster. This allows for the calculation of the Likelihood of Failure (LOF) for each VIS. More specifically, for one or more pipe segments contained within a single VIS, the sum of the LOFs of each pipe segment is defined as the LOF of the VIS. The LOF (Loss of Failure) for each pipe segment can be calculated using the individual failure likelihood calculation module 135. The functions of the individual failure likelihood calculation module 135 will be described later.
[0031] The identified VIS may be stored and managed in a database along with the calculated LOF value. Furthermore, failures related to the pipe segments included in the VIS may also be linked to that VIS, ensuring comprehensive data management and analysis capabilities.
[0032] The display module 140 is configured to display data indicating one or more clusters and the failure likelihood of each cluster on the screen of the display device. The display device may be located in the same place as the pipe failure prediction device 100, or in a different location, particularly a remote location.
[0033] On the screen, the VIS is converted into map tiles. The VIS map tiles may be similar to existing visualizations of pipes and faults. This allows users to visually explore VIS clusters on a geographical map, helping them to quickly identify high-risk areas. The map visualizations have been enhanced to include VIS identifiers, making it easier to understand the distinctions between VISs and the specific risks associated with each VIS cluster.
[0034] Data indicating one or more clusters and the failure likelihood of each cluster may be displayed on a map screen shown on a display to indicate the geographical location of the underground conduit network. For example, the LOF for each VIS can be displayed in the map screen of the pipe failure prediction system using the color assigned to the VIS. Alternatively, the map screen of the pipe failure prediction system may be provided with a graphical user interface (GUI) component that switches between showing the LOF for each pipe segment and the LOF for each VIS (see the toggle switch for "Construction Order Unit (VIS)" in Figure 11 as an example).
[0035] Furthermore, APIs may be provided to access VIS-related data based on identification values such as properties and VIS cluster LOF values. Such APIs enable integration with other systems and support detailed data export and analysis. Users can export VIS-related data and VIS LOF values in file formats suitable for offline analysis and reporting.
[0036] The user interface (UI) includes an option to switch between VIS layers on the map, providing a user-friendly mechanism for analyzing clustered data. For example, clicking on a VIS cluster could display detailed information in a sidebar.
[0037] Previous pipeline deterioration diagnostics displayed evaluations for each pipe segment with an ID, which sometimes resulted in overly detailed displays that didn't match the actual construction sections, making them difficult for users to use in planning construction projects. Actual construction sections are sometimes carried out in units of VIS (Vessel Identification Sections) separated by valves. As shown in Figure 1, the pipe failure prediction device 100 solves this problem by displaying LOF (Level of Failure) for each VIS.
[0038] Figure 6 shows an example of displaying LOFs at the cluster level on a map. In Figure 6, the percentile ranking of LOFs for each VIS is displayed from smallest to largest value. Each VIS is displayed with a color or shade corresponding to its percentile ranking. By referring to Figure 6, it becomes possible to prioritize repair plans based on the construction unit, starting with those with the highest failure likelihood (leakage risk).
[0039] Thus, the pipe failure prediction device 100 according to this disclosure automatically clusters all pipe segments that form a network between shut-off valves, creates a comprehensive VIS, and improves network visibility. By assigning LOFs for each cluster to each VIS, it becomes possible to better understand the risk distribution across the entire piping network. Overall, the pipe failure prediction device 100 described herein can be used to develop repair plans that balance risk assessment with enhanced asset allocation. In other words, it provides a clearer understanding of the pipe network at the road and neighborhood levels, rather than focusing solely on individual pipe segments. This broader perspective helps optimize resource allocation and planning for maintenance or replacement activities.
[0040] Next, we will describe another example of this disclosure relating to the integration of future population change forecasts and information on leakage risks in underground pipelines (especially water pipelines). Traditional assessments of the importance of underground piping were based on the location of facilities such as hospitals and schools, where the water supply must not be interrupted in emergencies. In recent years, in addition to the original functions of water supply facilities (water intake facilities, water storage facilities, water transport facilities, water purification facilities, water transmission facilities, main distribution pipes, water reservoirs, etc.), facilities that have a high potential to cause serious secondary disasters are also classified as important water supply facilities. However, with the increasing risk of leaks in underground pipes, replacing all of them while also addressing aging pipes would require enormous costs and time. Therefore, there is a need to streamline these processes and further optimize priorities.
[0041] Therefore, the pipe failure prediction device 100 according to this disclosure can have a function to display the degree of impact, indicating how many residents will be affected when the supply is interrupted. If we can visualize how many residents will be affected by water outages during an emergency, we can propose the option of prioritizing the replacement of pipelines that will affect the largest number of residents.
[0042] There is a need to provide governments and businesses with crucial insights into future demographic trends across various geographical regions. By overlaying projected demographic changes onto the LOFs of individual pipe segments and / or VISs on a map, it helps businesses optimize capital allocation and ensure that investments in pipe repair or installation are financially justified by future demand.
[0043] In another example of this disclosure, the loading module 110 in Figure 1 may be further configured to load current population data for each region, or the loading module 110 in Figure 1 may be further configured to load projected population data for each region at some point in the future. Furthermore, "region" may refer to administrative or business units such as municipalities, blocks, or districts, as well as grids on a map divided at appropriate intervals (e.g., 1 kilometer), or the entire display area. "Population data" or "population forecast data" includes, but is not limited to, one or more of the following: population itself, population density, or population growth rate. Furthermore, for simplicity, the term "population forecast data" can refer to both "current population data" and "population forecast data for a future period."
[0044] Population projection data is provided, for example, as a shapefile with polygon geometry by the Ministry of Land, Infrastructure, Transport and Tourism (MLIT) of Japan. This data is categorized by region and includes age- and gender-specific population estimates for each year: 2020, 2025, 2030, 2035, 2040, 2045, and 2050. Population projection data may be extracted from data held by these national and local governments, with age and gender information removed for simplification purposes, leaving only data on absolute numbers. This simplified data includes identifiers, prefectures, area codes, shapes, and population (for 2015, 2020, 2025, 2030, 2035, 2040, 2045, and 2050).
[0045] When each pipe segment (represented as multiple polylines or polyline geometry on the map) is displayed on the map, a buffer with a user-defined radius is generated around each pipe segment. This buffer represents the area that would be potentially affected if the pipe fails. For each pipe buffer, intersecting population polygons (polygons indicating regions that provide population data for population forecasting) are considered to be affected (threatened) by potential pipe failures. The total population within these polygons represents the number of customers affected during a service disruption. This logic is also applied to VIS clusters, aggregating the affected population across all pipe segments within the VIS.
[0046] In particular, the system may be configured to automatically link the underground piping data of each business company with the population density data read by the reading module 110. This function provides the ability to link population density data to the LOF of a pipe segment or VIS in the front-end viewer when the LOF results of the pipe segment or VIS are uploaded to the front-end viewer in step 11 of Figure 2.
[0047] The display module 140 can be further configured to display population projection data in addition to data indicating one or more clusters and the failure likelihood for each cluster. The LOF (Level of Field) for underground piping is displayed for each pipe segment or for each VIS (Vision System). The display module 140 allows population forecast data to be displayed on the map for each area, for example, by the color or shading of a semi-transparent mesh.
[0048] In particular, it allows us to visualize the future population supply. To do this, we overlay a mesh map based on future population projections onto the LOF (Level of Flow) of underground piping. The display module 140 may display translucent polygons representing population data, with their colors dynamically changing based on predicted population changes. This visual representation allows users to easily identify areas where the population is increasing or decreasing. Alternatively, or in addition to this, a dropdown menu may allow the user to select the year for which to view population projections, starting from the first year after the current year (e.g., 2025, 2030, 2035). This ensures that the data is appropriate and consistent with future planning periods. Population polygons are overlaid on existing LOF and / or VIS maps, allowing users to simultaneously visualize both failure risks and future population trends. This integration helps identify areas where investment in pipeline infrastructure may be more or less justified based on future population projections. The UI may dynamically update to display an estimated number of affected customers (due to water outages) based on the selected year and projected population data. This feature provides specific measures of the impact of pipeline failures or maintenance activities.
[0049] Referring to Figures 7, 8, 9, and 10, which can be displayed by the pipe failure prediction device 100 related to this disclosure, a display that overlays population forecasts onto the LOF of underground piping will be explained.
[0050] As an example, Figure 7 shows the LOF (Level of Focus) of underground piping overlaid with the projected population for 2030. The LOF of underground piping is displayed with a color or shade corresponding to the percentile ranking of the LOF, similar to Figure 6. The projected population for 2030 is displayed for each grid-like region on the map, with a color or shade corresponding to the percentile ranking of the projected population within that region. Note that the percentile ranking of the projected population in Figure 7 is based on areas divided into an equally spaced grid on the map, and is therefore equivalent to the percentile ranking based on projected population density. Similarly, Figures 8, 9, and 10 show the LOF of underground piping overlaid with projected population figures for 2035, 2040, and 2050, respectively.
[0051] In this way, it is possible to visualize on a map how the current population will change in the future and predict the LOF (Level of Flow) of underground piping. This makes it possible to assess importance while considering population projections. In particular, it becomes possible to formulate repair plans that take future revenue into account.
[0052] Furthermore, in addition to, or instead of, the projected future population, the rate of population change may be displayed. In particular, the rate of population change may be displayed using colors or shades applied to pipe segments or VIS. Referring to Figures 11, 12, and 13, which can be displayed by the pipe failure prediction device 100 related to this disclosure, a display that overlays population forecasts and population change rates onto the LOF of underground piping will be explained.
[0053] Figure 11 shows the locations of the pipe segments, with the projected population density for each area in 2030 indicated by the color or shade of the semi-transparent mesh, and the projected population change rate for each pipe segment in 2050 indicated by the color or shade of the pipe segment. Figure 12 shows an example of a GUI for filtering the display in Figure 11. In Figure 12, only pipe segments with a leakage probability (failure probability) in the range of 0-12% are displayed, and the filter is set so that the number of people in 2050 increases in the range of 676-1588, and the rate of change in 2050 is in the range of -25-19%. Figure 13 shows the display of Figure 11 with the filtering input from the GUI in Figure 12. In this way, it becomes possible to overlay future population distribution and rate of change with the occurrence of failures, which helps in the effective formulation of repair plans.
[0054] In another example of this disclosure, the pipe failure prediction device 100 may include an individual failure likelihood calculation module 135 in addition to, or instead of, the cluster failure likelihood calculation module 130. The individual failure likelihood calculation module 135 is configured to calculate the failure likelihood of a pipe segment based on machine learning of correlations between previously collected data. The cluster failure likelihood calculation module 130 can calculate the failure likelihood of each cluster by summing the failure likelihoods of each pipe segment calculated by the individual failure likelihood calculation module 135.
[0055] The collected data includes pipeline data and failure history. The collected data may also include environmental big data, which consists of various environmental information surrounding the pipelines. Pipeline data includes information such as pipe segment (water pipe) details (pipe diameter, length, material, construction year, etc.). Failure history includes, for example, leak history. Pipeline data and failure history are digitized, modified, and / or supplemented as needed. Environmental big data includes data on population, soil, rivers, transportation networks, earthquakes, etc., and is based on a database of a vast number of variables constructed across Japan, for example.
[0056] This section describes the long-term failure likelihood (LOF) prediction using the individual failure likelihood calculation module 135. Here, "long-term" generally refers to periods exceeding 5 years, typically 20, 30, 50, and 100 years into the future, but there are no upper or lower limits. Conventional methods are limited by the range of years observed in the training data and therefore cannot predict such long-term failure likelihoods. The long-term LOF prediction function of this disclosure is realized in the machine learning instance shown in Figure 2 by employing the method described below. This long-term forecasting capability enables businesses to better plan and allocate resources over the long term, ensuring that infrastructure investments align with future risks. In particular, businesses can proactively manage pipeline infrastructure based on long-term risk forecasts, optimized maintenance schedules, and replacement strategies.
[0057] First, once the business company's data is uploaded, it is cleaned and normalized to ensure compatibility with machine learning models. Based on environmental data such as soil characteristics, precipitation, population density, and transport characteristics, more than 100 additional variables (features) can be generated. These variables provide a comprehensive dataset for predicting future pipe failures.
[0058] The long-term LOF model used by the pipe failure prediction device 100 relating to this disclosure uses historical failure data (if available) to establish correlations between pipe attributes, environmental variables, and failure rates. This model applies advanced machine learning algorithms, described later, to predict the probability of pipe failure over long-term perspectives (20, 30, 50, and 100 years). This model explains time-dependent variables and estimates future risks based on current and historical data patterns.
[0059] The method described below allows us to predict the failure patterns of pipe segments of any length that will fail in any given future year. In other words, it models a random variable T that represents the number of years the pipe will survive in the future. Using this model, we can simulate pipe failures and estimate many useful statistics, such as failure probabilities and the expected number of failures. For this purpose, assuming the pipe did not fail until t-1 ("survived"), the probability of failure at year t is calculated.
number
[0060] Here we describe the Cox model. The Cox model is a methodology primarily used in the field of survival analysis to estimate hazard functions under the assumption of proportional hazards. Traditional applications of the Cox model rely on linear parameterization models with respect to covariates, but in this disclosure instead applies a learning-based approach via XGBoost's GBDT-based library, which has native support through a loss function tailored to the Cox proportional hazards model. In the model of this disclosure, for the hazard function,
number
number
[0061] In the XGBoost implementation of the Cox model, f(X) is trained using the GBDT algorithm with a well-defined loss function. The raw model output is Τ(X)=e f(X) To provide this, the output of the fitted GBM model can be directly plugged in for this part. However, since h0(t) must be estimated separately, we focus on how to use Breslow estimation and extend it to the desired approach.
[0062] Many treatments of Cox models avoid fully estimating h0(t). In fact, one of the great advantages of the proportional hazards assumption is that the scaling factor h0(t) is completely ignored and treated as constant with respect to X, and e f(X) The advantage is that relative risks can be compared simply by estimating e f(X) This alone doesn't tell us anything about the underlying absolute risk. Since the absolute risk is ultimately necessary, we also need to estimate h0(t).
[0063] The following approach uses Breslow estimation. Breslow estimation is a fitted model e f(X) Under the hazard ratio calculated by,
number
number
number
Number
Number
[0064] The individual failure likelihood calculation module 135 can be configured to calculate the Breslow estimator of the baseline hazard function using equations (4) and (5).
[0065] The original goal is to estimate the value of equation (1) rather than the hazard function. The hazard function does not teach anything about the underlying probability but teaches the actual failure probability for each year. In survival analysis, this is usually obtained from the survival curve S(t) representing the inverse cumulative distribution function of the random variable T. The cumulative hazard and S(t) are related by the following definitions.
Number
Number
Number
Number
Number
[0066] Although a method for calculating the survival curve has been described, it is necessary to consider the fact that the data does not actually fully conform to the assumptions of survival analysis. Importantly, in the Cox model, it is necessary to permit multiple events that are strictly permitted only one per subject.
[0067] The goal is to adequately predict the future (typically 100 years). To achieve this, it is necessary to change the rules of the models constructed so far. Since the Breslow estimator is a nonparametric estimator, it is necessarily restricted to the range of years observed in the training data, and the prediction is limited to the range of about t < 30 [years]. Generally, the time t satisfies t0 < t < t1, where t0 and t1 are constants representing the upper and lower limits, respectively. Typically, t0 = 0 and t1 = 30 (the unit is years).
[0068] To predict the future in the range of time t where t1 < t < t2 (t2 is a constant, typically t2 = 100 (the unit is years)), the following method is adopted. FIG. 14 plots the Breslow estimator of the baseline survival curve exp(-H0(t)) according to the Breslow estimators of the cumulative hazard function (Equations (4) and (5)) using the weights learned for 0 < t < 30. It is observed that the Breslow estimator of the baseline survival curve shown in FIG. 14 is approximately linear in the range of 0 < t < 30. FIG. 15 shows the regression line obtained by the least squares method in the range of 0 < t < 30.
[0069] Based on such findings, in the range of t0 < t < t1 (especially t0 = 0, t1 = 30 [years])
Number
[0070] The individual failure likelihood calculation module 135 can be configured to compute the extended Breslow estimator of the survival function by equation (9).
[0071] Estimated survival curve
number
number
number
number
[0072] By simply generating random numbers, we can simulate the results across all pipes each year and treat these as simulated failures. Since past failure counts are features included in feature set X, each time a simulated "failure" occurs, we increment the break count of X to reflect this updated state and rerun the inference on this newly updated data. Naturally, this yields a fairly high-variance output, so we repeat this simulation B=100 times over all 100 years to obtain the average result and thereby bootstrap the estimate to reduce the variability of the result. Since it is possible to directly access X, it is also possible to simulate pipe replacement by "updating" the selected pipe attribute value of X with the selected value and resetting the service life of the pipe together with the material type, diameter, etc.
[0073] FIG. 16 shows the average failure probability (see equation (1)) at t years (t ranges from 0 to 100) obtained by applying the long-term LOF prediction procedure for each material (AC, CAS, CON, DIP, OTH, PCRC, PEHDPE, PLASTIC, PVC, SP) of the pipe segment. That is, the covariate (feature) X indicates the material of the pipe segment, and the average failure probability is calculated from the learning in the range of 0 < t < 30.
[0074] FIG. 17 shows the average failure probability (see equation (1)) at t years (t ranges from 0 to 100) obtained by applying the long-term LOF prediction procedure for each joint (A, EF, GX, K, NS, RR, S2, T, TD, TLD, TS) of the cast iron pipe (DIP). That is, the covariate (feature) X indicates the joint of the pipe segment, and the average failure probability is calculated from the learning in the range of 0 < t < 30.
[0075] By these methods, the pipe failure prediction device 100 can predict the likelihood of failure in a much more distant future (typically 100 years ahead) than before. FIGS. 18, 19, and 20 are area-specific risk heat maps showing failure risks displayed by applying such a long-term LOF prediction procedure in the pipe failure prediction device 100 according to the present disclosure.
[0076] FIG. 18 is an area-specific risk heat map as of 2022. A total of 48 failure risks are predicted. FIG. 19 is an area-specific risk heat map as of 2032. A total of 133 failure risks are predicted. FIG. 20 is an area-specific risk heat map as of 2122 (100 years ahead). A total of 663 failure risks are predicted. According to the pipe failure prediction device 100, such long-term failure risk prediction and its display are possible.
[0077] The display module 140 may also display area-specific risk heatmaps overlaid on a map. As shown in Figures 11-13, long-term LOF forecasts may be classified into risk classes and ranked according to predefined thresholds and percentiles. Ranking may include, for example, the following risk classifications: <Highest Risk> Pipe segments with the top 1% of LOF. <Risk Level> Absolute risk classification based on predefined LOF buckets (for example, thresholds of 50%, 25%, 10%, and 5% corresponding to LOF values of 0.50, 0.25, 0.10, and 0.05). <Risk Rank> Relative risk classification based on percentile buckets (e.g., top 1%, top 3%, top 5%, bottom 95 percentile).
[0078] The UI functionality may be expanded as follows depending on the long-term LOF forecast. In the map control, the user selects a forecast period (20, 30, 50, or 100 years) from a dropdown menu on the screen. Such a menu allows for easy switching between different timeframes, helping users assess risks over various time periods. In the data display options, data can be displayed according to the risk classification selected by the user (highest risk, risk level, risk rank, etc.). These options allow users to filter the displayed data based on their specific needs, focusing on the highest risk segment or a broader risk classification. Add filters to narrow down the data displayed on the map. Users can filter the display by installation year, material, diameter, and other pipe attributes. This flexibility allows users to drill down into specific areas of interest and analyze segments based on multiple criteria.
[0079] The displayed map can be an interactive map. Zooming in and out of the map is possible using the mouse scroll wheel or the + / - controls in the lower left corner of the screen. When zoomed in, users can click and select individual pipe segments to view their risk rank (relative position within the risk distribution), historical failure count (historical data on pipe failures), top risk factors (major factors contributing to pipe failure risk based on a machine learning model), failure probability (probability of failure over a selected forecast period such as 20, 30, 50, or 100 years), and other detailed attributes (a complete profile of the pipe segment, including pipe ID, installation year, material, length, diameter, etc.).
[0080] The pipe failure prediction device 100 may also provide a user-accessible risk ranking table that ranks pipe segments based on LOF. The risk ranking table includes information for the top 1% of pipe segments with the highest LOF values, such as rank, pipe ID, location, installation year, diameter, length, material, number of past failures, LOF for the past year, and top risk factors. The pipe failure prediction device 100 may have a function that allows the user to export a report summarizing long-term LOF predictions, segmented by risk class and other attributes. This function supports detailed analysis and facilitates communication with stakeholders.
[0081] According to the pipe failure prediction device 100 with long-term LOF prediction capabilities, extending the prediction period allows businesses to anticipate future risks, develop long-term maintenance and replacement strategies, and reduce the likelihood of unexpected failures. Long-term LOF prediction provides a basis for more strategic capital allocation, ensuring that investments are targeted to high-risk areas over the long term. Access to detailed long-term risk predictions allows business managers to make more informed decisions about pipeline management and balance short-term needs with long-term goals.
[0082] Furthermore, even for municipalities and businesses with few failure (water leakage) records, it becomes possible to diagnose the likelihood of failure using a model (trained model) that has learned the leakage trends and / or patterns of municipalities other than the target municipality.
[0083] Figure 21 is a flowchart showing an example of a computer-based pipe failure prediction method according to one embodiment of this disclosure. The pipe failure prediction method 1000 will be described with reference to Figure 21. The piping failure prediction method 1000 includes the steps of reading connection data (S1100), automatically generating one or more clusters from the connection data (S1200), calculating the failure likelihood of each cluster (S1300), and displaying data indicating one or more clusters and the failure likelihood of each cluster on a display (S1400).
[0084] In the step of reading connection data (S1100), the connection data includes data indicating one or more pipe segments and data indicating one or more valves, where the end of each pipe segment is connected to the end of another pipe segment or a valve, indicating an underground piping network.
[0085] In the step (S1200) of automatically generating one or more clusters from connection data, each cluster is a VIS that includes pipe segments forming a connection network between valves.
[0086] In the step of calculating the failure likelihood of each cluster (S1300), the failure likelihood of the cluster is calculated by summing the failure likelihoods of the pipe segments included in each cluster.
[0087] In step (S1400), which involves displaying one or more clusters and the failure likelihood of each cluster on a display, one or more clusters and the failure likelihood of each cluster may be displayed on a map.
[0088] According to the piping failure prediction method 1000, the likelihood of failure for each cluster separated by valves is displayed. Overall, this allows for an understanding of the impact of failures on the VIS, which is the unit isolated during repairs, and enables the development of repair plans at the cluster level that can be shut off by valves.
[0089] Depending on the application to which the piping failure prediction method 1000 is applied, steps other than S1100 to S1300 may be added as appropriate. Furthermore, the execution order of each step may be changed or repeated as appropriate.
[0090] The pipe failure prediction method 1000 further includes the step of reading population prediction data for each region, and the population prediction data may be displayed on the display in addition to the one or more clusters and the failure likelihood of each cluster. The pipe failure prediction method 1000 may calculate the failure likelihood of the pipe segment based on machine learning of the correlation between previously collected data.
[0091] A method for predicting the occurrence of failures in an underground piping network using a computer, comprising the steps of: reading regional population forecast data; displaying data indicating one or more pipe segments on a display; and displaying the population forecast data and the failure likelihood of the one or more pipe segments, is also within the scope of this disclosure. In other words, when displaying population forecast data, it is acceptable to display the LOF for each pipe segment rather than the LOF for each VIS unit.
[0092] A method for predicting the occurrence of failures in an underground piping network using a computer, comprising the steps of: calculating the likelihood of failure of one or more pipe segments based on machine learning of correlations between previously collected data; and displaying data indicating the one or more pipe segments and the likelihood of failure of the one or more pipe segments on a display, is also within the scope of this disclosure. In other words, based on machine learning of the correlations between previously collected data, future LOFs may be displayed for each pipe segment rather than for the VIS unit.
[0093] The step of calculating the failure likelihood for each pipe segment may include calculating an estimator of the baseline cumulative hazard function by Breslow's estimation of the baseline hazard function. The step of calculating the failure likelihood for each pipe segment may include calculating an estimate of the baseline survival curve using an augmented Breslow estimate. This makes it possible to predict the likelihood of failure in the far future, and can be applied to the formulation of long-term maintenance plans.
[0094] Furthermore, the pipe failure prediction method of this disclosure may include each step performed by the pipe failure prediction device 100.
[0095] Figure 22 shows an example of the hardware configuration of an apparatus for performing the pipe failure prediction method according to this disclosure. The apparatus 200 in Figure 22 includes one or more processors 210 and one or more memories 220 that communicate with the one or more processors. The one or more memories contain computer-executable instructions, and when those instructions are executed by the one or more processors, the instructions cause the one or more processors to perform the pipe failure prediction method according to this disclosure. One or more processors 210 and one or more memory units 220 may be housed in the same chassis, or some or all of them may be in separate chassis. In particular, they may be distributed across a cloud server system. Furthermore, depending on the application, the device 200 may be appropriately equipped with additional components such as communication devices, input / output devices, display devices, and storage devices.
[0096] Furthermore, a computer program that, when executed by a computer, causes the computer to execute the pipe failure prediction method relating to this disclosure is also within the scope of this disclosure. A temporary or non-temporary computer-readable storage medium storing such a computer program and signals for transmitting such a computer program are also within the scope of this disclosure.
[0097] As used in this disclosure, the conjunctions “and,” “or,” and “and / or” are intended to indicate that there is one or more that they connect. In particular, the term “or” implies an inclusive, rather than exclusive, choice.
[0098] In the embodiments of this disclosure, unless otherwise specified or logically inconsistent, terms and / or descriptions in different embodiments are consistent and may be referenced to one another, and technical features of different embodiments may be combined to form new embodiments based on the internal logical relationships of the different embodiments.
[0099] In this disclosure, directional terms such as “top” and “bottom,” as well as “left” and “right,” are defined in relation to the direction in which the components are schematically positioned in the accompanying drawings. These directional terms are relative concepts and are used for relative explanation and clarification; please understand that they may change accordingly based on changes in the direction in which the components are positioned in the accompanying drawings.
[0100] It will be understood that the various numbers in the embodiments of this disclosure are used merely for the purpose of distinction to facilitate explanation and not to limit the scope of the embodiments of this disclosure. The sequence numbers of the processes described above do not imply execution order, and the execution order of the processes should be determined based on the functions and internal logic of the processes. Terms such as “first” and “second” are used to distinguish between similar subjects and do not need to be used to describe a specific order or sequence.
[0101] In addition, the components of the attached drawings in the embodiments of this disclosure are intended merely to show the operating principle of the pipe failure prediction device and system, the display content of the screen of the pipe failure prediction device and system, etc., and do not faithfully reflect the actual size relationships of the components, the screen layout, etc.
[0102] The above description represents only a specific implementation of the Disclosure and is not intended to limit the scope of protection of the Disclosure. Any modifications or substitutions that are readily conceivable by a person skilled in the art within the technical scope disclosed herein shall also fall within the scope of protection of the Disclosure. Accordingly, the scope of protection of the Disclosure shall be subject to the scope of protection of the claims. [Explanation of Symbols]
[0103] 11 Login 12 uploads 13 requests 14. Request Issuance 15. Request Issue 16 Road 17 Inserts 18 supply 19 Transfer 20 uploads 21 Login 100 Pipe failure prediction device 110 Loading Module 120 clustering modules 130 Cluster Failure Likelihood Calculation Module 135 Individual Failure Likelihood Calculation Module 140 Display Modules 200 equipment 210 processors 220 memory
Claims
1. A device for predicting the occurrence of failures in underground piping networks, A reading module configured to read connection data, wherein the connection data includes data indicating one or more pipe segments and data indicating one or more valves, and the end of each pipe segment is connected to the end of another pipe segment or a valve, and the reading module A clustering module configured to automatically generate data indicating one or more clusters from the aforementioned connection data, wherein each cluster includes a pipe segment that forms a connection network between valves, and the clustering module comprises: A cluster failure likelihood calculation module is configured to calculate the failure likelihood of each cluster by summing the failure likelihoods of the pipe segments included in each cluster, An apparatus including a display module configured to display data indicating one or more clusters and the failure likelihood of each cluster on a display.
2. The aforementioned loading module is further configured to load population forecast data for each region, The apparatus according to claim 1, wherein the display module is further configured to display population forecast data in addition to data indicating the one or more clusters and the failure likelihood of each cluster.
3. The apparatus according to claim 1, further comprising an individual failure likelihood calculation module configured to calculate the failure likelihood of the pipe segment based on machine learning of correlations between previously collected data.
4. The individual failure likelihood calculation module performs a Breslow estimation of the baseline hazard function for time t (where t0 < t < t1, and t0 and t1 are constants) and feature X. [Math 1] (Here [Math 2] represents the weights obtained by machine learning, δ j R(t) represents the total number of events at time j. j The baseline cumulative hazard function estimator is obtained by (where ) represents the set of individuals still exposed to risk at time j. [Math 3] The apparatus according to claim 3, configured to calculate as follows.
5. The individual failure likelihood calculation module calculates the baseline survival curve exp(-H) for time t (where t1 < t < t2, and t1 and t2 are constants). 0 (t)) is an estimator of the extended Breslow estimation. [Math 4] Calculate using (where t0 is a constant, [Math 5] The apparatus according to claim 3, wherein (where represents weights obtained by the least squares method in the range t0 < t < t1).
6. A method for predicting the occurrence of failures in an underground conduit network using a computer, A step of reading connection data, wherein the connection data includes data indicating one or more pipe segments and data indicating one or more valves, and the end of each pipe segment is connected to the end of another pipe segment or a valve. A step of automatically generating data indicating one or more clusters from the aforementioned connection data, wherein each cluster includes a pipe segment that forms a connection network between valves, The steps include: calculating the failure likelihood of each cluster by summing the failure likelihoods of the pipe segments included in each cluster; A method comprising the step of displaying one or more clusters and the failure likelihood of each cluster on a display.
7. This further includes the step of importing population projection data for each region. The method according to claim 6, wherein the display shows data indicating the one or more clusters, the failure likelihood of each cluster, and the population forecast data.
8. The method according to claim 6, wherein the failure likelihood of the pipe segment is calculated based on machine learning of correlations between previously collected data.
9. A method for predicting the occurrence of failures in an underground conduit network using a computer, Steps to load population forecast data for each region, A method comprising the steps of displaying data indicating one or more pipe segments on a display, population forecast data, and the failure likelihood of the one or more pipe segments.
10. A method for predicting the occurrence of failures in an underground conduit network using a computer, The steps include: calculating the failure likelihood of one or more pipe segments based on machine learning of correlations between previously collected data; A method comprising the steps of displaying data indicating the one or more pipe segments and the failure likelihood of the one or more pipe segments on a display.
11. The step of calculating the failure likelihood involves Breslow estimation of the baseline hazard function for time t (where t0 < t < t1, and t0 and t1 are constants) and feature X. [Math 6] (Here [Number 7] represents the weights obtained by machine learning, δ j R(t) represents the total number of events at time j. j The baseline cumulative hazard function estimator is obtained by (where ) represents the set of individuals still exposed to risk at time j. [Number 8] The method according to claim 10, which includes calculating the following.
12. The step of calculating the failure likelihood involves, for time t (where t1 < t < t2, and t1 and t2 are constants), the baseline survival curve exp(-H) 0 (t)) is an estimator of the extended Breslow estimation. [Number 9] Calculate using (where t0 is a constant, [Number 10] The method according to claim 10, wherein (where represents a weight obtained by the least squares method in the range t0 < t < t1).
13. One or more processors and A device comprising one or more memory that communicates with one or more processors, wherein the one or more memory includes computer-executable instructions, and when executed by the one or more processors, the instructions cause the one or more processors to execute the method according to any one of claims 6 to 12.
14. A computer program that, when executed by a computer, causes the computer to perform the method described in any one of claims 6 to 12.