Predicting pipe failure

JP2025172920A5Pending Publication Date: 2026-02-05FRACTA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025146926
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-10-09
Filing Date
2025-09-04
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Existing water utility companies lack accurate methods to prioritize the replacement of aging pipes, leading to unnecessary replacements and significant financial waste due to inaccurate prediction models.

Method used

A data-driven approach using machine learning and geospatial analysis to predict pipe leaks by selecting a subset of variables, constructing a mathematical model, and predicting likely leaking pipe segments based on historical data and external factors.

Benefits of technology

Improves leak prediction accuracy, reduces excavation costs, identifies risk factors, and optimizes pipe replacement strategies, thereby saving costs and time for water utility companies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a system and method for predicting a pipe failure.SOLUTION: A method includes: in response to login of a customer, uploading data (401, 402); issuing a request to a process manager of a management server (403); in response to log in of an operator of a pipe failure prediction company, issuing a request to an instance manager of the management server (404); issuing a request to a machine learning instance of a machine learning system (405); loading a raw file from a file server to a data process of the machine learning instance (406); inserting, by the data process, pipe and failure data into a GIS database and receiving, by geoprocessing, information from the GIS or country database (407); supplying the geoprocessed information to a predictor (408); transferring a prediction result to likelihood of failure (LOF) result on the file server (409); uploading it to a front end viewer (410); and allowing the log-in customer to view the LOF data (411).SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims priority to and incorporates by reference each of the following provisional applications:

[0002] U.S. Provisional Patent Application No. 62 / 649,058, filed March 28, 2018; U.S. Provisional Patent Application No. 62 / 658,189, filed April 16, 2018; U.S. Provisional Patent Application No. 62 / 671,601, filed May 15, 2018; U.S. Provisional Patent Application No. 62 / 743,477, filed October 9, 2018; U.S. Provisional Patent Application No. 62 / 743,483, filed October 9, 2018; and U.S. Provisional Patent Application No. 62 / 743,485, filed October 9, 2018.

[0003] This application is related to and incorporated by reference in its entirety by the following U.S. and PCT applications, filed on even date herewith and collectively referred to as the "co-pending patent applications": U.S. Patent Application No. 16 / 365,466 (Attorney Docket No. Fracta-002-US), U.S. Patent Application No. 16 / 365,522 (Attorney Docket No. Fracta-006-US), and International Patent Application No. ___ (Attorney Docket No. Fracta-006-PCT).

[0004] This patent application relates generally to improved systems and methods for predicting pipe damage. More particularly, some embodiments relate to improved systems and methods for predicting future breaks in pipes, such as water mains. [Background technology]

[0005] Over one million miles of water pipes in the United States are nearing the end of their useful lives and need to be replaced. Maintaining current levels of service to a growing population over the next 25 years will require an investment of at least $1 trillion. Ignoring this problem will result in even higher repair costs and increased service disruptions.

[0006] In the United States, approximately 50,000 water utilities do not have the resources to replace all of them due to limited budgets. Since it is impossible to replace all outdated pipes, it is important to strategically prioritize the replacement of pipes in the worst condition, with outdated but sound pipes being replaced in the future.

[0007] Replacement plans developed by water utilities are highly inaccurate and often not useful. The simplistic models developed by water utilities result in unnecessary replacement of pipes that still have years of life left. Over the next 25 years, this will result in millions of dollars being wasted. Summary of the Invention [Problem to be solved by the invention]

[0008] According to some embodiments, a method is described for predicting pipe leaks in a network of underground pipes for delivering fluids, such as fresh water or natural gas, to consumers. The method includes receiving, by a computer system, a first set of variables related to pipe leaks, the first set including at least 60 variables; automatically selecting, by the computer system, a second set of variables that is a subset of the first set of variables; constructing a mathematical model using machine learning based on the second set of variables; and predicting likely leaking pipe segments based on the model. According to some embodiments, the pipes can be used to deliver other types of fluids, such as wastewater, recycled water, brackish water, storm water, seawater, drinking water, steam, compressed air, oil, and natural gas. [Means for solving the problem]

[0009] According to some embodiments, the selecting includes building a model based on the first set of variables and assessing an importance associated with each variable in the first set based on the initial model. The importance assessment may be based on coefficients such as Gini coefficients or information gain coefficients. According to some embodiments, at least some of the variables in the first set are assigned to predetermined categories, and no more than a predetermined number of variables in each category are included in the second set.

[0010] According to some embodiments, the number of variables in the first set is at least 100, 300, 800, or 1000. According to some embodiments, the number of variables in the second set is less than 100, 60, 40, 20, or 12. According to some embodiments, the number of variables in the second set is less than 50% of the number of variables in the first set. According to some embodiments, the number of variables in the second set is less than 25% of the number of variables in the first set.

[0011] According to some embodiments, some of the first set of variables are generated by geospatial analysis. According to some embodiments, at least some of the first set of variables include data relating to pipelines that have not experienced leaks, and the selection is based at least in part on the non-leak data.

[0012] According to some embodiments, the network of pipes is for a first customer water utility company, and constructing the model includes using data from a second customer water utility company. According to some embodiments, constructing the model includes using data from a consolidated national water utility database. According to some embodiments, constructing the model is based at least in part on multiple models, each based on data from a different time interval.

[0013] According to some embodiments, a system for predicting pipe leaks in a network of underground pipes supplying fluid to customers is described, including a database storing a first set of variables related to pipe leaks, the first set being at least 60 variables, and a processing system configured to automatically select a second set of variables, the second set being a subset of the first set of variables, build a model using machine learning based on the second set of variables, and predict likely leaking pipe segments based on the model.

[0014] As used herein, all grammatical conjunctions "and," "or," and "and / or" are intended to indicate that one or more of the instances, objects, or entities they connect may occur or be present. Thus, as used herein, the word "or" in all instances denotes an inclusive or rather than an exclusive or.

[0015] To further clarify the above and other advantages and features of the subject matter of this patent specification, specific examples of embodiments thereof are illustrated in the accompanying drawings. It should be understood that elements or parts illustrated in one drawing may be substituted for equal or similar elements or parts illustrated in another drawing, and that the drawings are merely illustrative of exemplary embodiments and should not be considered as limiting the scope of this patent specification and the appended claims. The subject matter will be described and explained with additional specificity and detail using the accompanying drawings. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a schematic diagram illustrating aspects of a processing model used to predict pipe damage, according to some embodiments. [Figure 2] 1 is a diagram illustrating aspects of job planning based on a model used to predict pipe damage, according to some embodiments. [Figure 3A]1 is a diagram illustrating aspects of different water utility customer categories, according to some embodiments. [Figure 3B] 1 is a diagram illustrating aspects of different water utility customer categories, according to some embodiments. [Figure 3C] 1 is a diagram illustrating aspects of different water utility customer categories, according to some embodiments. [Figure 4] 1 is a schematic diagram illustrating aspects of an architecture for predicting pipe damage, according to some embodiments. [Figure 5] Schematic illustrating that the damage per unit length (LOF / length) can be different. [Figure 6] 1 is a diagram illustrating the use of an ensemble over time to calculate the probability of failure for the next five years, according to some embodiments. [Figure 7A] 7A and 7B are schematic diagrams illustrating the performance difference between a conventional technique (FIG. 7A) and a technique according to some embodiments (FIG. 7B). [Figure 7B] 7A and 7B are schematic diagrams illustrating the performance difference between a conventional technique (FIG. 7A) and a technique according to some embodiments (FIG. 7B). [Figure 8A] 1 is a diagram illustrating variables generated via geospatial analysis, according to some embodiments. [Figure 8B] 1 is a diagram illustrating variables generated via geospatial analysis, according to some embodiments. [Figure 9] 1 is a diagram illustrating the construction of an integrated utility database for predicting pipe breaks, according to some embodiments. [Figure 10] 1 is a schematic diagram illustrating aspects of automated variable selection for model-based pipe leak prediction, according to some embodiments. [Figure 11A] 1 is a diagram illustrating aspects of piping replacement optimization, according to some embodiments. [Figure 11B] 1 is a diagram illustrating aspects of piping replacement optimization, according to some embodiments. [Figure 12] An example of a potential plumbing replacement job. [Figure 13A] 1 shows one of three different possible pipe replacement jobs under consideration, according to some examples. [Figure 13B] 1 shows one of three different possible pipe replacement jobs under consideration, according to some examples. [Figure 13C] 1 shows one of three different possible pipe replacement jobs under consideration, according to some examples. [Figure 14A] 1 is a diagram illustrating aspects of predicting missing piping data values, according to some embodiments. [Figure 14B] 1 is a diagram illustrating aspects of predicting missing piping data values, according to some embodiments. [Figure 15A] 1 is a diagram illustrating the correlation between roads and pipes, according to some embodiments. [Figure 15B] 1 is a diagram illustrating the correlation between roads and pipes, according to some embodiments. [Figure 16] FIG. 1 is a block diagram illustrating the creation of virtual piping data, according to some embodiments. [Figure 17] 1 is a diagram illustrating exemplary plumbing data cleaning, according to some embodiments. [Figure 18] 10 is a diagram illustrating further processing of exemplary piping data, according to some embodiments. [Figure 19] 1 is a schematic diagram illustrating feature segmentation and extraction for use in machine learning, according to some embodiments. [Figure 20] 1 is a schematic diagram illustrating the creation of an L OF model through a combination of leakage and generic models, according to some embodiments. [Figure 21] 1 is a diagram illustrating aspects of job planning, according to some embodiments. [Figure 22] 1 is a schematic diagram illustrating a system for predicting pipe breaks, according to some embodiments. [Figure 23] 1 is a diagram illustrating predicting the number of breaks from LOF, according to some embodiments. [Figure 24]Schematic illustrating the predicted similarity in breakages for the past five years and the next five years. [Figure 25] 1 is a diagram illustrating that the predicted number of failures over the next N years can be calculated based on failure history, according to some embodiments. [Figure 26] 1 is a schematic diagram illustrating LOF probability calibration, according to some embodiments. [Figure 27] 1 is a plot showing an exemplary ranking of variable categories, according to some examples. [Figure 28] 1 is a diagram illustrating an example of calculating the importance of representative soil features, according to some embodiments. [Figure 29] Schematic diagram illustrating the normalization of the importance of representative features of the category. [Figure 30] 1 is a diagram illustrating the ranking of a number of representative features from 1 to 10 in importance, according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0017] Several examples of preferred embodiments are described in detail below. While several embodiments are described, it should be understood that the novel subject matter described in this patent specification is not limited to any one embodiment or combination of embodiments described herein, but encompasses numerous alternatives, modifications, and equivalents. Furthermore, while numerous specific details are set forth in the following description to provide a thorough understanding, some embodiments may be practiced without some or all of these details. Moreover, for purposes of explanation, certain technical matters known in the relevant art are not described in detail to avoid unnecessarily obscuring the novel subject matter described herein. It is apparent that one or more individual features of specific embodiments described herein can be used in combination with features of other described embodiments or with other features. Furthermore, like reference numerals and symbols in the various drawings represent like elements.

[0018] According to some embodiments, improved solutions for accurate prediction of pipe condition are described. The methods described herein apply a data-driven approach using a combination of information mining, classification, regression, and / or machine learning. The various systems and methods described herein offer numerous advantages over the prior art. The various systems and methods described herein have been found to result in substantial improvements in leak prediction performance. In particular, some of the embodiments described herein may result in one or more of the following improvements over the prior art: reducing or eliminating the need to excavate pipes to assess their condition, thereby significantly reducing costs; identifying risk of damage based on hundreds of variables, including soil properties, climate, coastal proximity, and railroad lines; identifying correlations that are difficult or impossible for humans to identify; increasing accuracy in predicting future pipe conditions; tolerating increased complexity in the leak prediction problem; and / or reducing the cost and time used in scaling pipe replacement for many water utility companies.

[0019] According to an aspect of the present disclosure, a system for prioritizing replacement of underground pipes includes a database storing information including pipe data, pipe damage data, and external data including geographically specific data, a memory storing at least one program having program instructions, a network interface coupled to at least one computer, and a processor coupled to the database, the network interface, and the memory, capable of executing the program instructions of the at least one program to cause the processor to (a) input and process the pipe data, pipe damage data, and external data from the database to create clean data for a network of pipes, (b) generate latent features in the cleaned data for use in a pipe life damage prediction model, (c) calculate significance of the latent features, (d) extract the most significant features, (e) apply the extracted features to likelihoods, i.e., probabilities, of a damage model created based on historical data and machine learning, (f) predict a future likelihood of failure for each pipe in the network of pipes, and (g) transmit the likelihood of damage for each pipe to a computer associated with a customer of the network of pipes.

[0020] In a further aspect of the present disclosure, the machine learning model is a mixed model having a model part build based on piping and external data without piping damage data and a model part based on piping data, piping damage data, and external data, in which case the prediction is made based on both model parts.

[0021] In another aspect of the present disclosure, one of the at least one programs stored in the memory is a web interface that allows the customer to upload to the database pipe data and pipe damage data for the customer's pipe network.

[0022] In a further aspect, the program further includes a presentation program that allows a user to be presented with a graphical representation of the pipes in the customer's network and the likelihood of each pipe break over a specified future time period.

[0023] In another aspect, the model for the future multi-year period is based in part on a model for pipes in the network over at least one prior multi-year period, and the external data includes data regarding the condition of the pipes, including soil data, weather data, and elevation data.

[0024] According to an aspect of the present disclosure, a method for prioritizing replacement of underground pipes includes inputting and processing pipe data, pipe damage data, and external data from a database to create clean data for a network of pipes, generating latent features in the clean data for use in a pipe life of damage prediction model, calculating the importance of the latent features, extracting the most important features, applying the extracted features to a likelihood of failure model formed based on historical data and machine learning, predicting a future likelihood of failure for each pipe in the network of pipes, and transmitting the likelihood of failure for each pipe to a computer associated with a customer of the network of pipes.

[0025] According to another aspect, the method uses a machine learning model that includes a model part build based on piping and external data without pipe break data, and a model part based on piping data, pipe break data, and external data, and the prediction is based on both model parts.

[0026] According to yet another aspect, the method uses at least one program stored in memory that is a web interface that allows a customer to upload pipe break data for the customer's network of pipes and to a database of pipe data, and further includes a presentation program that allows for presenting a graphical depiction of the pipes in the customer's network and the likelihood of a break for each pipe over a specified future time period.

[0027] According to yet another aspect, the model for the future multi-year period is based in part on models for pipes in the network over at least one previous multi-year period. Further, the external data includes data regarding the condition of the pipes, including soil data, weather data, and elevation data.

[0028] According to an aspect of the present disclosure, a computer program product has computer program logic stored therein that causes a server to determine a likelihood of failure for a network of pipes, the computer program logic including: input and processing logic that causes the server to input and process pipe data, pipe damage data, and external data from a database to create clean data for the network of pipes; generation logic that causes the server to generate latent features within the clean data for use in a pipe life failure prediction model; calculation logic that causes the server to calculate the significance of the latent features; extraction logic that causes the server to extract the most significant features; application logic that causes the server to apply the extracted features to a likelihood of failure model created based on historical data and machine learning; prediction logic that causes the server to predict a future likelihood of failure for each pipe in the network of pipes; and transmission logic that causes the server to transmit the likelihood of failure for each pipe to a computer associated with a customer of the network of pipes.

[0029] According to another aspect, the computer program product includes a machine learning model that is a mixed model having a model part build based on piping and external data without piping damage data and a model part build based on piping data, piping damage data, and external data, and the prediction is based on both model parts.

[0030] According to yet another aspect, the computer program product includes web interface logic that enables the server to enable a customer to upload pipe data and pipe damage data for the customer's pipe network to a database. In addition, the web interface logic further includes a presentation program that enables the server to present to the customer a graphical representation of the pipes in the customer's network and the likelihood of a break for each pipe over a specified future time period.

[0031] According to yet another aspect, the model for a future multi-year period is based in part on a model for pipes in the network over at least one prior multi-year period.

[0032] As used herein, the following terms have the following meanings: "Leaker" is a pipe that has been damaged N times (where N is an integer greater than 0); "Nonleaker" is a pipe that has never been damaged; "Public Data" refers to publicly available and / or available through government sources, such as soil data and weather data; "Utility Data" refers to data provided by water utilities (which can be further divided into "Pipe Data" and "Break Data"); "Pipe Data" refers to geographic pipe data including information on installation year, material, diameter, pressure, etc.; "Break Data" / "Break History" refers to a record of pipe damage including location, associated pipe ID, and date; "Prediction Model for xxxx (e.g., 2017)" refers to a prediction model for xxxx (e.g., 2017) xxxx (e.g., 2017) refers to a predictive model for predicting future pipe damage occurring from the first day of year xxxx in the next N (>0) years, where the model is constructed without using damage data occurring immediately after that date; "Model" refers to a mathematical equation such as f(x); "Modeling" refers to constructing a model such as f(x); "Ensemble" is a machine learning term that predicts an outcome based on the results of multiple models; "Features" correspond to x in y = f(x), where machine learning predicts a target feature from the features; and "Target features" correspond to y in y = f(x), where machine learning predicts a target feature from the features. It should be noted that the use of the term "public" when used in "public data" does not necessarily mean that the data is publicly available for free.Rather, it means that the data is available from a pooled resource such as a government agency (e.g., USGS soil data).

[0033] An important use of machine learning in the water industry is likelihood of failure ("LOF") analysis, also known as condition assessment. Many utilities assume that older pipes are in worse condition than newer pipes. However, older pipes, especially older cast iron pipes, often exhibit remarkable robustness despite being installed 80 to 100 years ago, while newer pipes installed only in more recent decades exhibit significant deterioration and are often close to failure. Therefore, simply replacing pipes of a certain age without taking into account several different variables can be wasteful.

[0034] In connection with predicting future pipe damage, the present disclosure may implement machine learning techniques. Machine learning may be used to build a model to represent a target, i.e., a function "f" of the form y=f(x), where x and y are feature and target features, respectively. In the present disclosure, machine learning may be used to derive the following model:

[0035] Probability of future pipe damage = f(pipe data + damage data + public data), or Probability of future pipe damage = f(pipe data + public data) Deriving a model based on available data is sometimes referred to as "training." During this training phase, labeled data can be used to iterate and build a model of how features correlate differently with target features. Once a model is built, it is then tested against validation data to assess the accuracy of the model in a process called cross-validation.

[0036] According to aspects of the present disclosure, a form of Random Forest process can be used, which can constitute a regression method composed of many individual decision trees. In particular, decision trees can be used where a series of true / false questions systematically place a piece of data into a particular category, thereby making a "decision" about which category the input belongs to. For example, a simple decision tree can determine whether a pipe is likely to break based on a series of true / false questions (such as questions about weather, location, etc.). Random Forest expands on this and uses many additional decision trees to derive an answer.

[0037] One tree might ask multiple true / false questions about pipe material, pipe diameter, temperature, etc. Another tree might consider the pipe's location and slope, and another might consider weather or traffic data. A random forest then computes multiple different trees that determine one outcome over another, i.e., whether the pipe will break, and the final answer or recommendation can be determined based on which outcome is identified by a majority of the trees or some other threshold.

[0038] FIG. 1 is a schematic diagram illustrating aspects of a processing model used to predict pipe breaks, according to some embodiments. In block 110, utility and public data are cleaned and geo-processed. Prior to performing machine learning analysis, data is collected. For example, the present disclosure provides a web interface through which water utilities can upload utility data, including pipe and damage data. The web interface can be configured to accept various formats of utility data, including, but not limited to, shapefiles, CSV, and GeoJSON. This data can be transferred to a utility database, where utility data from various locations can be stored. For example, utility data from across the country can be stored in a standardized format. The data can also be geo-processed so that it can be accessed or identified in relation to geographic data. The data can be combined with public data (weather, soil, transportation, etc.). For example, a server can be programmed or otherwise configured to access one or more national or nationwide databases and collect relevant public data. Additionally, the server can be configured to automatically associate the collected public data with specific portions of the utility data. For example, weather and soil data for a particular location can be associated with a particular pipe identified in the utility data. New variables can be generated in association with the collected data while maintaining a consistent format.

[0039] A random forest process can then be run multiple times on the collected utility and public data. In block 112, features and target features are separated. In block 114, feature importance is calculated. In block 116, important features are extracted. In block 118, a model is built based on historical data. When applying the random forest process, important variables can be automatically selected, thereby reducing the overall variable size and preventing overfitting of the data. In block 120, the blended model can be run, which generates accurate likelihoods of failure outcomes for both multiple-leaking and non-leaking pipes. From these results, a pipe replacement plan can be created that focuses on areas with the worst likelihood of failure ("LOF"). In particular, the blended model results can be used to generate LOF rankings, where pipes are identified with multiple rankings based on their predicted LOF. Furthermore, a financial simulation can be run that highlights savings from performing a job based on the determined LOF results.

[0040] FIG. 2 is a schematic diagram illustrating aspects of job planning based on models used to predict pipe breaks, according to some embodiments. Most water utilities have pipe and break data. However, sometimes the data is not digitized or some data is missing. According to aspects of the present disclosure, utility customers can be classified into categories based in part on the amount and type of pipe and break data accessible to the customer. Different models can then be used in association with different categories of utility customers. For example, customer categories can be divided into the following five cases: reuse models when the utility is too small or has no digitized breakage data; an ensemble of generic models to predict LOF in the next 1 year when there is sufficient pipe / virtual pipe data but only 1-2 years of breakage data; an ensemble of generic models to predict LOF in the next 3 years when there is sufficient pipe / virtual pipe data but only 3-5 years of breakage data; an ensemble of leak and generic models to predict LOF in the next 3 years when there is sufficient pipe / virtual pipe data but only 6-9 years of breakage data; and an ensemble of leak and generic models to predict LOF in the next 5 years when there is sufficient pipe / virtual pipe data but only 10 years of breakage data.

[0041] 3A-3C are schematic diagrams illustrating aspects of different water utility customer categories, according to some embodiments. FIG. 3A illustrates a reuse model. Some utilities have insufficient pipe and damage data, and there are several reasons why this may occur. One example is a utility that is too small to effectively operate the process of the present disclosure. Another example is a utility that has not digitized or recorded its pipe and damage data. For utilities that have digitized data but insufficient pipe and damage data, it may be possible to use a model built from another utility's data for the target utility. In this case, the utility may be identified based on whether the amount of available pipe and damage data is below a certain threshold.

[0042] Figure 3B illustrates the case of a generic model ensemble. A utility has sufficient pipe data but insufficient damage data. Often, this is because they only have data from the past few years, which may not be enough to build a comprehensive model. To solve this problem, a generic model is built using the utility's damage history. The number of years the model can predict is based on how much damage data the model has.

[0043] Figure 3C illustrates the case of an ensemble of generic and leak models. For utilities with sufficient pipe and damage data, it may be possible to build a unique model for the utility using the utility's pipe and damage data. This model can include both generic and leak models. The utility's data can be uploaded to an integrated utility database. Depending on the size of the utility, the model can be combined with other utility data to build an even more comprehensive model.

[0044] Figure 4 is a schematic diagram illustrating aspects of an architecture for predicting pipe breaks, according to some embodiments. The architecture includes a front-end interface, a management system, and a machine learning system. The front-end interface includes a page for uploading pipe data, damage data, and any auxiliary data; a page for viewing the results of the machine learning analysis in both a map view and auxiliary statistics; a page for downloading maps of cleaned data and machine learning results and downloading statistics; and an interface that allows small utilities to access Fracta's solution even without the correct data or software. The management system includes a management server for creating instances and processes, a database containing customer information, and a file server for storing files. The machine learning system (instance) includes scripts for spatial joins, geoprocessing, and machine learning, and a temporary database for storing files.

[0045] Referring to Figure 4, in step 401, a customer logs in and uploads data. In step 402, the customer's data is uploaded, and in step 403, a request is made to the process manager of the management server. In step 404, an operator (i.e., an operator of a pipe break prediction company) logs in and issues a request to the instance manager of the management server. In step 405, the instance manager issues a request to a machine learning instance of the machine learning system. In step 406, raw files from the file server are loaded into the data process of the machine learning instance. In step 407, the data process inserts pipe and damage data into a GIS database. A geoprocess receives information from the GIS and country databases. In step 408, the geoprocessed information is fed to a predictor. In step 409, the predicted results are transferred to the Likelihood of Break (LOF) results on the file server. In step 410, the LOF results are uploaded to the front-end viewer. In step 411, the customer logs in and views / downloads their LOF data.

[0046] When predicting the probability of a pipe failure within the next five years, aspects of the present disclosure may use the Poisson distribution of the particular data. For example, the probability of failure can be calculated as follows:

[0047] Prob=1-e (-損傷数 / 配管長さ) Instead of a binary classification (is the pipe damaged or not?), the Poisson distribution allows for probability calculations. To arrive at a final result, it is possible to use the following original form of the distribution:

[0048]

number

[0049] where "k" represents the number of events (e.g., damage) that occur within a time period, and "m" corresponds to the determined average number of events per time period. Figure 5 illustrates that damage per unit length (LOF / Length) may be different. A longer pipe 510 may have a different LOF / Length than a shorter pipe 512. Therefore, the length of the pipe (L) can be added to the equation.

[0050]

number

[0051] To calculate the probability of injury in the next five years (instead of one year), the equation can be modified to have five entries:

[0052]

number

[0053] Therefore, this probability is modified as follows:

[0054]

number

[0055] FIG. 6 is a schematic diagram illustrating the use of a time ensemble to calculate the likelihood of failure for the next five years, according to some embodiments. To create a failure likelihood ranking for each pipe segment corresponding to the next five years, the likelihood of failure per length may be calculated for each previous year for which data is available. For each year, the top variables affecting pipe condition may be noted. However, certain conditions may not be the same for a particular year. For example, a drought may occur in a given year, which may affect the importance of a certain variable. To reduce the impact of changes over time, average results for different time slices may be calculated. In this way, the impact of one event (such as a storm, drought, etc.) may be minimized. Furthermore, slowly changing correlations over time may be captured.

[0056] 7A and 7B are schematic diagrams illustrating the performance difference between the conventional technique ( FIG. 7A ) and the technique according to some embodiments ( FIG. 7B ). As illustrated in FIG. 7B , machine learning analysis according to some embodiments has the advantage of being able to accurately assess the condition of pipes that have never been damaged before. In contrast, as illustrated in FIG. 7A , conventional methods used by water utilities to assess the condition of pipes typically only highlight pipes that have leaked multiple times or new pipes that are unlikely to leak. Because utilities focus heavily on damage history, their analysis methods often do not accurately assess the condition of weak pipes that have never leaked. However, through the use of variables that represent the surrounding pipe condition, the system and method of the present disclosure can identify pipes that are more likely to break and assign them a higher break likelihood value. The use of a generic model that does not rely on specific damage data allows for a more accurate assessment of the condition of pipes that have never been damaged before.

[0057] 8A and 8B are schematic diagrams illustrating variables generated through geospatial analysis, according to some embodiments. To create variables representative of the conditions in which each pipe exists, more than a single feature may be required to represent each category. For example, the pH value of the nearest soil region (e.g., a geospatial polygon) may not adequately represent the complex pH variations around the pipe. FIG. 8A shows a representation of polygon data around a pipe of interest, and FIG. 8B shows a representation of more detailed raster data around the pipe of interest. To obtain more detailed data, several steps may be taken. First, the centroid (800) of each pipe segment may be extracted. N circles (e.g., circles 810 and 812) of different dimensions may then be created centered around the centroid of the pipe segment. Statistical values ​​(e.g., maximum, minimum, standard deviation, mean, average, difference between maximum and minimum, number of unique values) may then be calculated based on the geographic variables within the N circles. These values ​​may then be assigned to each pipe segment. By taking values ​​within N circles of multiple radii around the pipe segment, it is possible to determine even more information about the surrounding conditions. Potential variables that can be used are listed in Table 1.

[0058] [Table 1a]

[0059] [Table 1b]

[0060] [Table 1c]

[0061] FIG. 9 is a schematic diagram illustrating the construction of a consolidated utility database for predicting pipe breaks, according to some embodiments. Selected data from utility databases 910, 912, and 914 is consolidated into database 920. The consolidated database 920 is then used to construct a nationally scalable predictive model 930. By compiling all the pipe and damage histories of different utilities across the country, the machine learning models of the present disclosure can be applied to utilities with datasets that are too small or have many missing values. While a specific machine learning model for each utility is often desirable to provide accurate and localized results, there may not be enough data to construct such a model. For example, many small rural utilities often do not have the resources to digitize or otherwise record all of their pipe and damage data. Furthermore, utilities that do have such data may have many missing values ​​that make any analysis unreliable. However, with access to the consolidated utility database of pipe and damage data, a general-purpose machine learning model can be applied to these utilities. For example, a small Northern California utility could use a model built from data from a larger Northern California city. Additionally, this information can be incorporated for use in calculating values ​​for the virtual piping network, as described below.

[0062] Calculating the likelihood of failure for every pipe segment is a problem that often requires access to information about how that data has changed over time. However, for many variables (e.g., damage history, climate, population), only the most recent data may be available because older data has not yet been digitized or otherwise recorded. Furthermore, the relationship of certain variables with time may not be linear. It may be possible to manipulate time-related variables (e.g., pipe age) to better represent the changes in these variables over time. For example, time-related variables may be varied using functions such as logarithm, common logarithm, natural logarithm, square, cube, square root, cube root, exponent, negative exponent, arcsine, arccosine, arctangent, and sigmoid.

[0063] To accurately assess the condition of a pipe segment, damage data and corresponding subvariables (e.g., damage density, damage per mile, etc.) can be included in the machine learning analysis. However, over-reliance on damage data can lead to leakage issues, where the damage data overwhelms other variables. For example, the model may be very good at identifying pipes with multiple damages as faulty, but it may fail to accurately assess pipes with only a few or no damages. Some may suggest removing the damage data entirely, but doing so may result in pipes with multiple damages not being accurately classified as faulty. To address these issues, the present disclosure provides what are called hybrid models. For example, two predictive models can be constructed, one of which highly emphasizes damage (the leakage model) and the other does not incorporate damage history (the general model). Each pipe segment can then be classified as leaking or non-leaking. If the pipe is leaking, the result from the leakage model is assigned. If the pipe is non-leaking, the results from the generic model are used. By averaging both models, the final likelihood of failure per length can be calculated. Alternatively, a weighted average of each model can be taken.

[0064] FIG. 10 is a block diagram illustrating aspects of automated variable selection for model-based pipe damage prediction, according to some embodiments. The present disclosure allows for the collection and creation of over 1,000 variables that can be added to a usable feature set. However, not all of these variables may be used in a particular machine learning analysis. With a large feature set, there is a risk of overfitting. In overfitting, the machine learning model will describe the noise surrounding the model rather than the underlying relationships. The model may accurately describe the training data, but additional data may perturb the model. To prevent overfitting, ensemble methods may be run to reduce the variable set. First, in block 1010, the model is run on the full variable set. In block 1012, variable importance is obtained. The importance of each variable can be obtained, for example, using techniques such as the Gini coefficient or information gain, where higher coefficients correspond to higher importance. In block 1014, blocks 1010 and 1012 are repeated for different year slices. In block 1016, the model can be run again with N variables selected from the list of most important variables. In some embodiments, the variables can be categorized, so that similar variables are given the same category. When selecting variables, it may be beneficial to select only a certain number of variables within the same category. For example, if N variables from the soil pH category have already been selected, less important soil pH variables can be removed. For the "age" and "substance" variable categories, these rules do not need to be applied and all variables can be used.

[0065] According to some embodiments, pipe replacement work can be optimized based on pipe break prediction results. The results of the machine learning analysis of the present disclosure can provide utilities with better insight into the status of all pipes in their network. However, this information may not be sufficient for effective job planning (e.g., determining which pipes to replace). For example, pipes may be ranked from 1 to 5, with "5" representing the highest LOF and "1" representing the lowest LOF. A single rank 5 segment with a high probability of damage may be surrounded by rank 1 pipes with a low probability of damage. Most water utilities do not replace single segments, but rather replace entire areas or blocks of pipe. FIGS. 11A and 11B are schematic diagrams illustrating aspects of pipe replacement optimization, according to some embodiments. In FIG. 11A, a job planning flowchart is shown in which possible combinations of pipe segments are scanned to obtain a job with the required length. FIG. 11B shows the pipe segments scanned in the chart of FIG. 11A. Note that segment "4" was not selected in this case. 1-Π(1-Π) / ... si ) is the probability that one pipe will break; Π(1-P si )=(1-P si )(1-P si )...(1-P si ) is the probability that no pipes are damaged; and Psi =1-(1-P Li ) li is the LOF / segment of the ith pipe, and P Li is the LOF / length of the i-th pipe segment; and li is the length of the i-th pipe segment. In Figure 13A, LOF / construction is 0.87; in Figure 13B, LOF / construction is 0.92; and in Figure 13C, LOF / construction is 0.95. In this example, the job in Figure 13C is the highest priority job of the three.

[0066] This job planning procedure can be implemented as an automated script or in conjunction with other software. For example, a server can be configured to analyze predicted pipe conditions and automatically provide job plan suggestions to the customer based on the overall condition of multiple pipes. Using this tool, utilities can optimize their job planning process to focus on areas with the highest likelihood of damage.

[0067] FIGS. 14A and 14B are schematic diagrams illustrating aspects of predicting missing pipe data values, according to some embodiments. Some utilities may be missing information about their pipe network, such as pipe material or installation year. To address this issue, the missing values ​​can be assigned based on surrounding attributes, such as building, population, street data, and more. These values ​​can also be assigned based on the utility's own data if sufficient information is available. For example, correlations between public and utility data can be found to predict likely values ​​of currently unknown variables. FIG. 14A illustrates extracting correlations between the public database 1410 and the utility database 1412 to build a variable prediction model 1420. FIG. 14B illustrates using the variable prediction model 1420 to predict missing pipe attributes in database 1430 to create a more complete database 1432.

[0068] Figures 15A and 15B are schematic diagrams illustrating the correlation between streets and pipes, according to some embodiments. Pipe data 1510 is closely correlated with street data 1520 for the same geographic area. Many smaller utilities do not have reliable information about their pipes and do not have any Geographic Information System (GIS) data. To work with these utilities, the present disclosure provides for the creation of a virtual pipe network. A virtual pipe network is constructed based on road data to create the desired geospatial information. Information about material, diameter, and installation year is supplemented based on information provided by the utility and data obtained from work with other utilities. Figure 16 is a block diagram illustrating the creation of virtual pipe data, according to some embodiments. Block 1610 shows an area or region that has street data but no pipe data. Block 1620 shows the virtual pipe geometry created using the street data. Block 1630 illustrates the use of public data to predict missing pipe attributes.

[0069] The following is an exemplary case of how the disclosed systems and methods can be used by a customer utility company. In this example, ACME Water is a utility company interested in using software configured to operate in accordance with some embodiments disclosed herein. First, ACME Water can access a web portal and upload ACME Water's pipe and damage data. In this example, ACME Water has a lot of damage data and relatively complete pipe data. FIG. 17 is a schematic diagram illustrating the cleaning of example pipe data in accordance with some embodiments. The uploaded raw pipe and damage data 1710 is cleaned, such as by standardizing the data (e.g., standardizing data that appears under specific data fields) and identifying and correcting any bad data, resulting in cleaned pipe data 1712.

[0070] Figure 18 is a schematic diagram illustrating further processing of example plumbing data, according to some embodiments. Data 1712 can be prepared for machine learning processing. For example, public data 1810 (soil pH, elevation, etc.) can be accessed for the region in which ACME Water is located, and the accessed public data 1810 can be combined with the plumbing data (as shown in Figures 8A, 8B and the associated text). Additionally, time-based variables (e.g., year and corrections for year, such as square root) can be added to obtain processed plumbing data 1820 that accounts for changes in variable characteristics over time, as previously described.

[0071] According to some embodiments, the next step is to run a machine learning process on ACME's cleaned data. Figure 19 is a schematic diagram illustrating the partitioning and extraction of features used for machine learning according to some embodiments. Feature data 1910 is partitioned into features 1912 and target features 1914, and the features are correlated with the target features to calculate the importance of each feature (as shown in Figure 10 and the related description above). The features with their calculated importance are shown in Figure 19 as 1916. Important features can then be extracted from the dataset as 1918, while features with lower importance can be discarded.

[0072] To account for correlations that change over time, several models can be built on the data (as shown in FIG. 6 and the related text above). FIG. 20 is a schematic diagram illustrating the creation of an LOF model through the combination of a leakage model and a generic model, according to some embodiments. To build a model that predicts the years 2018-2022, a model based on known past data is built, while a model that predicts the years 2009-2017 is built (using data from 2004 to 2017). For every "slice" of years, two models can be built: a generic model and a leakage model (as described in FIGS. 3A-3C and the related text above). In FIG. 20, leakage model 2020 is built using leakage model data 2022, and generic model 2030 is built using generic model data 2032. The generic model data 2032 can include all features as input, except for damage. Leakage model data 2022 includes damage and all other features. Using the average of these constructed models, two predictive models for 2018-2022 can be constructed: a general model 2024 and a leakage model 2034. These models can be combined to produce the final mixed model LOF / length (as continuous values) result 2040.

[0073] The LOF / length model 2040, created from the generic model and the leak model, can assign an LOF / length rating or rank to every pipe. According to some embodiments, these results can be ranked, and the customer (ACME Water) can view the LOF / length results for their network on a web interface. According to some embodiments, the customer interface allows filtering and sorting of pipes based on the LOF results or other assigned variables. With these LOF / length predictions (as continuous values), a pipe replacement plan can be created for ACME Water. Figure 21 is a schematic diagram illustrating aspects of job planning, according to some embodiments. A server running job planner software can create job plans based on the highest LOF / length in an area. As previously mentioned, the job planner can associate pipes with the highest grouped LOF / length.

[0074] 22 is a schematic diagram illustrating a system for predicting pipe breaks, according to some embodiments. The server can include a processor 2212 coupled to memory 214, a database 2216, input / output devices 2218, and one or more networks 2220. Additionally, the described methods can be implemented in conjunction with the example systems shown above.

[0075] Memory 2214 may store program instructions for different programs run by the various servers described herein, including the front-end web interface server and back-end server shown and described in connection with Figure 4, which are used to implement the systems, interfaces, methods, and computer-implemented processes described herein. Processor 2212 executes the program instructions to cause the processor, platform, server, computer, or other device to interact with other elements coupled to it directly or indirectly via a network or bus.

[0076] Processor 2212 can be coupled directly or via the Internet, a local area network, a wide area network, a wireless network, or other networks to various databases, customer devices, administrator devices, and other devices. Databases 2216 and 2226 can be third-party databases containing piping, public, and private information having an impact on the likelihood of one or more pipe or pipe section failures, as described herein. Databases 2216 and 2226 can be maintained by a third party or by the same entity that implements the systems and methods described herein.

[0077] One or more customers or subscribers can be coupled to the system shown in FIG. 22 via a network. For example, a user can be coupled to the processor 2212 and front-end and back-end systems described herein via a wireless network, such as a 3G, 4G, or 5G wireless network, via a WiFi network, or via some other network connection. Customer or subscriber information can be stored by the system in a database (such as database 2216 and / or 2226), and each customer can be provided with access to information about the pipe network specific to that customer, which may be, for example, a local or regional water utility. There may be multiple customers and water utilities using the same platform and interfacing with the platform via a web interface that provides each customer with access to data and allows the customer to create projects specific to their pipe network. Alternatively, each customer may have an application that interacts with the system to exchange data about the customer's pipe network and projects. The machine learning aspect of the system can leverage pipe and pipe damage data across all pipe networks in the system to facilitate more accurate break likelihood predictions over time.

[0078] According to some embodiments, predictions are made for LOF / segment, LOF / length, and the number of damages in the next N years. Note that if the predicted LOF / segment (sometimes referred to herein simply as "LOF") for the next N years is correct, then calculating the predicted number of damages is straightforward. The sum of the LOFs will equal the predicted number of damages in the next N years, assuming that none of the pipes will have multiple damages in the next N years.

[0079] However, LOF usually contains a certain amount of error, resulting from overconfidence or underconfidence in the prediction model. This can lead to deviations from the predicted number of damages in N years, calculated by summing LOFs. For example, if the number of damages in the most recent five years is 120, the predicted number of damages may be 180 (>>120). While this may be true, it represents such a large increase that there is some probability that this is an unreasonably high number. Furthermore, this can occur even when the sorted LOF order is very good ("good sorted LOF order" means "good prioritization of good and bad pipes." Prioritization is important because it is highly relied upon when planning pipe replacement). The above-mentioned problems have been found to be caused by inaccuracies in the LOF scale.

[0080] 23 is a schematic diagram illustrating predicting the number of injuries from LOF according to some embodiments. The predicted number of injuries in year N may be the sum of the LOFs for the next N years, as shown in FIG. 23. There is usually a gap between the predicted number of injuries and the expected results. If N years is a short or medium period, it can be assumed that there is not much difference between the injuries in the most recent N years and the injuries in the next N years.

[0081] Figure 24 is a diagram illustrating the expected similarity in injuries for the past five years and the next five years. As shown, the number of injuries in the past five years and the predicted number of injuries in the next five years are similar. Figure 25 is a diagram illustrating that the predicted number of injuries in the next N years can be calculated from injury history, based on some embodiments. In this case, the predicted number of injuries in the next five years is a total of 120 NBH.

[0082] 26 is a schematic diagram illustrating LOF probability calibration, according to some embodiments. LOF is the predicted number of damages from the damage history (predicted NBH ), predicted number of injuries from LOF (predicted N LOF ) can be calibrated. According to some embodiments, the LOF for leakers and non-leakers can be calibrated in the same way. According to some embodiments, the LOF can be used to calculate the BRE (Business Risk Exposure), which is an estimate of the damage if the pipe breaks. The BRE is the product of the LOF and the COF (Consequence of Break). Even if the COF were perfectly correct, the BRE would be overestimated, so the LOF would be well calibrated.

[0083] According to some embodiments, the following algorithm can be used for the calibration:

[0084] i) Calculate the predicted number of damages based on the utility's damage history (1) Remove outliers from the damage history by the following process: (a)Nb n If the following formula is satisfied, the number of damages in the latest year (Nb n ) may not be assumed to have all the damage for that year: Nb n <c i Nb n-1 n: latest year 0 <c i <1.0 (b) Use the four-point range method to remove outliers. If the number is outside the range, it can be removed: [q I -c2(IQR),q J +c2(IQR)] q I : Ith quartile, q J :Jth quartile, I< <J IQR=q j -q I , 1.0≦c2≦1.5 (c) Remove the number if the number is less than c3 × M.

[0085] M: Average, median, maximum or minimum number of injuries in the last N years 0 <c3<1 (2) Use the remaining numbers to obtain a predicted number of injuries from the injury history by using averages, weighted averages, linear regression, exponentials, logarithms, or machine learning.

[0086] ii) Calibrate LOF / segment from the ratio between the predicted number of damages from itself and damage history and LOF. Calibrated LOF / segment = LOF / segment x x x = predicted N BH / Predicted N LOF Predicted N BH : Predicted number of damages based on damage history Predicted N LOF : Estimated number of injuries from LOF If calibrated LOF / segment >= 1.0, then calibrated LOF / segment is set to 1-c, where c is > 0 and very close to 0 (so, for example, LOF / segment is 0.99999...).

[0087] iii) LOF / Length Calibration: Calibrated LOF / Length = 1-(1-Calibrated LOF / Segment) 1 / L L: Pipe length As mentioned above, the machine learning model described thus far can use over 1,000 variables to calculate the importance of each variable based on techniques such as the Gini coefficient or information gain. According to some embodiments, the variables can be grouped into 14 categories, such as soil properties, topography, climate, population, buildings, transportation, water areas, land division, coastline, age, damage history, diameter, material, and pressure. Examples of "soil properties" include pH, CaCO3, bulk density, and water content. The model can use several variables from the same category to calculate LOF.

[0088] The model can be considered a "black box" if the importance of its features is unknown. A utility company may want to understand which attributes affect pipe deterioration because they may be able to use the knowledge to maintain and manage their pipes. The categories can be ranked from 1 to 10 depending on importance, or the categories can be sorted by importance and assigned according to the sorting order. "10" represents the most important feature and "1" represents the least important. Figure 27 is a plot showing an example ranking of variable categories, according to some embodiments.

[0089] According to some embodiments, feature importance can be automatically calculated and categorized by following the techniques shown in FIG. 10 and described above in the associated text.

[0090] 28 is a diagram showing an example of calculating representative feature importance for soils, according to some embodiments. The representative importance for each category can be the maximum, average, weighted average, or median of the importance within each category.

[0091] 29 is a schematic diagram illustrating the normalization of the representative feature importance of the categories. The representative importance for each of the categories (e.g., the 14 categories listed above) can be normalized using, for example, the following formula:

[0092] X normalized =(XX min ) / (X max -X min ) X normalized =(X-μ) / σ X=[x 土壌 ,x 地勢 ,...] x カテゴリー :Representative importance of each category FIG. 30 is a schematic diagram illustrating the ranking of multiple representative features by importance from 1 to 10. In the illustrated example, there are 14 representative feature categories. Each representative importance is then ranked based on its normalized value on a scale of 1 to 10. According to some embodiments, after normalizing the importance, the normalized importance is divided evenly into 10 levels. In one example, if the maximum normalized importance is 0.8 and the minimum normalized importance is 0.0, the following normalized importance ranges can be assigned values ​​from 1 to 10 as follows: "10" for 0.72 to 0.8; "9" for 0.64 to 0.72; "8" for 0.56 to 0.64; "7" for 0.48 to 0.56; "6" for 0.40 to 0.48; ...; "2" for 0.08 to 0.16; and "1" for 0.0 to 0.08. However, when using this even splitting technique, some assigned importance values ​​may be skipped. To avoid this, it is possible to use a sorted value based assignment as described above.

[0093] According to some embodiments, the range of forecast years provided by the model depends on the utility's available damage history. For example, the machine learning model can provide a 5-year LOF if the utility has a sufficiently long damage history. However, if the available damage history is insufficient, the model can only provide a 3-year LOF. While some utilities do not have a long enough damage history to predict 5 years, most utilities would like to know the 5-year LOF. Furthermore, many utilities prefer a short-term LOF, such as a 1-year LOF, to optimize current operations, while planning to use a long-term LOF, such as a 3- or 5-year LOF, for replacement planning. To help utilities with their current operations, it is possible to approximately calculate an N-year LOF from an M-year LOF, assuming that damage behavior does not change.

[0094] For example, the following method can be used to predict the 5-year LOF from utility data and to approximately calculate the 1- and 3-year LOF: First, predict the LOF (P) for year M. Then, the probability that the pipe will not be damaged in year M is 1-P. The probability that the pipe will not be damaged in year N is (1-P). N / M Therefore, the LOF for year N is 1-(1-P) N / M The following is an example:

[0095] 5-year LOF=P (projected from utility data) 1 year LOF = 1-(1-P) 1 / 5 3-year LOF = 1-(1-P) 3 / 5 While the above-described examples primarily relate to networks of underground pipes, according to some embodiments, many of the techniques described above may be applied to other types of networks. According to some embodiments, the systems and methods described herein are applied to networks of above-ground utility poles and / or electric wires used to deliver power to consumers, such as between underground nodes. According to some further embodiments, instead of or in addition to electric wires, utility poles themselves may also be treated as target assets. When applying the techniques described herein to other types of networks and / or assets, a different set of environmental variables may be used. For example, for above-ground electric wires, a subset of environmental variables may be used rather than all the environmental variables used for underground pipes. In this case, it may be possible to remove soil from the set of variables if the wire is above ground. The meaning of a break event should also be redefined. For example, for electric wires, a break may mean a broken wire, a deterioration in condition or strength, etc. For utility poles, a break may mean damage, a deterioration in condition, or a deterioration in strength.

[0096] Although the foregoing has been described in some detail for purposes of clarity, it will be apparent that certain changes and modifications can be made without departing from the principles thereof. It should be noted that there are many alternative ways of implementing both the processes and apparatus described herein. Accordingly, these examples should be considered as illustrative and not restrictive, and the invention described herein should not be limited to the details set forth herein; it can be modified within the scope of the appended claims and their equivalents.

Claims

1. A method for predicting pipe leakage, implemented by a computer system, comprising: and constructing and predicting a mathematical model based on utility data and public data related to the geographic data; the utility data includes data based on piping data and damage data; When the pipe to be predicted is a pipe that has never experienced a leak, predicting the likelihood of a leak from the pipe based on at least the pipe data among the public data and the utility data. Leak prediction methods.

2. A leakage prediction method as described in claim 1, wherein the piping data includes at least some of the data regarding the installation location, installation time, material, diameter, and fluid pressure of the piping.

3. A leakage prediction method as described in claim 1 or claim 2, wherein the public data includes at least one of data regarding weather, meteorology, soil, transportation, topography, and population.

4. A piping leakage prediction device, a receiver for receiving utility data and public data related to the geographic data; a model construction unit that constructs a mathematical model based on the utility data and the public data; a leak prediction unit configured to predict the likelihood of a leak in at least a portion of the piping based on the mathematical model; the utility data includes data based on piping data and damage data; the leak prediction unit predicts the likelihood of a leak from the pipe based on at least the pipe data among the public data and the utility data when the pipe to be predicted is a pipe that has never leaked. Leak prediction device.

5. A pipe leak prediction system comprising: at least one first server that provides utility data related to geographic data; at least one second server that provides public data; and a pipe leak prediction device, The leakage prediction device includes: a receiving unit for receiving the utility data and the public data; a model construction unit that constructs a mathematical model based on the utility data and the public data; a leak prediction unit configured to predict the likelihood of a leak in at least a portion of the piping based on the mathematical model; the utility data includes data based on piping data and damage data; When the pipe to be predicted is a pipe that has never experienced a leak, predicting the likelihood of a leak from the pipe based on at least the pipe data among the public data and the utility data. Leak prediction system.

6. A program that causes a computer to function as a piping leakage prediction device, The computer receiving means for receiving utility data and public data related to the geographic data; a model building means for building a mathematical model based on the utility data and the public data; acting as a leak prediction means configured to predict the likelihood of a leak in at least a portion of the piping based on the mathematical model; the utility data includes data based on piping data and damage data; the leakage prediction means predicts the likelihood of leakage from the pipe based on at least the pipe data among the public data and the utility data when the pipe to be predicted is a pipe that has never leaked; program.