Processing data to predict pipe breaks
By generating complementary data and creating a standardized database, the method addresses data consolidation and quality issues in water utility networks, enhancing predictive analysis and maintenance planning for underground pipes.
Patent Information
- Application Number
- JP2020552731
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-10-09
- Filing Date
- 2019-03-26
- Publication Date
- 2025-09-25
- Estimated Expiration
- 2039-03-26
AI Technical Summary
Water utility companies in the US face challenges in consolidating and maintaining the quality of their pipe network data, which is often stored in various formats and contains errors, limiting accurate analysis and utility effectiveness.
A method and system for improving pipe data by receiving and generating complementary data using geolocation transfer, machine learning, and creating a standardized database, including environmental data, to manage underground pipe networks effectively.
Enhances data accuracy and computational efficiency, enabling predictive analysis and maintenance planning for underground pipes, reducing human error and computational costs.
Smart Images

Figure 0007744131000002 
Figure 0007744131000003 
Figure 0007744131000004
Abstract
Description
[Technical Field]
[0001] This patent application claims priority to and incorporates by reference each of the following provisional applications:
[0002] U.S. Provisional Application No. 62 / 649,058, filed March 28, 2018; U.S. Provisional Application No. 62 / 658,189, filed April 16, 2018; U.S. Provisional Application No. 62 / 671,601, filed May 15, 2018; U.S. Provisional Application No. 62 / 743,477, filed October 9, 2018; U.S. Provisional Application No. 62 / 743,483, filed October 9, 2018; and U.S. Provisional Application No. 62 / 743,485, filed October 9, 2018.
[0003] This patent application is related to and incorporates by reference the following U.S. and PCT applications filed on even date herewith, collectively referred to as the "co-pending patent applications": U.S. Patent Application No. 16 / 365,466 (Attorney Docket No. Fracta-002-US); U.S. Patent Application No. 16 / 365,522 (Attorney Docket No. Fracta-006-US); and International Patent Application No. PCT / US19 / 24139 (Attorney Docket No. Fracta-002-PCT).
[0004] This patent application generally relates to improved systems and methods for processing data related to underground pipe networks. More particularly, some embodiments relate to methods and systems for improving pipe data related to underground pipe networks that carry fluids to consumers. Some embodiments relate to methods and systems for creating a standardized database containing environmental data that may be associated with underground pipe networks. [Background technology]
[0005] Water utility companies in the United States have accumulated millions of data points about their systems, collecting information about their pipe networks and recording damage that occurs over time. This data is useful to these companies for everything from understanding the age and configuration of their systems to planning pipe replacement and addressing damage.
[0006] However, the approximately 50,000 water utilities operating in the United States do not consolidate their data, and much of that data is stored in a variety of formats, limiting their ability to combine their data with other forms of information and often making it impossible to perform accurate analyses of their systems.
[0007] Furthermore, most utilities do not maintain the quality of their data sets. The data some utilities have may contain errors and affect the accuracy of their analyses. For example, some utilities do not geolocate the location of previous pipe damage or identify which pipes have previously been damaged. These issues significantly limit the accuracy of analyses related to utility data as well as the effectiveness of the utility's work.
[0008] Additionally, various organizations across the United States collect millions of data points related to environmental and geospatial attributes such as soil properties, population, weather, and others. These data points are useful for use in geospatial analysis and correlating relationships between these variables and other characteristics. For example, these datasets can be extremely useful in calculating the likelihood of a water main break across the country. Unfortunately, however, these datasets are not available in their raw form. Summary of the Invention [Problem to be solved by the invention]
[0009] According to some embodiments, a method for improving pipe data related to a network of underground pipes for carrying fluids to consumers is described. According to some embodiments, the pipes may be used to carry other types of fluids, such as wastewater, recycled water, brackish water, storm water, seawater, drinking water, steam, compressed air, oil, and natural gas. The method includes receiving a set of pipe data including a plurality of pipe attributes for one or more of the networks of underground pipes, identifying one or more pipes in at least one of the networks for which one or more pipe attributes are missing or inaccurate, automatically generating complementary data for the one or more pipe attributes resulting in an improved set of pipe data, and managing one or more aspects of one or more of the networks of underground pipes based on the improved set of pipe data. [Means for solving the problem]
[0010] According to some embodiments, the complementary data is automatically generated based on a geolocation transfer process using one or more sources of information, such as GPS, addresses, and / or geocodes. The set of piping data can include pipe damage data having locations for the pipe damage but not specific pipe sections for the damage. The geolocation transfer process can assign a pipe section to at least some of the pipe damages that is not the closest pipe section to the respective location for the pipe damage. According to some embodiments, the complementary data can include using parcel data to associate geographic positions with some of the pipe sections.
[0011] According to some embodiments, automatically CompletionExamples of the missing or inaccurate pipe attribute data that may be present include installation year, pipe material, and pipe diameter. According to some embodiments, the likelihood of failure of one or more pipe sections can be predicted based on the improved set of pipe data. The prediction can be based on a model constructed with machine learning using the improved set of pipe data.
[0012] According to some embodiments, a system for improving pipe data associated with a network of underground pipes for carrying fluids to consumers is described. The system includes a database storing a set of pipe data including a plurality of pipe attributes for one or more of the networks of underground pipes; and a processing system configured to identify one or more pipes in at least one of the networks for which one or more pipe attributes are missing or inaccurate and generate complementary data for the one or more pipe attributes resulting in an improved set of pipe data; and is capable of managing one or more aspects of one or more of the networks of underground pipes based at least in part on the improved set of pipe data. According to some embodiments, the system may also include a front-end system including an uploader for receiving the set of pipe data from a customer managing the one or more aspects of at least one of the networks of underground pipes, and a viewer capable of displaying at least some of the improved set of pipe data to simplify the management. According to some embodiments, the processing system is further configured to predict the likelihood of a leaking pipe segment in the network and replace a pipe section in one or more aspects of the managed network of underground piping based at least in part on the predicted likelihood of a leak.
[0013] According to some embodiments, a method for creating a standardized database containing environmental data is described. The method may include accessing a first plurality of separate environmental raster-type datasets; combining at least some of the first plurality of separate environmental raster-type datasets to generate a second plurality of larger environmental-type datasets; vectorizing at least some of the second plurality of larger environmental-type datasets to generate a third plurality of larger vector datasets; combining at least some of the third plurality of larger vector datasets to form a combined vector dataset; and completing one or more missing attributes in the combined vector dataset to generate the standardized database. In some cases, the combined raster data does not overlap and / or is not coextensive with one another.
[0014] According to some embodiments, outlier data is removed from the standardized database based on falling outside one or more predetermined limits. The standardized database may include one or more new environmental variables, such as soil density, population density, and national gradient, that were not included in the original environmental raster data.
[0015] The standardized database can be used by multiple customers to provide computational cost savings to each customer compared to using a custom-generated database. According to some embodiments, the standardized database is used to predict the likelihood of a leak in a pipe segment in a network of underground pipes that carry fluids to consumers. The prediction can be made using a model built using machine learning based on the standardized database.
[0016] As used herein, the grammatical conjunctions "and," "or," and "and / or" are all intended to indicate that one or more of the objects or subjects at which they are connected may occur or be present. Thus, as used herein, the term "or" in all instances is intended to indicate an inclusive or rather than an exclusive or.
[0017] To further clarify the above and other advantages and features of the subject matter of the present patent application, specific examples of embodiments thereof are illustrated in the accompanying drawings. It should be understood that elements or parts illustrated in one drawing may be substituted for equivalent or similar elements or parts illustrated in another drawing, and that the drawings merely illustrate exemplary embodiments and, therefore, should not be considered as limiting the scope of the present patent specification or claims. The subject matter herein will be described and explained with additional specificity and detail using the accompanying drawings. [Brief explanation of the drawings]
[0018] [Figure 1] FIG. 1 is a block diagram illustrating aspects of data cleaning for a water utility company, according to some embodiments. [Figure 2] 1 is a schematic diagram illustrating a possible architecture for a wrangling system for a water utility company, according to some embodiments. [Figure 3] 10 is a schematic illustrating the automatic filling of installation year data with other piping values according to some embodiments. [Figure 4] FIG. 10 is a diagram illustrating aspects of automatically populating pipe installation year data using parcel data, according to some embodiments. [Figure 5] 1 is a table showing reference material installation year ranges based on several examples. [Figure 6] 10 is a diagram illustrating the supplementation of pipe material information based on installation year and water utility industry information, according to some embodiments. [Figure 7]10 is a diagram illustrating automatically filling in pipe diameter data based on other pipe characteristics, according to some embodiments. [Figure 8] 1 is a diagram illustrating geocoding of injury data, according to some embodiments. [Figure 9] 1 is a diagram illustrating the linking of piping data with damage data, according to some embodiments. [Figure 10] Schematic diagram showing data received from a customer in an exemplary case. [Figure 11] 1 is a schematic diagram illustrating the process of cleaning plumbing data for an exemplary case customer. [Figure 12] 1 is a diagram illustrating the process of cleaning damage data for an exemplary case customer. [Figure 13] (A) and (B) are schematic diagrams illustrating potential problems with raster data. [Figure 14] FIG. 1 is a block diagram illustrating a process for coalescing raster data, according to some embodiments. [Figure 15] FIG. 1 is a diagram illustrating vectorizing raster data, according to some embodiments, whereby the vector data can be used to use national data for analysis. [Figure 16] 1 is a diagram illustrating splitting raster data in preparation for vectorization, according to some embodiments. [Figure 17] 1 is a schematic diagram illustrating vertical concatenation of CSVs according to some embodiments. [Figure 18] 1 is a flowchart illustrating a hierarchy for imputing missing attributes in environmental data, according to some embodiments. [Figure 19] Schematic diagrams showing examples of new environmental variable generation according to some embodiments. [Figure 20] FIG. 1 is a block diagram illustrating generating national density data, according to some embodiments. [Figure 21]1 is a schematic diagram illustrating a raster matrix used for gradient calculations, according to some embodiments. [Figure 22] FIG. 10 is a block diagram illustrating the creation of new variables from line data, according to some embodiments. [Figure 23] 1 is a diagram illustrating an example of recategorizing zoning data, according to some embodiments. [Figure 24] 1A and 1B are schematic diagrams illustrating an example of approximating square-shaped data, according to some embodiments. [Figure 25] FIG. 10 is a block diagram illustrating an example of integrating grid shape vectors according to some embodiments. [Figure 26] 1 is a schematic diagram illustrating a geodatabase system architecture, according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0019] Several examples of preferred embodiments are described in detail below. While several embodiments are described, it should be understood that the novel subject matter described in this patent specification should not be limited to any one embodiment or combination of embodiments described herein, but encompasses numerous alternatives, modifications, and equivalents. Furthermore, while numerous specific details are set forth in the following description to provide a thorough understanding, some embodiments may be practiced without some or all of these details. Furthermore, for convenience of explanation, certain technical matters known in the relevant art have not been described in detail to avoid unnecessarily obscuring the novel subject matter described herein. It is apparent that one or more individual features of specific embodiments described herein can be used in conjunction with or in combination with features of other described embodiments. Furthermore, like numbers and symbols in the various drawings represent like elements.
[0020] According to some embodiments, a standardized process for receiving and converting utility pipe and damage data is described. To accomplish this, an automated approach is provided for data processing based on research in statistical methods as well as relevant water industry practices. The various systems and methods described herein, according to some embodiments, provide numerous improvements over the use of prior art. The various systems and methods described herein have been found to provide substantial improvements in the quality of data related to underground pipe networks. For example, some aspects of the present disclosure may provide one or more of the following improvements over the prior art: reducing errors in datasets through standardization and the use of programmatic procedures; imputing missing data based on statistical methods as well as historical water industry data; identifying and correcting errors in both pipe and damage data; and providing geolocation of damage data for map display and identifying the pipe where the damage occurred.
[0021] According to some embodiments, a system for prioritizing the replacement of underground pipes is described. The system includes a database storing information including pipe data, pipe damage data, and external data including geographically specific data; a memory storing at least one program having program instructions; a network interface coupled to at least one computer; and a processor coupled to the database, the network interface, and the memory. The processor is capable of executing the program instructions of the at least one program to cause the processor to receive a set of pipe data including a plurality of pipe attributes; identify one or more pipes from the plurality of pipes in which one or more pipe attributes are missing or inaccurate; access parcel data related to geographic locations corresponding to the one or more pipes; generate complementary data for the one or more pipe attributes based on the parcel data; and transmit cleaned pipe data including the complementary data.
[0022] According to some further embodiments, the processor is also configured to receive damage data; access geographic data associated with the damage data; generate geolocation data for the received damage data based on the accessed geographic information; and predict a likelihood of a future failure for at least a subset of the plurality of pipes based at least in part on at least one of the interpolation data and the geolocation data.
[0023] According to some further embodiments, the damage data includes a street address associated with the occurrence of a pipe damage, and the geographic data has a geographic location associated with the street address. Additionally, the processor may be further configured to identify a particular pipe from a plurality of pipes associated with the street address and generate geolocation data by associating the damage data with the particular pipe.
[0024] According to some other embodiments, the processor may be further configured to automatically reformat the received pipe data and damage data. Reformatting may include converting pipe attributes to a standardized acronym or numeric standardized form. The parcel data may include one or more dates of construction for one or more structures associated with a geographic location corresponding to the one or more pipes, and the one or more attributes may include a pipe installation date. The pipe attributes may include at least one of a pipe ID, a pipe material, and a pipe size.
[0025] According to some other embodiments, the processor is further configured to generate complementary data for the first pipe attribute based on second pipe attributes, the first and second pipe attributes being associated with the same pipe or with pipe having one or more similar attributes. The processor may be further configured to determine pipes similar to the one or more pipes for which one or more attributes are missing or incorrect, and the complementary data is based at least in part on attributes of pipes similar to the one or more pipes. The processor may be further configured to provide a display of a visual representation of the pipe attributes determined to be missing or incorrect and to provide a display of a visual representation of data entries including the complementary data.
[0026] According to some embodiments, it is possible to create a national database filled with transformed, optimized, merged, and supplemented environmental data. A programmatic approach using geoprocessing is described. The various systems and methods described herein have proven to be substantial improvements. The creation of a standardized database, such as that described herein, containing national data has proven useful for many types of analysis. The improvements resulting from such a database include one or more of the following: substantial increases in computational efficiency and reduced uptime for applications; substantial reductions in real-time software crashes due to pre-processing of data; and the ability to include and use data from a wider range of sources.
[0027] As used herein, the following terms have the following meanings: "wrangling" is a data cleaning process; "utility data" is data provided by a water utility; the category of utility data may include pipe data and break data; "pipe data" is geographic pipe data that may include information regarding installation year, material, diameter, etc.; "break data" / "break history" is pipe break records that may include location-related pipe ID and date information; "parcel data" is data containing information about a particular piece of land, including when buildings or other structures on that piece of land were constructed; "shapefile" is a type of file format that may contain a directory containing geographic information, including shape information and supplemental information (such as projection or column attributes); "geocoding" is the process of associating specific objects or events with one or more geographic locations, such as latitude and longitude points; "raster data" is geospatial information stored as pixel data / pixel arrays and including pixel resolution and at least one point of geographic location (e.g., raster data can be overlaid on a geographic map), examples include tif and geotiff; "vector data" is geospatial information including geometry and attributes in a tabular format, examples include shp, gpkg, and geojson; "geodatabase" is a relational database that can handle geographic information, for example, allowing data to be selected based on location; "WKT geometry" is geometry information in the form of Well-Known-Text, examples include LINESTRING(30 10, 10 30, 4040), where "CRS" is a coordinate reference system, also known as SRS (Spatial Reference System), is used to locate geographic entities and consists of three elements: a coordinate system, a datum, and a projection. One of the most common datums in use is the World Geodetic System, as revised in 1984 (WGS84), which is used for applications such as GPS; "line data" is geographic line data, including roads (major roads, highways, etc.), railways (conventional railways, light railways, subways, intercity railroads, etc.), linear waters (rivers, streams, canals, drainages, etc.), coastlines, and more; "point data" is geographic point data, including rail stations (railroad stations, rail depots, toll stations, etc.), bus stops (bus stops, bus stations, etc.); and "public data" is data that is publicly available and / or available through government sources, such as soil data, weather data, etc. It should be noted that the use of the term "public" when used in "public data" does not necessarily mean that the data is generally available to the public free of charge. Instead, it is one in which the data is available from a pooled source such as a government agency (eg, USGS soil data).
[0028] When it comes to data storage and processing of pipe and damage data, methods used by U.S. water utilities and others are inconsistent and inaccurate. These methods can result in poor analysis and gaps in the data. Automating data processing, including cleaning steps, is important to ensure utility data is accurate and complete. The automated process described herein can prevent human error and can be implemented more quickly than traditional manual techniques. The process described herein is also scalable, allowing the same set of rules and assumptions to be applied to multiple utilities. Using a standardized format helps ensure that data is compatible across different utilities, thus enabling more powerful analyses with more available data. Many of the systems and methods described herein also provide for the generation of new data. The generated data can be used as part of an analysis of a utility's infrastructure. For example, the generated data can be used to more accurately determine the probability of future pipe damage and to identify the types of maintenance and replacement work that should be performed.
[0029] Data wrangling (which can also be referred to as data cleaning) is a critical but potentially time-intensive process for analyzing large infrastructure such as utility pipes. Data cleaning alone can often take up more than 80% of the time for an analysis project. Part of the reason for this length of process is the fact that input data can arrive in a multitude of different formats. Occasionally, the data may be exceptionally clean with few gaps. However, in most cases, it is necessary to thoroughly clean the data by correcting values and imputing missing data.
[0030] In the case of data from water utilities in the United States, there are many possibilities for how the data is stored and recorded. For example, when recording the name of the pipe material, there is no consistent standard among all utilities for how to record this information. One utility may refer to cast iron piping as "CIP," while another may refer to it as "cast iron." It is important that this information is consistent when performing analysis across multiple water utilities. Another issue with utility data is the prevalence of errors. Because pipe and damage information is entered manually, there is a high potential for human error. Poorly entered data will result in inaccurate results, and small errors can have a significant impact on further analysis.
[0031] Furthermore, sometimes water utility companies do not collect enough information to complete an analysis. In these cases, it is necessary to generate this data. For example, not all utilities capture geolocation information for damage data, which can be important for running geospatial analyses and identifying which pipes have damage. To mitigate this, damage data may be located based on street address. Another example of insufficient data is pipe data that does not identify the year the pipe was installed. Such information is important for running year-based analyses and for completing missing material data. Therefore, the systems and methods described herein enable the generation of additional information based on utility data and available public data. For example, the system can complete the date a pipe was installed by analyzing the year or years that one or more nearby buildings were constructed.
[0032] FIG. 1 is a block diagram illustrating aspects of data cleaning for a water utility company, according to some embodiments. To automatically process utility data, the data can be uploaded to one or more servers, as shown in block 110. For example, a water utility company can upload pipe and damage data in any number of formats, including, for example, shapefiles, CSV files, Excel files, gdb (geodatabase) files, or GeoJSON. Additionally, the water utility company can fill out a questionnaire that includes questions about their data and the acronyms they use. Once the utility data is uploaded, it can be reviewed by an operator, who can determine if additional data is needed. The data can be uploaded to a file server, and an instance of an automated data wrangling tool can be spawned on it.
[0033] In block 112, the automated data wrangling tool can fill in and correct various information for the pipe data. Missing or inaccurate information includes pipe material, diameter, installation year, surface area, and length, all of which can be automatically filled in and corrected by the data wrangling tool. For example, if a utility's data does not include the installation year for a particular pipe, the automated data wrangling tool can use parcel data to complete the pipe installation year. In particular, the parcel data may include a date or dates on which buildings on or nearby the pipe were constructed. All pipes within a certain distance of these buildings can be assigned an installation date that corresponds to the construction date of the building. The data processing tool can standardize column names and material acronyms.
[0034] Upon cleaning the utility's damage data, the automated data wrangling tool supplements and corrects the damage data and matches the damage to the appropriate pipe in block 114. If the damage data is not geolocated, the automated data wrangling tool can geolocate the damage given an address or latitude and longitude.
[0035] Once the automated process is complete, both the pipe and damage data can be output at block 116 for customer download or in preparation for further analysis.
[0036] FIG. 2 is a schematic diagram illustrating a possible architecture for a wrangling system for water utility data, according to some embodiments. According to some embodiments, information can be acquired and provided to customer utility companies via a web interface. For example, in step 201, one or more front-end web servers can be configured to provide utility customers with secure access (via secure login) to a web interface that allows each utility customer to upload their utility data. Additionally, the web interface includes a viewer (205) that allows the utility customer to view / download information, such as analysis and statistics, of their utility data. The web interface can include multiple pages. For example, the front-end system can provide pages for uploading pipe data, damage data, and refill data; text boxes for entering material acronyms and column names; and a page for downloading wrangled data, including a map of the cleaned data. A management system can provide a management server for instance and process creation, a database containing client information, and a file server for hosting raw and cleaned files. A wrangling instance can provide automated wrangling scripts and temporary and / or permanent databases for hosting files.
[0037] Prior to processing of the pipe and damage data, the data can be uploaded to a file server where it is processed. According to some embodiments, a web portal is provided through which the utility company can upload pipe data, damage data, and supplemental information such as material acronyms and column names. For example, utility data can be uploaded as a shapefile, CSV file, Excel file, or gdb (geodatabase) file. The supplemental information provided by the utility includes data column definitions, a list of known materials used, and identification of unknown material acronyms.
[0038] FIG. 3 is a schematic diagram illustrating the automatic filling of installation year data with values from other pipes, according to some embodiments. Without installation year data, it may be difficult to accurately analyze pipe conditions or complete other pipe characteristics. According to some embodiments, utility data, such as a utility's pipe installation year data, is automatically filled and / or corrected. In the example shown in FIG. 3 , a utility has some pipe installation year data. The server is configured to automatically fill in the remainder using patterns from other data sources. In such cases, the server can identify rows with missing installation year values. These rows can be filled based on the installation years of other rows containing installation year values. The server can analyze pipes with similar attributes to a particular pipe with missing information, such as a missing installation date. The server can determine installation dates for pipes located within a certain distance from the particular pipe and / or constructed of the same material and / or the same diameter as the particular pipe. The server can then assign data fields for the particular pipe with dates corresponding to pipes with similar attributes. For example, the server can fill in a data field for a cast iron pipe with a missing installation year by using the median installation year of other non-empty cast iron pipes of a similar diameter. If the data includes an installation date instead of a year, the server can automatically convert the date to a year for easier use, or vice versa.
[0039] According to some embodiments, if the utility does not have installation year data or has installation year data that is missing within a certain threshold, this data can be supplemented using parcel data. Parcel data can be collected, including information related to the date a building or other structure was constructed on a particular parcel of land. The parcel data can be collected from privately or publicly available data. For example, a server can be configured to conduct a search of various databases or online sources to identify parcel data associated with parcel data on or near a particular pipeline of interest. The server can also be configured to receive parcel data directly from a utility company or other user via a web page. The parcel data can be converted to a particular format and wrangled to correct errors present in the parcel data. For example, parcel data from multiple sources can be compared to each other, and corresponding construction dates can be used based on the dates provided by the multiple sources. In particular, if different construction dates are provided for a particular building, a median date can be used. Alternatively, the server can be configured to use the earliest or latest available construction date for a particular type of building or structure. Additionally, empty construction date values in the parcel data can be filled with the median construction date for surrounding structures.
[0040] As shown in FIG. 3 , the user interface can provide an indication that certain utility data entries are empty and an indication for data entries that have been completed based on other data. For example, the utility data 310 provided by the utility company in FIG. 3 does not identify installation years for pipes with pipe IDs 2 and 6. Therefore, the server can be configured to present the utility data in a manner that indicates that data is missing by highlighting these data fields in red or by providing some other visual indication that data is missing. Furthermore, once missing installation dates are completed via the automated data processing tool, the server can be configured to present the completed data fields with a visual indication that the data consists of completed data. For example, in 312, the completed installation year data for pipes with pipe IDs 2 and 6 is highlighted in yellow, indicating that these dates were not included in the original utility data but were completed. Furthermore, the server can be configured to provide additional information regarding the manner in which certain data has been completed. In particular, the user can be allowed to select data entries provided by the user interface. When a particular data entry is selected, the user interface can provide information indicating where the data originated or how the data was complemented. For example, when a complementing data entry is selected by the user, a pop-up window can appear listing the data on which the complementing data entry was based.
[0041] 4 is a schematic diagram illustrating aspects of automatically filling in pipe installation year data using parcel data, according to some embodiments. Parcel construction years can be used to estimate installation years for utility pipe data. In the example shown in FIG. 4, the median construction year of parcels located within a certain radius around the pipe 410 is used to fill in the installation year for each pipe to generate the complementary data entry 412.
[0042] FIG. 5 is a table showing standard material installation year ranges, according to some embodiments. According to some embodiments, the system not only fills in missing data but also corrects it. FIG. 5 shows the year ranges in which a particular material was used for utility piping. By using standard installation year ranges from the American Water Works Association ("AWWA") or a similar organization, the server can be configured to identify piping of a certain material with an incorrect installation year. For example, the server can identify and correct mismatched 1925 plastic piping to generate a corrective data entry to replace the incorrect data entry. According to some embodiments, other ranges can be used. For example, the ranges can be customized for a particular customer. According to some embodiments, different ranges can be used for different customers.
[0043] 6 is a diagram illustrating the supplementation of pipe material information based on installation year and water utility industry information, according to some embodiments. The data processing system has the ability to correct and supplement pipe material data based on the utility's own data and data collected from other sources. This allows for deeper analysis and allows utilities to better understand their own data.
[0044] Material information can be supplemented using installation year data from either the utility's own data or from third-party or public data, including data collected autonomously by the server. Installation year ranges from AWWA (or similar organization) baseline information provided by the utility, such as material acronyms, and previous material acronyms collected previously, can be used to successfully complement material values. For example, the server can be configured to identify installation year ranges based on the identified material of the pipe, as shown in Figure 5.
[0045] In some cases, it is possible to use a general category to classify materials such as cast iron or plastic, rather than a specific name such as Cement Lined Cast Iron or High Density Polyethylene. This can be done to ensure that less common materials can also be classified and used for analysis in relation to the general category.
[0046] If there is a discrepancy between installation year and material, the system disclosed herein can correct for this based on information from AWWA (or a similar organization), e.g., 1990 cast iron piping is changed to ductile iron.
[0047] 7 is a schematic diagram illustrating the automatic filling of pipe diameter data based on other pipe characteristics, according to some embodiments. Pipe dimensions, including pipe diameter, are important information to have for performing various types of analysis. Therefore, the disclosed system can fill in any missing diameter information or correct inaccurate diameter information.
[0048] If a pipe or piping is missing diameter information, the server of the disclosed system can be configured to automatically fill in this data using other piping attributes. This can be done by the server accessing a database of information associating various piping characteristics with specific diameters or specific ranges of diameters. For example, a 1940s cast iron piping that is missing a diameter entry can be filled in with a value typical of other 1940s cast iron piping.
[0049] If there is a pipe with a missing diameter for a particular material, the server can complete the diameter of the pipe as the most common diameter for pipes used by a particular water utility company or the most common diameter for nearby pipes. For example, if the most common diameter for pipes used by multiple water utility companies is 8 inches, the diameter entry for a water utility company can be completed with 8 inches.
[0050] According to some embodiments described above, data collected from a customer utility company, data from AWWA (or a similar organization), and other sources are collected to supplement and correct pipe data based on standard rules. According to some further embodiments, machine learning can also be used to supplement and correct pipe material, diameter, and installation year data. According to some embodiments, a machine learning algorithm is built solely from the utility company's data. For example, if a utility is missing 30% of its material data, the remainder can be supplemented by finding correlations between material and other pipe attributes, such as installation year or diameter. Installation year and diameter data can be supplemented and corrected following a similar process.
[0051] According to some other embodiments, a machine learning algorithm is built from utility company data and other public data, such as population or zoning. If 30% of a utility's substance data is missing, the remainder can be filled in with a machine learning model that uses other pipe attributes related to the public data. For example, pipe substance data can be filled in with a machine learning model that finds correlations between pipes containing substance values and other public data, such as population or zoning.
[0052] Using these automated methods, pipe leak prediction and job planning can be provided to utilities with large amounts of missing data. Gaps in a utility's own data are not as significant because their internal databases of public and other utility data may be sufficient to fill the gaps.
[0053] Utilities typically keep records of pipe damage. However, not all utilities record this data in a manner that allows it to be easily accessed or displayed, such as by displaying the damage data on a map. When processing the damage data, the server can be configured to geolocate the utility damage data and place it on a map.
[0054] FIG. 8 is a schematic diagram illustrating geocoding of damage data, according to some embodiments. When recording damage data, most utilities record the street address of the damage. Using this information, the server associates every damage with a specific latitude and longitude. According to some embodiments, a geocoding API, such as Google Maps API, is used to localize every damage to an address provided by the utility company. Furthermore, the damage data can be analyzed to remove damages that are unlikely to be located. This can be achieved by creating a buffer zone around the pipe network and removing damages that fall outside the boundaries of this buffer zone. For example, the buffer zone can be defined as all pipes within 300 meters of an existing pipe. Furthermore, if a utility's damage data does not include location information but at least includes the ID of the pipe where the damage occurred, the damage can be geolocated to the center of the pipe. As illustrated in FIG. 8, damage data 810 is provided by the utility company in the form of a pipe ID, damage date, and street address. This information is then converted by the processor into geographically located damage data 812, which can be displayed on a map that can be provided to the user as part of the user interface. A point or circle displayed in the map can represent where the damage occurred. This location can be directly over the geographic location of a street address or over the geographic location of the pipe closest to the street address. According to some embodiments, the system also allows the user to select a particular point, which will display additional information about the particular damage to the user. For example, the map can show a thumbtack or other indicator that the damage occurred at a particular pipe location.Selecting the displayed indicator allows the user to provide information regarding the date the damage occurred.
[0055] FIG. 9 is a schematic diagram illustrating the merging of pipe data and damage data, according to some embodiments. Connecting damage data to the pipe where it originated can provide useful insight into what type of material and vintage are responsible for most damage. Unfortunately, many utilities do not record this information. However, the automated data processing tool can do this. Using geospatial methods, the system can match the geocoded damage to the nearest pipe. If the damage data originally contains information about pipe attributes such as installation year, diameter, material, etc., and one or more of the attributes do not match those of the nearest pipe, a second pipe is selected. This process can be repeated until the distance between the damage point and the center of the pipe segment reaches a certain threshold or until no pipe can be selected. If no pipe is selected, the damage point is assumed to have occurred on an abandoned pipe. The system then attaches pipe data, such as pipe ID, material, and age, to the damage. Further analysis can then be performed on the damage to predict the likelihood of future damage on the same or other pipes. Furthermore, a maintenance and replacement plan can be generated and provided to the utility company based on the analyzed damage data. As shown in Figure 9, pipe data and damage data can be combined so that damage is associated with a particular pipe.
[0056] 10 is a schematic diagram showing data received from an exemplary customer. In this exemplary case, ACME Water is a utility that has poor quality pipe and damage data and is interested in using the system disclosed herein to process its pipe and damage data. First, ACME Water can access a web portal and use the web portal's user interface to upload its pipe data 1010 and damage data 1012. For example, ACME can upload a pipe shapefile, a damage CSV, and provide field names, material acronyms, and a list of materials used 1014.
[0057] This data can be received by a back-end system and reviewed by an operator, as shown in Figure 2. The operator can notice that the piping data does not have installation year data and select the appropriate parcel data from a database. Alternatively, the back-end management system can automatically identify missing data entries and locate relevant data, such as parcel data, available online or in one or more databases, without any operator input. The piping data can then be processed by the data processing tool to clean the data, including reformatting the data and completing missing or inaccurate data.
[0058] FIG. 11 is a schematic diagram illustrating the process of cleaning piping data for an example case customer. The backend system can convert the received utility file into a standard format. This may include converting specific identifiers to pre-specified acronyms and converting dates into a uniform format by providing all dates as four-digit years or converting dates with non-numeric months to a numeric year, month, and day format. Using the parcel data, the backend system can fill in the installation year and other missing data for the piping, such as material and diameter. The backend system can also review the material data and change acronyms to specific naming conventions. In one aspect, the system can use installation year data to fill in missing material data and correct any inconsistent values. Additionally, the system can review diameter data to fill in and correct values based on material and installation year data.
[0059] Figure 12 is a schematic diagram illustrating the process of cleaning damage data for an example case customer. The damage data may be received as a CSV that may or may not be geocoded. If addresses for all damages are recorded, the system can use the addresses to geocode all damages by associating a latitude and longitude with each point. Once this is done, the system can use the cleaned pipe data to attach the damage to the nearest pipe. The system can then project the damage data into a standard coordinate reference system.
[0060] Once both the pipe and damage data has been processed by the system disclosed herein, the processed data can be returned to ACME Water for its own use. The data can be accessed via a downloadable file. Additionally, a user interface can be provided in the web portal, where the processed data can be viewed in tabular output format and in the form of a displayed map. For example, the location of the pipe and damage data can be overlaid on a map and displayed to the user via the web portal or via a downloadable file. Additionally, the disclosed system can generate descriptive statistics and perform further analysis on this newly processed data. In particular, the processed data can be analyzed using machine learning processes described in a co-pending patent application. As described in the co-pending patent application, utility data can be analyzed to create various predictive models capable of predicting future pipe damage. In accordance with the present disclosure, the processed data described herein can be used in one or more End The processed data can be used to create a predictive model for future pipe damage. In this manner, predictions of future pipe damage can be based on cleaned data that includes imputed data values due to missing or inaccurate data to provide a more accurate predictive model. Furthermore, the processed data can be used in job planning algorithms to generate improved pipe maintenance and / or replacement plans. For example, the processed data can be used to identify particular pipes that have a likelihood of failure above a certain probability threshold, or for which maintenance and / or replacement would provide a certain amount of projected savings.
[0061] Further details regarding geoprocessing of data are provided below. Before data can be geoprocessed, it must first be collected. Data can come from a variety of sources, such as public organizations like the U.S. Department of Agriculture or private sources like Open Street Map. This data is then assembled and stored.
[0062] Next, geoprocessing is performed to convert this data into a format that can be used for data analysis. The collected raster data is merged and vectorized. The collected vector data is also merged or split into separate parts to reduce computational costs. The data is cleaned to impute missing values and remove outliers. Furthermore, depending on the situation, the data resolution can be improved for more granular results or generalized to reduce computational costs. Additional variables useful for analysis are then generated. The data is then re-categorised and renamed. Finally, this national data is linked with other data that can be used for analysis.
[0063] After completing these steps, a unified set of data containing nationwide environmental data along with target features is available for analysis.
[0064] Environmental data, such as elevation, is provided in raster format. Figures 13A and 13B are schematic diagrams illustrating potential problems with raster data. There is no single file for elevation across the entire United States. Instead, this elevation data is split into multiple raster datasets representing different areas of the United States. Merging these datasets together can be complicated because there is no guarantee that the rasters will not overlap (e.g., Figure 13A). Furthermore, the top left point of one raster may start in a different location than another raster, meaning that the gap between two rasters (e.g., Figure 13B) may not necessarily be N pixels, where N is an integer.
[0065] FIG. 14 is a block diagram illustrating the process of combining raster data according to some embodiments. In block 1410, an upper-left point and a lower-right point covering the entire United States are selected. In block 1412, the pixel resolution of the raster data is converted, if necessary (this step can be interchanged with block 1414). Changing raster resolution is described in more detail below. In block 1414, each raster data set is expanded to the upper-left point and the lower-right point (this step can be interchanged with block 1412). Pixels in the expanded portion can be filled with an appropriate value, such as 0. This fill value depends on the data being processed. If the raster is expanded to V × integer, where V is a floating-point number, the nearest value or the average over a certain area is used. In block 1416, all raster data is overlaid using the max / min / average / median value of each pixel and using that value as the resulting pixel value. If it is not possible to load all the raster files into memory, a group of rasters is created and blocks 1410, 1412, 1416, and 1418 are performed recursively. In block 1418, the resulting raster data is saved on disk or in a GeoDB. For example, this geoprocessing step is used for elevation data, slope data, soil density (by county / state), etc.
[0066] FIG. 15 is a schematic diagram illustrating vectorization of raster data according to some embodiments. To use nationwide data for analysis, vector data can be used. Vector data is useful because it is flexible for use in many types of geospatial analysis. To convert raster data to vector data, a process known as vectorization is performed. When raster data is vectorized, pixels are converted into multiple square polygons. Vectorization is often performed because spatial joins are commonly used to perform geographic analysis. Vector data is often the optimal data format for that analysis.
[0067] However, one problem is that the vectorized raster data will be significantly larger in size than the original raster, because all vector object vertices must be stored explicitly. When vectorizing large raster data, the memory size of the resulting vector data must be taken into consideration.
[0068] A typical desktop PC (e.g., one with 32-64 GB of memory) would not be able to accommodate all of the resulting data when vectorizing raster data on a national scale (the resulting data size can be 100 GB-1 TB or more).
[0069] To solve this problem, the raster data can be split into multiple rasters and each individual raster can be vectorized. All of the split rasters can be multi-processed and the resulting raster data can be stored on disk or in a Geo database. When storing these rasters, take care to avoid overwriting previously saved files and give each file a unique name.
[0070] FIG. 16 is a schematic diagram illustrating the splitting of raster data in preparation for vectorization, according to some embodiments. First, the maximum dimensions, i.e., the x and y dimensions, are set for the split raster data. For example, if the original raster data is 1,000,000 x 1,000,000 pixels, the raster is split based on the x and y dimensions. If the x dimension = 1,000 and the y dimension = 1,000, the raster is split into 1,000,000 (= 1,000 x 1,000) rasters. Next, the split rasters are individually vectorized with the multi-process option. Finally, the resulting vectors are stored on disk or in the GeoDB. The resulting vectors can be merged programmatically, as described below, or by inserting each vector into the same table in the GeoDB. This geoprocessing step can be used to vectorize national elevation raster data, national gradient raster data, national climate raster data, etc.
[0071] As was the case with raster data, it is sometimes preferable to combine multiple large vector datasets to obtain a single vector dataset. This can be done to reduce the number of resulting files that need to be combined. For example, data collected by county may contain 1,000 files and data collected by state may contain 50 files, while national data may contain only one file.
[0072] However, some vector data formats have an upper limit on the number of features that can be combined. According to some embodiments, this limit can be circumvented by performing the following steps: First, load each vector by converting its CRS to EPSG:4326. Then, convert all files to CSV with WKT geometry. Then, vertically concatenate the CSV files as shown in Figure 17. Then, convert the CSV files back to the original vector format or other format. Then, add the CRS (EPSG:4326) to the vector data. Finally, save the resulting vector files on disk or in a Geo database. This geoprocessing step is used to obtain a single file for the national zoning data and 50 files for the state population data.
[0073] Environmental data is usually provided as vector data. Some of these datasets keep geometry and attributes separate, so they should be combined. In both of these cases, the attributes of the vector data may sometimes have missing values.
[0074] FIG. 18 is a flowchart illustrating a hierarchy for imputing missing attributes in environmental data, according to some embodiments. Missing data can be filled in through the use of other environmental data. For example, a predictive model can be built with environmental data, such as building data and population data (assuming population has some missing values). Population can be provided in vector format and its shape can be block data. To build a predictive model, a model is built using the first centroids of multiple blocks of population data, using environmental data (such as building counts) near the centroids and available population values. The resulting model can then be used to predict missing population values based on environmental data near the centroids of the blocks whose population values are missing.
[0075] For example, this geoprocessing step can be used for soils, population data, climate, elevation, slope, etc. Furthermore, these same methods can be applied to data that is not necessarily environmental but is still useful for analysis, such as supplementing installation year, diameter, material, etc. for pipe data.
[0076] Environmental data retrieved from government sources may contain outliers, such as a pH value of 17, a temperature value of 1,000, or a gradient value of 120. These values are all clearly outliers. To correct for this, these values are programmatically removed.
[0077] The thresholds for these values are based on subject matter (physics, chemistry, etc.) knowledge. For example, knowing that the threshold for plausible pH values is between 0 and 14. These thresholds can also be created based on the following process: (1) calculate the mean and standard deviation of the attribute; (2) set the upper bound to the mean + N*standard deviation (where N is an integer); (3) set the lower bound to the mean - N*standard deviation (where N is an integer); (4) if the value is greater than the upper bound, it is replaced or interpolated with the upper bound (described in more detail below); and (5) if the value is less than the lower bound, it is replaced or interpolated with the lower bound (described in more detail below). For example, this geoprocessing step can be used for soil, population, climate, elevation, slope, etc.
[0078] To build an effective analysis, high-resolution data is required. However, sometimes high-resolution data is not available. For example, climate raster data is sometimes low-resolution, with a single pixel dimension of 400m x 400m. To distinguish between environmental factors, it is possible to generate higher-resolution data from the original lower-resolution data. To do so, it is possible to (1) split the raster pixels and (2) apply a smoothing filter. For example, this geoprocessing step can be used for climate, elevation, slope, etc.
[0079] 19A-19C are schematic diagrams showing examples of generating new environmental variables, according to some embodiments. Sometimes, more data is needed to develop a better analysis. It is possible to create new types of environmental data based on other available environmental data. For example, density information can be constructed from soil data and population data, resulting in soil density and population density. In FIG. 19A, soil density is constructed from geographic information related to soil type and represents the amount of different types of soil in an area. In FIG. 19B, population density is constructed from population values in an area and represents the density of the population in an area. In FIG. 19C, slope data can be generated from elevation. It is also possible to construct proximity of target features to other features. For example, proximity of different types of line data can be generated, such as roads (major roads, highways, etc.), railways (regular rail, light rail, subway, intercity rail, etc.), linear water (rivers, streams, canals, drainage, etc.), and coastlines. Furthermore, proximity of non-environmental data, such as damage, can also be generated.
[0080] FIG. 20 is a block diagram illustrating generating nationwide density data according to some embodiments. In block 2010, nationwide soil data is prepared. In block 2012, an ROI (radius of interest) window is activated. In block 2014, soil vectors in the ROI are selected. In block 2016, soil polygon centroids are selected. In block 2016, the soil centroids are rasterized. In block 2020, a raster heat map is generated from the raster soil centroids. In block 2022, steps 2010, 2012, 2014, 2016, 2018, and 2020 are repeated until the window is activated across the entire country. In block 2024, all raster heat maps are merged (as described below). In block 2026, the merged raster heat map can be vectorized (as described below).
[0081] According to some embodiments, kernel density estimation (KDE) is used to create the heatmap. KDE is a technique for estimating an unknown probability distribution of a random variable based on a sample of points taken from that distribution. Various normal distributions (known as kernel shapes) are used to estimate the values of the unknown points. The distance between the known and unknown points is used as a parameter. When KDE is applied, the density values of the unknown points decrease smoothly according to a Gaussian probability density function.
[0082] Assuming the population data is in vector format and the total population is assigned to each polygon, (1) population data is prepared (by county, state, etc.), missing population values are imputed or filled with zeros, and outliers are removed; (2) the area of each polygon is calculated; (3) the total population is divided by the calculated area of the corresponding polygon to generate the population density, and if the area is 0, the population density is also 0; and (4) a new attribute field representing population density is added to the population vector data, or new vector data is generated that has the same polygons as the population vector data and has attributes corresponding to population density.
[0083] Assuming the slope data is a raster file, national slope data can be generated as follows: (1) prepare national raster elevation data (as described above in Merging Raster Files); (2) calculate the slope using the commonly used equation below; (3) smooth the resulting raster data; and (4) vectorize the resulting raster data (as described above). The following equation can be used to calculate the slope, which refers to the raster matrix in Figure 21:
[0084]
number
[0085] The above method can be used to generate different types of elevation data, such as mean elevation, median elevation, maximum elevation, minimum elevation, elevation standard deviation, etc. The above steps can be applied to all types of elevation data, and all of the resulting slope data can be integrated according to the process described below in connection with integrating gridded vectors. Furthermore, all elevation and resulting slope data can be integrated in the same or similar manner.
[0086] Figure 22 is a block diagram illustrating the creation of new variables from line data, according to some embodiments. In block 2210, if the line data is in vector format, the line vectors can be coalesced as described above. In block 2212, the line vector data is rasterized (if the data is in vector format). In block 2214, a proximity map is generated (e.g., using a popular geoprocessing library such as OGR) and a proximity raster is received. Then, in block 2216, the resulting proximity raster is vectorized. If the resulting raster is very large, it is vectorized according to the process described in section 3-2.
[0087] To use point data (such as for bus stop information or pipe damage data), the following steps can be performed: (1) rasterize the points; (2) generate a raster heat map from the point raster; and (3) vectorize the raster heat map (as described above). This geoprocessing step is useful for incorporating either proximity or density information, such as for bus stop and bus station proximity or generating damage density from damage point data.
[0088] FIG. 23 is a schematic diagram illustrating an example of recategorizing zoning or zoning data, according to some embodiments. When using zoning data, sometimes the received data may be too specific to use for analysis. For example, some zoning data, such as national park areas, may be very unusual in the area found to be used for data analysis. Similar types of zoning data can be grouped for improved analysis. For example, retail and commercial areas can be recategorized as commercial areas; industrial, quarry, and military areas can be recategorized as industrial areas (as shown in the example of FIG. 23); forest, pasture, shrub, weed, wilderness, and national park areas can be recategorized as natural areas; and agricultural, park, orchard, and vineyard areas can be recategorized as artificial natural areas.
[0089] An important benefit of generating geoprocessed environmental data is its use in data analysis. For example, it can be used to predict the likelihood of a water main break. For LOF prediction, a machine learning model is built, which uses pipe, damage, and environmental data as input variables. The environment of each pipe needs to be represented using a spatial join. However, the computational cost of spatial joins and calculating which polygons are overlapped by another is very high. If it is necessary to check polygons that are overlapped by pipes multiple times, the total computational cost increases dramatically.
[0090] However, by optimizing the vector data structure, this process can be sped up. This method can be used for square-shaped environmental data and all environmental data that can be approximated as square-shaped environmental data. For example, soil data can be arbitrarily shaped, but the soil data can be approximated as square-shaped soil data by filling the soil data area with small tiles.
[0091] 24A and 24B are schematic diagrams illustrating an example of approximating square-shaped data, according to some embodiments. In the illustrated example, in FIG. 24A, the soil has three properties (P1, P2, P3). The data can be converted into three rasters, with a pixel corresponding to each property, or can be directly converted into a single square-shaped vector by performing spatial joins. This is useful for soil, climate, elevation, slope, proximity, heat maps (soil density), etc. Originally, this square-shaped environmental data may be in raster format. For example, if a rail, road, and station proximity map is used, these three raster data files need to be vectorized and then spatially joined three times.
[0092] Figure 25 is a schematic diagram illustrating an example of merging grid-shaped vectors, according to some embodiments. By implementing the illustrated process, it is possible to reduce the number of spatial joins performed. At block 2510, rasters with the same pixel resolution are expanded to the same dimensions and the same top-left point. At block 2512, a grid vector is created that completely overlaps the original raster data (starting from the same top-left geographic point, and each grid dimension is the same as the pixel dimensions of the raster data). At block 2514, pixel values from each raster are loaded and assigned to each grid. At block 2516, the resulting vector is obtained on disk or in a GeoDB. In some cases, the resulting vector can be spatially joined to some other vector.
[0093] The above example outputs the resulting lattice vector, which contains the three attributes. By using this process, the resulting lattice vector contains all the information contained in the multiple rasters.
[0094] Note that spatial joins may not be necessary when assigning pixel values from raster data to grid vector data, because the resulting grid vector will completely overlap the raster data. Values can be assigned based on pixel and grid index. The resulting vector data can then be used to spatially join other data, such as piping data.
[0095] Without the illustrated technique, three spatial joins would be required, as in the example above. However, by using this technique, only one spatial join is required.
[0096] It should be noted that to improve its usefulness, national data should be easily understandable. However, attributes of vector data can have many different names. For example, soil attributes can be pH, CaCO3, and other names that may not be as obvious. To clarify their use, the following naming conventions can be used: (1) the addition of a prefix such as "pssn," where the first two letters represent "public" and "standard" and the last two letters represent "soil" and "numeric property"; (2) the addition of numbering and a postfix. For example, pH and CaCO3 are converted to pssn00 and pssn01 (with prefix and numbering). Total population and population density are converted to pspp00 and pspp02 (with prefix and numbering).
[0097] When performing spatial joins of environmental data to target features, LOF analysis generates very large tables that, without these naming conventions, contain the original attribute names as column names, making it difficult to sort, filter, and identify features.
[0098] Figure 26 is a schematic diagram illustrating a geodatabase system architecture according to some embodiments. The system includes an application interface portion and a database portion. The application interface portion includes the ability to query geospatial data and the ability to import geospatial data. The database portion includes a single database containing geospatial data and multiple databases containing geospatial data.
[0099] Although the above-described embodiments primarily relate to data relating to networks of underground pipes, according to some embodiments, many of the techniques described above may be applied to data relating to other types of networks. According to some embodiments, the systems and methods described herein are applied to data wrangling and / or environmental data relating to networks of above-ground utility poles and / or electrical wires used to deliver power to customers, such as between underground nodes.
[0100] According to some further embodiments, utility poles themselves can be treated as target assets, rather than or in addition to power lines. For example, utility poles can be treated as target assets for data wrangling. In such cases, the techniques described herein can be applied to wrangling data related to power lines and to wrangling data related to utility poles. Examples of utility pole data include, for example, pole diameter, pole material, pole installation year, etc.
[0101] Although some details have been set forth above for purposes of clarity, it will be apparent that certain changes and modifications can be made without departing from the principles of the present invention. It should be noted that there are many alternative ways of implementing both the processes and apparatus described herein. Therefore, the present examples should be considered as illustrative and not restrictive, and the invention described herein should not be limited to the details set forth herein, but can be modified within the scope of the appended claims and their equivalents.
Claims
1. A method for improving utility company piping data relating to a network of underground piping for transporting fluids to consumers by a computer system including a processor, comprising: receiving, by the computer system, pipe data for one or more of the networks of underground pipes, the pipe data including at least pipe ID, pipe material, pipe size, and installation year as pipe attributes; the computer system identifying, based on the piping data, one or more piping in which at least one of the piping attributes is missing or inaccurate; and the computer system automatically generating complementary data for the at least one of the missing or inaccurate pipe attributes, thereby forming the improved set of pipe data for a network of underground pipes for transporting fluids to consumers; Imputing missing installation year data for a pipe includes calculating the year based on installation years of geographically nearby pipes included in the pipe data, and / or based on parcel data related to the location of the pipe and including construction years of nearby buildings, and / or based on a standard material installation year range; Imputing missing pipe material data includes determining materials based on installation year and water utility industry information; Completing the missing dimensional data of the pipe includes determining the diameter of the pipe based on the diameter or other attributes of the pipe, including the material; A method that encompasses the above.
2. The method described in claim 1, wherein the step of automatically generating complementary data uses one or more information sources selected from the group consisting of GPS, address, and geocode to associate the piping data with the location address of the damaged piping.
3. The method of claim 2, wherein the piping data includes piping damage data, at least some of which includes a street address for the piping damage rather than a specific piping section for the damage.
4. The method of claim 1, wherein the step of automatically generating complementary data includes using machine learning to correct for and / or complete at least some of the missing or inaccurate piping attributes.
5. The method of claim 1, further comprising a step of predicting the likelihood of damage to one or more piping sections based on the piping data.
6. 6. The method of claim 5, wherein the predicting step is based at least in part on a predictive model constructed by machine learning using at least some of the piping data.
7. A computer system for improving utility company piping data relating to a network of underground piping for transporting fluids to consumers, comprising: means for receiving and storing pipe data for one or more of said networks of underground pipes, the pipe data including at least pipe ID, pipe material, pipe size, and installation year as pipe attributes; means for identifying, based on the piping data, one or more piping for which at least one of the piping attributes is missing or inaccurate; and means for automatically generating complementary data for said at least one of said pipe attributes that is missing or inaccurate, thereby forming said improved set of pipe data for a network of underground pipes for transporting fluids to consumers; generating complementary data for the missing installation year of the pipe by calculating the year based on the installation years of geographically nearby pipes included in the pipe data, and / or based on parcel data related to the location of the pipe and including the construction years of nearby buildings, and / or based on the installation year range of standard materials; generating missing material completion data for the pipe by determining the pipe material based on installation year and water utility industry information; means configured to generate complementary data for missing dimensions of the pipe by determining the diameter of the pipe based on other attributes of the pipe, including the diameter or material of the pipe; A system that encompasses.
8. The system described in claim 7, further comprising a front-end system including an uploader for receiving the piping data from a customer, and a viewer for displaying at least some of the improved set of piping data.
9. 8. The system of claim 7, further comprising means for predicting the likelihood of a leak in a pipe segment in the network.
10. 10. The system of claim 9, wherein the means for predicting comprises a predictive model constructed by machine learning using at least some of the refined set of piping data.
Citation Information
Patent Citations
Underground pipeline data synchronization method and underground pipeline data synchronization device
CN105069694A
Pipe damage probability prediction method based on BP neural network
CN106022518A
Method and equipment for predicting damage to sewer pipe of sewer pipe network
JP2006183274A
System and method for filling gaps of missing data using source specified data
US6862540B1
Detecting small leaks in pipeline network
US9395262B1