Urban drainage pipeline defect type prediction method based on machine learning
Through a machine learning-based approach, using support vector machines and random forest classification models, the problem of low efficiency in urban drainage pipeline defect detection was solved, efficient and accurate defect type prediction was achieved, reducing costs and improving the operating efficiency of the urban drainage system.
Patent Information
- Application Number
- CN202510709641.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-16
AI Technical Summary
Existing urban drainage pipe defect detection methods are inefficient, costly, and time-consuming, making it difficult to identify defect types efficiently and accurately.
A machine learning-based method is used to establish support vector machine and random forest classification models through data collection, preprocessing, feature encoding and model training to predict whether there are defects in urban drainage pipelines and their types.
It improves the efficiency and accuracy of defect detection, reduces material and manpower costs, can quickly identify pipeline defects and their types in a short period of time, and improves the reliability and operational efficiency of urban drainage systems.
Smart Images

Figure CN120654134A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of urban drainage pipeline defect type prediction, and in particular to a method for predicting urban drainage pipeline defect types based on machine learning. Background Art
[0002] Urban drainage pipes are like the blood vessels of a city, closely linked to its public health, safety, and economic well-being. Clogged drainage pipes can lead to the entire city's collapse. As China's urbanization continues to expand, the scale of urban drainage pipes is also increasing. Large-scale pipeline expansion increases the time and manpower required to effectively inspect all pipes. Therefore, efficient and accurate drainage pipe inspection is a crucial component of urban construction. Current methods for detecting drainage pipe defects include endoscopic sonar, periscopes, and closed-circuit television (CCTV). CCTV inspection systems are the most widely used on construction sites. While CCTV inspection provides an effective method for detecting pipe defects, it requires specialized pipeline robots, consumes considerable time and effort to analyze the inspection results, and is typically limited to a small portion of the entire urban drainage system. Consequently, currently used inspection methods still suffer from low efficiency, high costs, and time-consuming processes. Therefore, utilizing modern technologies to improve the efficiency and accuracy of urban drainage pipe defect detection has become a pressing issue. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a method for predicting the defect type of urban drainage pipes based on machine learning, which can effectively improve the efficiency and accuracy of defect detection in urban drainage pipes.
[0004] The technical solution adopted by the present invention to solve the above technical problems is:
[0005] A method for predicting defect types in urban drainage pipes based on machine learning, comprising:
[0006] Step 1: Data collection:
[0007] (1-1) Using pipeline detection instruments, we obtain a characteristic dataset of urban drainage pipelines, denoted as S, where S = {(x0, y0), (x1, y1), (x2, y2) ... (x m ,y m )},in:
[0008] x represents the sample attributes, including ground type (DM), surrounding disturbance type (RD), ground load type (ZH), longitude (JD), latitude (WD), pipe age (GL), pipe diameter (GJ), pipe length (GC), pipe material type (GH), depth (SD), slope (PD), and whether rainwater and sewage are mixed (YW);
[0009] y represents the category, including normal category and defect category;
[0010] m=0,1,2……;
[0011] (1-2) The above defect categories are denoted as y', including rupture denoted as PL, deformation denoted as BX, dislocation denoted as CK, corrosion denoted as FS, leakage denoted as SL, disconnection denoted as TJ, deposition denoted as CJ, scaling denoted as JG, collapse denoted as TT, and obstacle denoted as XW; further forming a data subset of defect categories, denoted as S', S' = {(x 0, y'0),(x 1, y'1),(x 2, y'2)......(x m, y' m )};
[0012] Step 2: Data preprocessing:
[0013] Clean the above dataset S and data subset S' and delete missing data and data with outliers;
[0014] Step 3: Feature encoding:
[0015] (3-1) Label encoding is performed on the cleaned dataset S and the data subset S', where the "defect category" and "normal category" in the dataset S are mapped to 0 and 1 respectively;
[0016] (3-2) Map the defect categories y' in the data subset S' to 0, 1, 2, 3, 4, 5, 6, 7, 8, and 9 respectively;
[0017] (3-3) One-hot encoding is performed on DM, RD, ZH, GM and YW in the sample attributes x in the cleaned dataset S and the data subset S';
[0018] Step 4: Model training:
[0019] (4-1) Support Vector Machine Classification Model (SVM) Training:
[0020] In the support vector machine classification model training, the core hyperparameters are optimized through grid search using the GridSearchCV class of Scikit-learn. Specifically, they include: the regularization coefficient C, with candidate values of 0.1, 1, 10, and 100; controlling the generalization ability of the model; the kernel function parameter gamma, with candidate values of 0.001, 0.01, 0.1, and 1; adjusting the nonlinear mapping of the RBF kernel; the kernel function type kernel, with candidate values of linear / 'rbf'; and determining the complexity of the classification boundary.
[0021] The grid search optimization process evaluates the prediction accuracy of the above different parameter combinations on the validation set through five-fold cross-validation, and finally finds the optimal hyperparameter combination, which is:
[0022] The dataset S obtained in step 3 above is divided into a training set and a validation set in a ratio of 4:1. Four training sets are randomly selected to train the support vector machine classification model. The remaining training set data is used as the validation set to verify the accuracy of the trained support vector machine classification model. This cycle is repeated five times to obtain five prediction accuracy rates, and then the arithmetic mean is calculated. The above process is performed once for each grid search, and each time an arithmetic mean accuracy rate is reached, the model hyperparameter combination with the highest arithmetic mean accuracy rate is finally selected as the hyperparameter combination with the best prediction performance of the support vector machine model on this dataset.
[0023] The support vector machine classification model is configured using the hyperparameter combination with the best performance, thereby obtaining a trained support vector machine classification model;
[0024] (4-2) Random Forest Classification Model (RandomForest) Training:
[0025] In the training of the random forest classification model, the core hyperparameters are optimized through grid search using the GridSearchCV class of Scikit-learn, including: n_estimators, with candidate values of 5, 100, and 200; it controls the number of decision trees in the forest and affects the ensemble learning ability of the model; max_depth, with candidate values of 5, 10, and None; it limits the maximum depth of a single tree and prevents overfitting; max_features, with candidate values of "auto", "sqrt", and "log2"; it determines the feature sampling ratio when splitting nodes; min_samples_split, with candidate values of 1, 2, 5, and 10; min_samples_leaf, with candidate values of 1, 2, and 4; it sets the minimum number of samples for leaf nodes and improves generalization; class_weight, with candidate values of None and "balanced"; it adjusts the weight distribution when the class samples are unbalanced.
[0026] The grid search optimization process evaluates the prediction accuracy of the above different parameter combinations on the validation set through five-fold cross-validation, and finally finds the optimal hyperparameter combination, which is:
[0027] The data subset S' obtained in step 3 above is divided into a training subset and a validation subset in a ratio of 4:1. Four training subsets are randomly selected to train the random forest classification model. The remaining data is used as the validation subset to verify the accuracy of the trained random forest classification model. This cycle is repeated five times to obtain five prediction accuracy rates, and then the arithmetic mean is calculated. The above process is performed once for each grid search, and each time an arithmetic mean accuracy rate is reached, the model hyperparameter combination with the highest arithmetic mean accuracy rate is finally selected as the hyperparameter combination with the best prediction performance of the random forest classification model on this data subset.
[0028] Use the hyperparameters with the best performance to configure the random forest classification model, thus obtaining the trained random forest classification model.
[0029] Step 5: Defect prediction:
[0030] The data to be predicted is brought into the support vector machine classification model trained in step 4 above to predict whether there is a defect in the urban drainage pipe section. If the prediction result is "defect", the predicted data is further brought into the random forest classification model trained in step 4 to continue predicting the type of the "defect" and finally obtain the defect type prediction result.
[0031] The specific operation of the above step 2 is: use the Python-based data analysis and processing library Pandas to clean the above dataset S and data subset S', wherein the dropna method in the Pandas library is used to delete missing data, and the IsolationForest algorithm in the Pandas library is used to delete data with outliers.
[0032] After step 2 preprocessing, the dataset S and the data subset S' are organized into two-dimensional tables and saved as CSV files for subsequent data reading and processing.
[0033] The defect category y' in the data subset S' is mapped to 0, 1, 2, 3, 4, 5, 6, 7, 8, and 9 respectively, where 0 corresponds to the defect category PL, 1 corresponds to the defect category BX, 2 corresponds to the defect category CK, 3 corresponds to the defect category FS, 4 corresponds to the defect category SL, 5 corresponds to the defect category TJ, 6 corresponds to the defect category CJ, 7 corresponds to the defect category JG, 8 corresponds to the defect category TT, and 9 corresponds to the defect category ZW.
[0034] In step 3 above:
[0035] The LabelEncoder in the scikit-learn library is used to encode the labels of the dataset S and the data subset S'.
[0036] The One-Hot Encoder provided by the machine learning library Scikit-learn is used to perform One-Hot encoding on the sample attribute x in the dataset S and the data subset S'.
[0037] Before training the support vector machine classification model, the StandardScaler method is used to standardize the dataset S obtained in step 3. The parameters of the support vector machine classification model are optimized by the grid search method GridSearchCV and the five-fold cross-validation method. Finally, the optimal hyperparameter configuration of the support vector machine classification model (SVM) is selected as follows: the value range of C is 1.0; the value of kernel is 'rbf'; and the value of gamma is 0.1.
[0038] Taking the model prediction accuracy as the optimization indicator, the grid search method GridSearchCV and the five-fold cross-validation method are used to optimize the parameters of the random forest classification model. Finally, the hyperparameter configuration of the random forest classification model with the best performance is selected as follows: n_estimators is 100, max_depth is 10, class_weight is 'balanced', min_samples_split is 5, min_samples_leaf is 2, and max_features is 'auto'.
[0039] The pipeline detection instrument is CCTV.
[0040] Compared with the existing technology, the advantages of the present invention are: by using the urban drainage pipeline inspection data collected by existing means and establishing a data set, and using it to train a machine learning model, it is possible to achieve efficient and highly accurate prediction of whether there are defects in the urban drainage pipeline and the type of defects. This method does not require complex detection equipment and is simple and fast, which can greatly improve the efficiency of urban drainage pipeline defect detection and greatly reduce the material and manpower costs of this work. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 is a flow chart of the present invention;
[0042] Figure 2 It is an implementation flow chart of the present invention. DETAILED DESCRIPTION
[0043] One of the defects is selected below and described in further detail in conjunction with the present invention.
[0044] Step 1: Data collection:
[0045] Using CCTV inspection equipment, a city's drainage pipes were inspected to obtain a pipeline feature dataset. This data consists of two components: categories y and sample attributes x. Category y is divided into normal categories (ZC) and defect categories (QX). Defect categories include 10 types: rupture (PL), deformation (BX), dislocation (CK), corrosion (FS), leakage (SL), disconnection (TJ), deposition (CJ), scaling (JG), collapse (TT), and obstruction (XW). Sample attributes x include 12 types: ground type (DM), surrounding disturbance type (RD), ground load type (ZH), longitude (JD), latitude (WD), pipe age (GL), pipe diameter (GJ), pipe length (GC), pipe material type (GH), depth (SD), slope (PD), and whether there is a mixed connection between rainwater and sewage (YW). Ultimately, data from 1,000 pipeline segments were collected, forming the corresponding dataset S and data subset S'.
[0046] Step 2: Data preprocessing:
[0047] The library used is Pandas, a Python-based data analysis and processing library. The specific method is as follows: First, use the dropna method in the Pandas library to delete data with missing values. For example, in the pipe age (GL) attribute, if certain data rows do not have specific pipe age values filled in, these rows are deleted. Then, use the IsolationForest algorithm in the Pandas library to identify and delete data containing outliers. For example, for slope (PD), if the slope of most pipes is concentrated between 0.01 and 0.1, but a small number of data have a slope of 10 or higher, these are likely outliers and are detected and deleted using the IsolationForest algorithm.
[0048] Data organization and storage format: Organize the preprocessed dataset S and the data subsets S' for each defect category into two-dimensional tables and save them as CSV files for subsequent data reading and processing. For example, the training set used in the following application can be saved as "train_data.csv," the test set as "test_data.csv," and the crack (PL) data subset as "PL_data.csv." Data subsets for other defect categories can also be saved using similar naming conventions.
[0049] Step 3: Feature encoding:
[0050] (1) Label encoding: Label Encoder in the scikit-learn library is used to encode the “defect category” and “normal category” in the dataset S, mapping the normal category (ZC) to 0 and the defect category (QX) to 1; for the data subset S' of the defect category, Label Encoder is also used to map its corresponding defect categories to 0-9, where 0 corresponds to the defect category PL, 1 corresponds to the defect category BX, 2 corresponds to the defect category CK, 3 corresponds to the defect category FS, 4 corresponds to the defect category SL, 5 corresponds to the defect category TJ, 6 corresponds to the defect category CJ, 7 corresponds to the defect category JG, 8 corresponds to the defect category TT, and 9 corresponds to the defect category ZW;
[0051] (2) One-hot encoding: The sample attributes DM, RD, ZH, GM, and YW in the dataset S and the data subset S' are one-hot encoded using the encoder OneHotEncoder provided by Scikit-learn. For example, the ground type (DM) may have several different types, such as soil, cement, asphalt, etc. After one-hot encoding, each type will correspond to a binary feature column. If the ground type of a section of the pipeline is soil, the corresponding soil feature column will take the value of 1, and the columns corresponding to other types will take the value of 0. This can convert categorical variables into numerical features that the model can handle, avoiding the model's misunderstanding of the order relationship between different types.
[0052] Step 4: Model training:
[0053] (1) Data preparation: The collected dataset S is divided into a training set and a test set in a ratio of 4:1. The training set contains 800 data segments and the test set contains 200 data segments.
[0054] (2) Support Vector Machine Classification Model (SVM) Training:
[0055] Data preparation and standardization: First, extract the data used for model training from the training set, including sample attributes and corresponding labels (0 or 1, indicating whether a defect exists). Then, use the StandardScale method to standardize the data, scaling the values of the sample attributes to a similar range. For example, each attribute has a mean of 0 and a standard deviation of 1. This can improve the training efficiency and accuracy of the SVM model.
[0056] Parameter optimization and training: Using the five-fold cross-validation method, the training set is divided into five parts. In each training process, four parts of the data are used for training, and the remaining part of the data is used for verification. Using the grid search method GridSearchCV, the parameters of the SVM model are optimized within the framework of five-fold cross-validation. The value range of the regularization coefficient C is set to 1.0-10.0, the candidate values of the kernel function type kernel are common kernel functions such as 'rbf', and the candidate values of the kernel function parameter gamma are 'scale'. After multiple iterative searches and verifications, the optimal parameters of the SVM model are finally determined to be: the regularization coefficient C is 1.0, the kernel function type kernel is 'rbf', and the kernel function parameter gamma is 0.1. This optimal parameter combination is used to train the SVM model to obtain a trained SVM classification model.
[0057] (3) Random Forest Classification Model (RandomForest) Training
[0058] Data preparation: The data subset S' obtained in step 3 above includes sample attributes and corresponding defect category labels (0, 1..., 9);
[0059] Parameter Optimization and Training: Using the same five-fold cross-validation method, the training subset was divided into five parts. Four parts of the training subset were used for training the random forest classification model during each training session, with the remaining part used for validation. Using model prediction accuracy as the optimization metric, the parameters of the random forest classification model were optimized using the grid search method GridSearchCV and five-fold cross-validation. Candidate parameter ranges included n_estimators (the number of decision trees) between 50 and 200, and max_depth (the maximum depth of the decision tree) between 5 and 15. After searching and validating, the optimal parameters for the random forest classification model were found to be: n_estimators = 100, max_depth = 10, class_weight = 'balanced', min_samples_split = 5, min_samples_leaf = 2, and max_features = 'auto'. This optimal parameter combination was then used to train the random forest classification model, resulting in a trained random forest classification model.
[0060] Step 5: Defect prediction:
[0061] The urban drainage pipeline data to be predicted is first input into the trained support vector machine classification model to predict whether the section of pipeline has defects. If the prediction result is "defect", that is, it belongs to the defect category (QX), the data to be predicted is further input into the trained random forest classification model to further predict which specific defect type the "defect" belongs to, such as rupture (PL), deformation (BX), etc., and finally obtain the defect type prediction result.
[0062] Based on the defect type prediction results obtained above, relevant departments can formulate targeted repair and maintenance plans and promptly repair problematic pipelines, effectively improving the reliability and operational efficiency of urban drainage systems and reducing the risk of accidents and economic losses that may be caused by pipeline defects.
[0063] By comparing the model prediction results with the actual pipeline defect situation (real defect information is obtained through other reliable means such as manual inspection and pipeline robot inspection), and calculating indicators such as prediction accuracy, it is evaluated that the accuracy of this method can reach more than 90%.
[0064] Statistics on the time, manpower and material costs required for defect inspection of urban drainage pipes before and after the adoption of the invented method were collected. For example, traditional manual inspection methods may take several weeks, a large amount of manpower and professional equipment to conduct a comprehensive inspection of pipes in a certain area. However, after adopting this method, through data collection and model prediction, the pipes that may have defects and their defect types can be quickly identified in a relatively short period of time, thereby greatly reducing the workload and time cost of manual inspection and improving overall work efficiency.
[0065] In practical applications, the use of this method can help city managers better maintain and manage urban drainage systems, discover potential pipeline problems in advance, and avoid the adverse effects of accidents such as road collapse and sewage leakage caused by pipeline defects on urban traffic, residents' lives and the environment, which has significant social and economic benefits.
Claims
1. A method for predicting urban drainage pipe defect types based on machine learning, characterized in that include: Step 1: Data collection: (1-1) Using pipeline detection instruments, we obtain a characteristic dataset of urban drainage pipelines, denoted as S, where S = {(x0, y0), (x1, y1), (x2, y2) ... (x m ,y m )},in: x represents the sample attributes, including ground type (DM), surrounding disturbance type (RD), ground load type (ZH), longitude (JD), latitude (WD), pipe age (GL), pipe diameter (GJ), pipe length (GC), pipe material type (GH), depth (SD), slope (PD), and whether rainwater and sewage are mixed (YW); y represents the category, including normal category and defect category; m=0,1,2......; (1-2) The above defect categories are denoted as y', including rupture denoted as PL, deformation denoted as BX, dislocation denoted as CK, corrosion denoted as FS, leakage denoted as SL, disconnection denoted as TJ, deposition denoted as CJ, scaling denoted as JG, collapse denoted as TT, and obstacle denoted as XW; further, a data subset of defect categories is formed, denoted as S', S' = {(x0,y'0), (x1,y'1), (x2,y'2) ... (x m ,y' m )}; Step 2: Data preprocessing: Clean the above dataset S and data subset S' and delete missing data and data with outliers; Step 3: Feature encoding: (3-1) Label encoding is performed on the cleaned dataset S and the data subset S', where the "defect category" and "normal category" in the dataset S are mapped to 0 and 1 respectively; (3-2) Map the defect categories y' in the data subset S' to 0, 1, 2, 3, 4, 5, 6, 7, 8, and 9 respectively; (3-3) One-hot encoding is performed on DM, RD, ZH, GM and YW in the sample attributes x in the cleaned dataset S and the data subset S'; Step 4: Model training: (4-1) Support vector machine classification model training: In the support vector machine classification model training, the core hyperparameters are optimized through grid search using the GridSearchCV class of Scikit-learn. Specifically, they include: the regularization coefficient C, with candidate values of 0.1, 1, 10, and 100; controlling the generalization ability of the model; the kernel function parameter gamma, with candidate values of 0.001, 0.01, 0.1, and 1; adjusting the nonlinear mapping of the RBF kernel; the kernel function type kernel, with candidate values of linear / 'rbf'; and determining the complexity of the classification boundary. The grid search optimization process evaluates the prediction accuracy of the above different parameter combinations on the validation set through five-fold cross-validation, and finally finds the optimal hyperparameter combination, which is: The dataset S obtained in step 3 above is divided into a training set and a validation set in a ratio of 4:
1. Four training sets are randomly selected to train the support vector machine classification model. The remaining training set data is used as the validation set to verify the accuracy of the trained support vector machine classification model. This cycle is repeated five times to obtain five prediction accuracy rates, and then the arithmetic mean is calculated. The above process is performed once for each grid search, and each time an arithmetic mean accuracy rate is reached, the model hyperparameter combination with the highest arithmetic mean accuracy rate is finally selected as the hyperparameter combination with the best prediction performance of the support vector machine model on this dataset. The support vector machine classification model is configured using the hyperparameter combination with the best performance, thereby obtaining a trained support vector machine classification model; (4-2) Random forest classification model training: In the training of the random forest classification model, the core hyperparameters are optimized through grid search using the GridSearchCV class of Scikit-learn, including: n_estimators, with candidate values of 5, 100, and 200; it controls the number of decision trees in the forest and affects the ensemble learning ability of the model; max_depth, with candidate values of 5, 10, and None; it limits the maximum depth of a single tree and prevents overfitting; max_features, with candidate values of "auto", "sqrt", and "log2"; it determines the feature sampling ratio when splitting nodes; min_samples_split, with candidate values of 1, 2, 5, and 10; min_samples_leaf, with candidate values of 1, 2, and 4; it sets the minimum number of samples for leaf nodes and improves generalization; class_weight, with candidate values of None and "balanced"; it adjusts the weight distribution when the class samples are unbalanced. The grid search optimization process evaluates the prediction accuracy of the above different parameter combinations on the validation set through five-fold cross-validation, and finally finds the optimal hyperparameter combination, which is: The data subset S' obtained in step 3 above is divided into a training subset and a validation subset in a ratio of 4:
1. Four training subsets are randomly selected to train the random forest classification model. The remaining data is used as the validation subset to verify the accuracy of the trained random forest classification model. This cycle is repeated five times to obtain five prediction accuracy rates, and then the arithmetic mean is calculated. The above process is performed once for each grid search, and each time an arithmetic mean accuracy rate is reached, the model hyperparameter combination with the highest arithmetic mean accuracy rate is finally selected as the hyperparameter combination with the best prediction performance of the random forest classification model on this data subset. Use the hyperparameters with the best performance to configure the random forest classification model, thus obtaining the trained random forest classification model. Step 5: Defect prediction: The data to be predicted is fed into the support vector machine classification model trained in step 4 above to predict whether there is a defect in this section of urban drainage pipe. If the prediction result is "defect", the predicted data is further fed into the random forest classification model trained in step 4 to continue predicting the type of the "defect". Finally, the defect type prediction result is obtained.
2. The method for predicting urban drainage pipe defect types based on machine learning according to claim 1 is characterized in that The specific operation of the above step 2 is: use the Python-based data analysis and processing library Pandas to clean the above dataset S and data subset S', wherein the dropna method in the Pandas library is used to delete missing data, and the IsolationForest algorithm in the Pandas library is used to delete data with outliers.
3. The method for predicting urban drainage pipe defect types based on machine learning as claimed in claim 1, characterized in that The defect category y' in the data subset S' is mapped to 0, 1, 2, 3, 4, 5, 6, 7, 8, and 9 respectively, where 0 corresponds to the defect category PL, 1 corresponds to the defect category BX, 2 corresponds to the defect category CK, 3 corresponds to the defect category FS, 4 corresponds to the defect category SL, 5 corresponds to the defect category TJ, 6 corresponds to the defect category CJ, 7 corresponds to the defect category JG, 8 corresponds to the defect category TT, and 9 corresponds to the defect category ZW.
4. The method for predicting urban drainage pipe defect types based on machine learning according to claim 1, characterized in that In step 3 above: The LabelEncoder in the scikit-learn library is used to encode the labels of the dataset S and the data subset S'. The One-Hot Encoder provided by the machine learning library Scikit-learn is used to perform One-Hot encoding on the sample attribute x in the dataset S and the data subset S'.
5. The method for predicting urban drainage pipe defect types based on machine learning as claimed in claim 1, characterized in that Before training the support vector machine classification model, the StandardScaler method is used to standardize the dataset S obtained in step 3. The parameters of the support vector machine classification model are optimized by the grid search method GridSearchCV and the five-fold cross-validation method. Finally, the optimal hyperparameter configuration of the support vector machine classification model (SVM) is selected as follows: the value range of C is 1.0; the value of kernel is 'rbf'; and the value of gamma is 0.
1.
6. The method for predicting urban drainage pipe defect types based on machine learning as claimed in claim 1, characterized in that Taking the model prediction accuracy as the optimization indicator, the grid search method GridSearchCV and the five-fold cross-validation method are used to optimize the parameters of the random forest classification model. Finally, the hyperparameter configuration of the random forest classification model with the best performance is selected as follows: n_estimators is 100, max_depth is 10, class_weight is 'balanced', min_samples_split is 5, min_samples_leaf is 2, and max_features is 'auto'.
7. The method for predicting urban drainage pipe defect types based on machine learning as claimed in claim 1, characterized in that The pipeline detection instrument is CCTV.
Citation Information
Cited By
Urban drainage pipeline defect detection method based on CCTV video
CN122067029A