A pipeline defect trenchless repair decision method and system based on reinforcement learning

CN116739375BActive Publication Date: 2026-08-11MCC SOUTHERN CITY CONSTR ENG TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-09
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0004]针对现有技术中的缺陷,本发明提供一种基于强化学习的管道缺陷非开挖修复决策方法及系统,解决传统管段人工检测与评估时效性差以及成本高的问题,具有精准高效、成本低等优点

Benefits of technology

[0053]1、通过采用自适应边缘检测算法实现了大规模声纳轮廓图的精准提取,采用SVM多分类器实现了对大规模管道缺陷等级的初步评估,节省了大量人力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116739375B_ABST
    Figure CN116739375B_ABST
Patent Text Reader

Abstract

This invention discloses a trenchless pipeline defect repair decision-making method and system based on reinforcement learning. The method includes: acquiring a sonar profile map of the pipeline; using an SVM-based multi-classifier-based preliminary defect level assessment model to provide a preliminary assessment result of the pipeline defect level; the pipeline defect level includes structural and functional defect severity levels; screening pipelines based on the preliminary defect level assessment results; acquiring television inspection images of the screened pipelines; using a GNN-SVM-based final defect level assessment model to provide a final assessment result of the pipeline defect level; obtaining the comprehensive pipeline defect severity level; constructing a PPO-based repair strategy model; acquiring and executing the optimal repair strategy; and correcting the two assessment models based on the pipeline defect severity level verified during the pipeline repair process. This invention solves the problems of poor timeliness and high cost of traditional manual pipeline inspection and assessment, and has the advantages of accuracy, efficiency, and low cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of drainage pipeline repair technology, specifically to a trenchless pipeline defect repair decision-making method and system based on reinforcement learning. Background Technology

[0002] With rapid economic development, the scale of urban infrastructure construction is expanding, and the total mileage of urban underground pipe networks is increasing. Consequently, the number of municipal pipelines reaching their service life and requiring urgent repair and replacement each year is also rising. Drainage pipelines built years ago have developed numerous structural defects due to changes in surface load, groundwater erosion, pipe corrosion, and aging pipe materials. Some of these defects have seriously affected the normal operation of the pipelines, even causing damage that endangers road structural safety. Therefore, it is urgent to inspect and repair these defective pipelines to eliminate urban safety hazards and ensure the normal operation of the "city's veins." Trenchless engineering is an environmentally friendly technology (EST) for underground facilities. Trenchless technology for municipal pipelines has advantages such as low overall cost, short construction period, minimal environmental impact, no disruption to traffic or residents' lives, and good construction safety. It is increasingly favored and widely used in the construction of underground pipelines such as municipal drainage pipelines, communication cables, and gas pipelines, yielding significant economic and social benefits.

[0003] Currently, urban pipeline sections are inspected and maintained according to technical specifications. However, this method relies on multiple assessments by a large number of people, resulting in poor timeliness and high labor costs. With the development of artificial intelligence technology, various industries are beginning to use AI methods to overcome traditional technical bottlenecks. Therefore, how to intelligently assess large-scale pipeline defects and quickly determine a technically feasible and economically reasonable trenchless repair method is a problem that the industry urgently needs to solve. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a trenchless pipeline defect repair decision-making method and system based on reinforcement learning, which solves the problems of poor timeliness and high cost of traditional manual pipeline inspection and evaluation, and has the advantages of accuracy, efficiency and low cost.

[0005] The technical solution adopted in this invention is as follows:

[0006] According to a first aspect of the present invention, a trenchless repair decision-making method for pipeline defects based on reinforcement learning is provided, comprising the following steps:

[0007] S100: Obtain the sonar profile of the pipeline and use the SVM-based multi-classifier-based preliminary defect level assessment model to give the preliminary assessment results of the pipeline defect level; the pipeline defect level includes the structural defect level and the functional defect level.

[0008] S200: Based on the preliminary assessment results of pipeline defect levels, the pipelines are screened, and the television inspection images of the screened pipelines are obtained. The final assessment results of pipeline defect levels are given using the GNN-SVM-based defect level final assessment model.

[0009] S300: Based on the final assessment results of the pipeline defect level, a comprehensive pipeline defect level is obtained through joint evaluation.

[0010] S400: Based on the comprehensive defect level of the pipeline, construct a repair strategy model based on PPO, and use this repair strategy model to obtain the optimal repair strategy;

[0011] S500: Executes the optimal repair strategy to repair the pipeline, and revises the preliminary defect level assessment model based on SVM multi-classifier and the final defect level assessment model based on GNN-SVM based on the pipeline defect severity level verified during the pipeline repair process.

[0012] Furthermore, the preliminary assessment process for pipeline defect levels includes:

[0013] S101: Use an adaptive edge detection algorithm to extract features from the sonar profile of the pipeline;

[0014] S102: The extracted feature information is used as the input of the SVM multi-classifier, and the corresponding pipeline defect level assessment result is used as the output of the SVM multi-classifier to obtain the trained preliminary defect level assessment model based on the SVM multi-classifier.

[0015] S103: Input the sonar profile of the pipeline to be evaluated into the preliminary assessment model of defect level based on SVM multi-classifier to obtain the preliminary assessment results of the pipeline defect.

[0016] Furthermore, feature extraction of the sonar contour map of the pipeline is performed using an adaptive edge detection algorithm, including:

[0017] S11: Filter and reduce noise in the sonar profile of the pipeline;

[0018] S12: Select the verification operator to calculate the gradient value and direction information of the contour map;

[0019] S13: Gradient magnitude calculation;

[0020] S14: Non-maximum suppression processing;

[0021] S15: Hysteresis boundary threshold segmentation;

[0022] S16: Obtaining pipeline contour information, which includes pipeline contour lines, pipe diameter, and pipeline sludge depth information.

[0023] Furthermore, the final assessment process for the pipeline defect level includes:

[0024] S201: Input the filtered pipeline television inspection images into the GNN graph neural network to extract graph features containing multidimensional defect features;

[0025] S202: The extracted graph features are used as the input of the SVM multi-classifier, and the corresponding pipeline defect level assessment results are used as the output of the SVM multi-classifier to obtain the trained GNN-SVM-based final defect level assessment model.

[0026] S203: Input the television inspection image of the pipeline to be evaluated into the final evaluation model of the defect level based on GNN-SVM to obtain the final evaluation result of the pipeline defect.

[0027] Furthermore, based on the final assessment results of the pipeline defect levels, a joint evaluation is conducted to obtain the overall pipeline defect severity level, including:

[0028] The degree of structural defects is categorized as L1, and the degree of functional defects as L2.

[0029] When L1≤L2, the overall defect level of the pipeline is L2;

[0030] When L1 > L2, the overall defect level of the pipeline is L1.

[0031] Furthermore, the construction of a PPO-based remediation strategy model includes:

[0032] S401: Selection of action and state space; Let the set of pipeline comprehensive defect severity levels at time t constitute the state space S. t The corresponding repair action space is A. t :

[0033] S t ={s1,s2,s3,s4,…,s n} (2)

[0034] A t ={a1, a2, a3} (3)

[0035] Among them, s n This represents the overall defect level of the pipeline at the nth assessment; a1, a2, and a3 represent no repair, partial repair, and overall repair, respectively.

[0036] S402: Design of the reward function; The reward function J is designed as a function related to each constraint term, as shown below:

[0037]

[0038] In the formula, y1, y2, y3, y4, y5 and y6 represent the amount of construction work, applicable pipe materials, repair scale, land area, flow impact and project cost factor, respectively; sigmoid() represents the normalization function;

[0039] S403: Design of the policy function; policy Given a continuous probability distribution, based on the current state, within the action space, according to... The algorithm selects the repair action with the highest probability, and the strategy optimization algorithm iteratively optimizes the strategy based on the reward value of each interaction. parameters Therefore, the optimal policy is selected, and the parameters of the optimal policy function are... The update formula is:

[0040]

[0041]

[0042] Among them, G t This represents the time from time t to the final state time t ref The total reward obtained from the entire collection, J t This represents the reward at time t. This represents the parameter at time t+1. Let represent the parameters at time t, and η represent the hyperparameters.

[0043] Furthermore, both the structural defect severity level and the functional defect severity level include normal, minor defect, moderate defect, severe defect, and major defect.

[0044] Furthermore, the selection principle for pipelines is to exclude pipelines where both the degree of structural defects and the degree of functional defects are normal.

[0045] Furthermore, s n It is any one of the following: minor defect, moderate defect, serious defect, and major defect.

[0046] According to a first aspect of the present invention, a reinforcement learning-based trenchless repair decision system for pipeline defects is provided, comprising:

[0047] The preliminary assessment module for pipeline defect levels is used to obtain sonar profile maps of pipelines and to provide preliminary assessment results of pipeline defect levels using a preliminary defect level assessment model based on an SVM multi-classifier. Pipeline defect levels include structural defect severity levels and functional defect severity levels.

[0048] The final step evaluation module for pipeline defect levels is used to screen pipelines based on the preliminary evaluation results of pipeline defect levels, obtain the television inspection images of the screened pipelines, and give the final evaluation results of pipeline defect levels using the GNN-SVM-based defect level final evaluation model.

[0049] The pipeline comprehensive defect severity assessment module is used to jointly evaluate the pipeline based on the final assessment results of the pipeline defect severity level to obtain the pipeline comprehensive defect severity level.

[0050] The repair strategy module is used to construct a repair strategy model based on PPO according to the overall defect severity level of the pipeline, and to obtain the optimal repair strategy using this repair strategy model;

[0051] The model correction module is used to execute the optimal repair strategy to repair the pipeline, and to correct the preliminary defect level assessment model based on SVM multi-classifier and the final defect level assessment model based on GNN-SVM based on the pipeline defect severity level verified during the pipeline repair process.

[0052] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0053] 1. By adopting an adaptive edge detection algorithm, accurate extraction of large-scale sonar contour maps was achieved, and an SVM multi-classifier was used to achieve preliminary assessment of the defect level of large-scale pipelines, saving a lot of manpower.

[0054] 2. The accuracy of the assessment results was ensured by verifying the preliminary assessment results using pipeline television inspection images;

[0055] 3. An optimization objective combining a reward function and multiple constraints was established, and the PPO algorithm was used to find the optimal repair strategy, making it easier for decision-makers to scientifically and quickly select a technically feasible and economically reasonable repair method.

[0056] 4. The evaluation model is modified using the repair results, forming a closed loop of detection and evaluation, which improves the reliability of the method. Attached Figure Description

[0057] Figure 1 This is a flowchart of a trenchless pipeline defect repair decision-making method based on reinforcement learning provided by the present invention;

[0058] Figure 2 This is a technical roadmap for a trenchless pipeline defect repair decision-making method based on reinforcement learning provided by the present invention;

[0059] Figure 3 This is a schematic diagram of the preliminary assessment process for large-scale pipeline defect levels provided by the present invention;

[0060] Figure 4This is a schematic diagram of a two-dimensional spatial SVM provided by the present invention;

[0061] Figure 5 This is a schematic diagram of the final evaluation process for large-scale pipeline defect levels provided by the present invention;

[0062] Figure 6 This is a schematic diagram of the GNN structure provided by the present invention;

[0063] Figure 7 This is a schematic diagram of a trenchless pipeline defect repair decision system based on reinforcement learning provided by the present invention. Detailed Implementation

[0064] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.

[0065] Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of this invention. Furthermore, the technical features of the various embodiments or individual embodiments provided by this invention can be arbitrarily combined to form feasible technical solutions. Such combinations are not constrained by the order of steps and / or structural composition patterns, but must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by this invention.

[0066] Figure 1 A flowchart of a trenchless pipeline defect repair decision-making method based on reinforcement learning is provided for this invention, as shown below. Figure 1 As shown, the method includes:

[0067] S100: Obtain sonar profile maps of large-scale pipelines and use an SVM-based multi-classifier-based preliminary defect level assessment model to provide preliminary assessment results of the defect levels of large-scale pipelines.

[0068] S200: Based on the preliminary assessment results, large-scale pipelines are screened, and television inspection images of the screened pipelines are obtained. Using the GNN-SVM-based defect level final assessment model, the final assessment results of the defect level of the large-scale pipelines are given.

[0069] S300: The overall defect level of the pipeline is obtained by jointly evaluating the degree of structural defects and the degree of functional defects.

[0070] S400: Based on the overall defect level of the pipeline, construct a repair strategy model based on PPO, and use this strategy model to obtain the optimal repair strategy;

[0071] S500: Executes the optimal repair strategy and, based on the degree of pipeline defects verified during the repair process, corrects the preliminary defect level assessment model based on SVM multi-classifier and the final defect level assessment model based on GNN-SVM.

[0072] Figure 2 This is the overall technical roadmap for this method.

[0073] In some embodiments of this application, such as Figure 3 As shown, the preliminary assessment process for the level of defects in large-scale pipelines is as follows:

[0074] S101: Use an adaptive edge detection algorithm to extract features from historical sonar contour maps;

[0075] S102: The extracted feature information is used as the input of the SVM multi-classifier, and the historical pipeline defect level assessment results are used as the output of the SVM multi-classifier to obtain a trained preliminary defect level assessment model based on the SVM multi-classifier.

[0076] S103: Input the sonar profile of the pipeline to be evaluated into the preliminary assessment model of defect level based on SVM multi-classifier to obtain the preliminary assessment results of the pipeline defect.

[0077] The pipeline defect level includes structural defect severity level and functional defect severity level. Structural defects refer to damage to the pipeline structure itself, affecting its strength, stiffness, and service life. Functional defects refer to changes in the pipeline's cross-sectional area, affecting its flow performance. Both structural and functional defect severity levels include normal, minor, moderate, severe, and major defects. For example, the structural defect severity levels of normal, minor, moderate, severe, and major defects correspond to no defects, minor defects with minimal structural impact, defects significantly exceeding minor defects and affecting pipeline operation, severe defects affecting structural condition, and major defects leading to severe or imminent pipeline damage. The SVM multi-classifier-based preliminary defect level assessment model, after inputting the pipeline's sonar profile map, outputs the structural and functional defect severity levels for that pipeline.

[0078] Furthermore, the adaptive edge detection algorithm for obtaining the pipeline contour includes the following steps:

[0079] S11: Filter and reduce noise in the sonar profile of large-scale pipelines;

[0080] S12: Select the verification operator to calculate the gradient value and direction information of the contour map;

[0081] S13: Gradient magnitude calculation;

[0082] S14: Non-maximum suppression processing;

[0083] S15: Hysteresis boundary threshold segmentation;

[0084] S16: Obtain pipeline contour information, including pipeline contour lines, pipe diameter, pipeline sludge depth lines, and other contour information.

[0085] The contour information includes the pipe outline, pipe diameter, and pipe sludge depth line.

[0086] Furthermore, the construction process of the SVM multi-classifier is as follows:

[0087] SVM (Support Vector Machine) is itself a binary classifier, such as... Figure 4 As shown. The SVM algorithm was originally designed for binary classification problems. When dealing with multi-class problems, it is necessary to construct a suitable multi-class classifier. This invention provides a method for constructing an SVM multi-class classifier, mainly by combining multiple binary classifiers to achieve the construction of the multi-class classifier, as follows: During training, samples of one class are sequentially assigned to one class, and the remaining samples are assigned to another class. In this way, samples of k classes are used to construct k SVMs. During classification, unknown samples are classified into the class with the largest classification function value.

[0088] In some embodiments of this application, such as Figure 5 As shown, the final assessment process for large-scale pipeline defect levels is as follows:

[0089] S201: Input the filtered historical TV inspection map of the pipeline into GNN (Graph Neural Networks) to extract graph features containing multidimensional defect features;

[0090] S202: The extracted graph features are used as the input of the SVM multi-classifier, and the historical pipeline defect level assessment results are used as the output of the SVM multi-classifier to obtain a trained defect level assessment model based on GNN-SVM.

[0091] S203: Input the television inspection image of the pipeline to be evaluated into the final evaluation model of the defect level based on GNN-SVM to obtain the final evaluation result of the pipeline defect.

[0092] Similarly, pipeline defect levels include structural defect severity levels and functional defect severity levels; both structural and functional defect severity levels include normal, minor, moderate, severe, and major defects. The final defect level evaluation model based on GNN-SVM, after inputting the pipeline's television inspection image, will output the pipeline's structural and functional defect severity levels. The television inspection image of the pipeline is simply the image of the pipeline itself.

[0093] Furthermore, the selection principle for pipelines is to eliminate pipelines from large-scale pipelines where both the degree of structural defects and the degree of functional defects are normal.

[0094] like Figure 6 As shown, the basic principle of a graph neural network is: inputting data (a graph) into the network (GNN) will produce output data (also a graph). Compared with the input graph, the vertices, edges, and global information of the output graph will be changed.

[0095] In some embodiments of this application, the degree of structural defect is denoted as L1, and the degree of functional defect is denoted as L2. The principle of joint evaluation based on the degree of structural defect and the degree of functional defect is as follows:

[0096] When L1≤L2, the overall defect level of the pipeline is L2;

[0097] When L1 > L2, the overall defect level of the pipeline is L1.

[0098] The results of the comprehensive defect assessment of the pipeline are shown in Table 1:

[0099] Table 1 Evaluation Results of Comprehensive Pipeline Defects

[0100]

[0101] In some embodiments of this application, the Proximal Policy Optimization (PPO) algorithm is a policy-based reinforcement learning algorithm that uses two neural networks. By inputting the current state of the "agent" into the neural network, corresponding "actions" and "rewards" are obtained. The state of the "agent" is then updated based on the "actions." According to the objective function containing "rewards" and "actions," gradient ascent is used to update the weight parameters in the neural network, thereby obtaining the "action" judgment that maximizes the overall reward value. The general form of the proximal policy optimization problem is:

[0102]

[0103] In the formula: For policy network functions, For policy network parameters; objective function The agent is in state S at time t. t At that time, according to the strategy Perform action A t Then, the reward and expected value obtained by all trajectories τ formed; β t ∈(0,1) is the reward discount factor, which can evaluate the subsequent rewards for performing the current action, and decreases as the time step increases; the constraint term is the optimization objective function of PPO under the constraint of average KL divergence, which minimizes the expected reward. In strategy The dominant function under these conditions; In strategy The expected value of the trajectory τ formed below; δ is the confidence region.

[0104] In this embodiment, the construction of the PPO-based remediation strategy model is specifically as follows:

[0105] S401: Selection of action and state space. Let the set of comprehensive pipeline defect levels at time t constitute the state space S. t The corresponding repair action space is A. t :

[0106] S t ={s1,s2,s3,s4,…,s n} (2)

[0107] A t ={a1, a2, a3} (3)

[0108] Among them, s n The overall pipeline defect level at the nth assessment, s n It can be any one of the following states: "minor defect", "moderate defect", "serious defect" and "major defect"; a1, a2 and a3 represent "no repair", "partial repair" and "overall repair" respectively.

[0109] S402: Design of the reward function. The reward function J is designed as a function related to each constraint term, as shown below:

[0110]

[0111] y1, y2, y3, y4, y5, and y6 represent the construction workload (person-days), applicable pipe material, repair scale, land area, flow impact, and project cost factor, respectively; sigmoid() represents the normalization function.

[0112] S403: Design of the policy function, policy Given a continuous probability distribution, based on the current state, within the action space, according to... The algorithm selects the repair action with the highest probability, and the strategy optimization algorithm iteratively optimizes the strategy based on the reward value of each interaction. parameters Therefore, the optimal policy is selected, and the parameters of the optimal policy function are... The update formula is:

[0113]

[0114]

[0115] Among them, G t This represents the time from time t to the final state time t ref The total reward obtained from the entire collection, J t This represents the reward at time t. This represents the parameter at time t+1. Let represent the parameters at time t, and η represent the hyperparameters.

[0116] In some embodiments of this application, the overall pipeline defect level is the result of online assessment. After step S400 is defined, manual verification and repair are performed. If the overall pipeline defect level of offline assessment is found to be inconsistent with the overall pipeline defect level of online assessment, the preliminary SVM defect level assessment model and the GNN-SVM defect level assessment model are corrected according to the overall pipeline defect level of offline assessment to ensure that the false negative rate of online assessment is less than 10% and the false positive rate is less than 10%.

[0117] Figure 7 This invention provides a schematic diagram of a trenchless pipeline defect repair decision system based on reinforcement learning. The system includes a preliminary assessment module 601 for pipeline defect levels, a final assessment module 602 for pipeline defect levels, a comprehensive pipeline defect severity assessment module 603, a repair strategy module 604, and a model correction module 605, wherein:

[0118] The preliminary assessment module 601 for pipeline defect levels is used to obtain sonar profile maps of large-scale pipelines and to give preliminary assessment results of the defect levels of large-scale pipelines using a defect level preliminary assessment model based on SVM multi-classifier.

[0119] The final evaluation module 602 for pipeline defect levels is used to screen large-scale pipelines based on the preliminary evaluation results, obtain the television inspection images of the screened pipelines, and use the GNN-SVM-based defect level final evaluation model to give the final evaluation results of the defect levels of large-scale pipelines.

[0120] The pipeline comprehensive defect assessment module 603 is used to obtain the pipeline comprehensive defect level based on a joint evaluation of the degree of structural defects and the degree of functional defects.

[0121] Repair strategy module 604 is used to construct a repair strategy model based on PPO according to the overall defect degree of the pipeline, and use the strategy model to obtain the optimal repair strategy.

[0122] The model correction module 605 is used to execute the optimal repair strategy and correct the preliminary defect level assessment model based on SVM multi-classifier and the final defect level assessment model based on GNN-SVM according to the degree of pipeline defects verified during the repair process.

[0123] It is understood that the reinforcement learning-based trenchless pipeline defect repair decision system provided by this invention corresponds to the reinforcement learning-based trenchless pipeline defect repair decision method provided in the foregoing embodiments. The relevant technical features of the reinforcement learning-based trenchless pipeline defect repair decision system can be referred to the relevant technical features of the reinforcement learning-based trenchless pipeline defect repair decision method, and will not be repeated here.

[0124] This invention provides a trenchless pipeline defect repair decision-making method and system based on reinforcement learning. The method includes: 1) establishing a preliminary defect level assessment model based on an SVM multi-classifier, which can use sonar profile maps of the pipeline to perform preliminary defect assessment; 2) establishing a final defect level assessment model based on a GNN-SVM, which can use television inspection images of the pipeline to perform final assessment; 3) establishing a joint evaluation system for structural and functional defect levels, realizing comprehensive pipeline defect level assessment; 4) establishing a repair strategy model based on PPO, which can obtain the optimal repair strategy according to the comprehensive pipeline defect level; 5) using the repair results to correct the assessment model, forming a closed loop between detection and assessment, improving the reliability of the method. This method has advantages such as accuracy, efficiency, and low cost.

[0125] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0126] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0127] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0128] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0129] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0130] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0131] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A trenchless pipeline defect repair decision-making method based on reinforcement learning, characterized in that, Includes the following steps: S100: Obtain the sonar profile of the pipeline and use the SVM-based multi-classifier-based preliminary defect level assessment model to give the preliminary assessment results of the pipeline defect level; the pipeline defect level includes the structural defect level and the functional defect level. S200: Based on the preliminary assessment results of pipeline defect levels, the pipelines are screened, and the television inspection images of the screened pipelines are obtained. The final assessment results of pipeline defect levels are given using the GNN-SVM-based defect level final assessment model. S300: Based on the final assessment results of the pipeline defect level, a comprehensive pipeline defect level is obtained through joint evaluation. S400: Based on the comprehensive defect level of the pipeline, construct a repair strategy model based on PPO, and use this repair strategy model to obtain the optimal repair strategy; S500: Executes the optimal repair strategy to repair the pipeline, and corrects the preliminary defect level assessment model based on SVM multi-classifier and the final defect level assessment model based on GNN-SVM based on the pipeline defect severity level verified during the pipeline repair process. The construction of a PPO-based remediation strategy model includes: S401: Selection of Action and State Space; Let the set of pipeline comprehensive defect severity levels at time t constitute the state space. The corresponding repair action space is : ; ; in, This represents the overall pipeline defect level during the nth assessment. , and These represent no repair, partial repair, and overall repair, respectively. S402: Design of the reward function; The reward function J is designed as a function related to each constraint term, as shown below: ; In the formula, , , , , and These represent the workload, applicable pipe materials, repair scale, land area, flow impact, and project cost factor, respectively. Represented as a normalization function; S403: Design of the policy function; policy Given a continuous probability distribution, based on the current state, within the action space, according to... The algorithm selects the repair action with the highest probability, and the strategy optimization algorithm iteratively optimizes the strategy based on the reward value of each interaction. parameters Thus, the optimal strategy is selected, and the optimal strategy function parameters are... The update formula is: ; ; in, Represents the time from time t to the final state. The total rewards obtained from the entire collection, This represents the reward at time t. This represents the parameter at time t+1. The parameter at time t, This represents hyperparameters.

2. The trenchless pipeline defect repair decision-making method based on reinforcement learning according to claim 1, characterized in that, The preliminary assessment process for pipeline defect levels includes: S101: Use an adaptive edge detection algorithm to extract features from the sonar contour map of the pipeline; S102: The extracted feature information is used as the input of the SVM multi-classifier, and the corresponding pipeline defect level assessment result is used as the output of the SVM multi-classifier to obtain the trained preliminary defect level assessment model based on the SVM multi-classifier. S103: Input the sonar profile of the pipeline to be evaluated into the preliminary assessment model of defect level based on SVM multi-classifier to obtain the preliminary assessment results of the pipeline defect.

3. The trenchless pipeline defect repair decision-making method based on reinforcement learning according to claim 2, characterized in that, Feature extraction of the sonar contour map of the pipeline using an adaptive edge detection algorithm includes: S11: Filter and reduce noise in the sonar profile of the pipeline; S12: Select the verification operator to calculate the gradient value and direction information of the contour map; S13: Gradient magnitude calculation; S14: Non-maximum suppression processing; S15: Hysteresis boundary threshold segmentation; S16: Obtaining pipeline contour information, which includes pipeline contour lines, pipe diameter, and pipeline sludge depth information.

4. The trenchless pipeline defect repair decision-making method based on reinforcement learning according to claim 1, characterized in that, The final assessment process for pipeline defect levels includes: S201: Input the filtered pipeline television inspection images into the GNN graph neural network to extract graph features containing multidimensional defect features; S202: The extracted graph features are used as the input of the SVM multi-classifier, and the corresponding pipeline defect level assessment results are used as the output of the SVM multi-classifier to obtain the trained GNN-SVM-based final defect level assessment model. S203: Input the television inspection image of the pipeline to be evaluated into the final evaluation model of the defect level based on GNN-SVM to obtain the final evaluation result of the pipeline defect.

5. The trenchless pipeline defect repair decision-making method based on reinforcement learning according to claim 1, characterized in that, Based on the final assessment results of the pipeline defect levels, a joint evaluation is conducted to obtain the overall pipeline defect severity level, which includes: The degree of structural defects is categorized as follows: The functional defect level is ; when At that time, the overall defect level of the pipeline was: ; when At that time, the overall defect level of the pipeline was: .

6. The trenchless pipeline defect repair decision-making method based on reinforcement learning according to any one of claims 1 to 5, characterized in that, Both structural defect severity levels and functional defect severity levels include normal, minor defect, moderate defect, severe defect, and major defect.

7. The trenchless pipeline defect repair decision-making method based on reinforcement learning according to claim 6, characterized in that, The selection principle for pipelines is to exclude pipelines where both the degree of structural defects and the degree of functional defects are normal.

8. The trenchless pipeline defect repair decision-making method based on reinforcement learning according to claim 6, characterized in that, It is any one of the following: minor defect, moderate defect, serious defect, and major defect.

9. A trenchless pipeline defect repair decision system based on reinforcement learning, characterized in that, include: The preliminary assessment module for pipeline defect levels is used to obtain sonar profile maps of pipelines and to provide preliminary assessment results of pipeline defect levels using a preliminary defect level assessment model based on an SVM multi-classifier. Pipeline defect levels include structural defect severity levels and functional defect severity levels. The final step evaluation module for pipeline defect levels is used to screen pipelines based on the preliminary evaluation results of pipeline defect levels, obtain the television inspection images of the screened pipelines, and give the final evaluation results of pipeline defect levels using the GNN-SVM-based defect level final evaluation model. The pipeline comprehensive defect severity assessment module is used to jointly evaluate the pipeline based on the final assessment results of the pipeline defect severity level to obtain the pipeline comprehensive defect severity level. The repair strategy module is used to construct a repair strategy model based on PPO according to the overall defect severity level of the pipeline, and to obtain the optimal repair strategy using this repair strategy model; The model correction module is used to execute the optimal repair strategy to repair the pipeline, and to correct the initial defect level assessment model based on SVM multi-classifier and the final defect level assessment model based on GNN-SVM based on the pipeline defect severity level verified during the pipeline repair process. The construction of a PPO-based remediation strategy model includes: S401: Selection of Action and State Space; Let the set of pipeline comprehensive defect severity levels at time t constitute the state space. The corresponding repair action space is : ; ; in, This represents the overall pipeline defect level during the nth assessment. , and These represent no repair, partial repair, and overall repair, respectively. S402: Design of the reward function; The reward function J is designed as a function related to each constraint term, as shown below: ; In the formula, , , , , and These represent the workload, applicable pipe materials, repair scale, land area, flow impact, and project cost factor, respectively. Represented as a normalization function; S403: Design of the policy function; policy Given a continuous probability distribution, based on the current state, within the action space, according to... The algorithm selects the repair action with the highest probability, and the strategy optimization algorithm iteratively optimizes the strategy based on the reward value of each interaction. parameters Thus, the optimal strategy is selected, and the optimal strategy function parameters are... The update formula is: ; ; in, Represents the time from time t to the final state. The total rewards obtained from the entire collection, This represents the reward at time t. This represents the parameter at time t+1. The parameter at time t, This represents hyperparameters.

Citation Information

Patent Citations

  • AUV pipeline cycle management method based on image feature depth reinforcement learning

    CN109407682A

  • Pipeline defect identification method and device, terminal equipment and storage medium

    CN113284109A