Pipeline inner wall maintenance state evaluation method based on intelligent image recognition

By combining a temporal differential autoencoder with a spatiotemporal graph convolutional network, the technical problem of dynamic image acquisition and evaluation in existing pipeline inner wall detection methods is solved. This enables efficient, precise, and automated evaluation of the pipeline inner wall condition, improving the detection coverage and discrimination accuracy. It is suitable for intelligent inspection and maintenance decision-making for various types of pipelines.

CN120913015AInactive Publication Date: 2025-11-07TIANJIN BINHAI NEW AREA ZHONGDA PETROLEUM TECHNOLOGY SERVICES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510871818.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-11-07
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing pipeline inner wall inspection methods are difficult to achieve large-scale continuous inspection, which poses risks of omission and misjudgment. Furthermore, they are difficult to capture subtle structural changes and temporal dynamic characteristics, resulting in insufficient assessment accuracy and failing to meet the intelligent assessment needs of urban pipe networks and industrial pipelines.

Method used

A temporal differential autoencoder and a spatiotemporal graph convolutional network are used, combined with dynamic image acquisition, image enhancement and preprocessing, temporal differential feature extraction, and spatiotemporal structural feature modeling. Through multimodal feature fusion, the maintenance status of the inner wall of the pipeline is automatically evaluated. A spatiotemporal graph structure of multiple frames of dynamic images is constructed, and a multi-layer spatiotemporal graph convolutional network is used to extract high-level structural features, and feature fusion and hierarchical evaluation are performed.

Benefits of technology

It enables efficient, precise, and automated assessment of the condition of the pipeline inner wall, improves the coverage and accuracy of detection, supports multi-dimensional result query and traceability, is highly adaptable, and is suitable for intelligent inspection and maintenance decision-making of various types of pipelines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913015A_ABST
    Figure CN120913015A_ABST
Patent Text Reader

Abstract

The invention discloses a pipeline inner wall maintenance state evaluation method based on intelligent image recognition, and the method comprises the following steps: S1, obtaining the dynamic image data of the surface of the pipeline inner wall, and dividing the dynamic image data into a plurality of image subsequences; s2, performing image enhancement and preprocessing on the image subsequences to form a preprocessed image sequence; s3, inputting the preprocessed image sequence into an auto-encoder model with a time sequence difference mechanism, and outputting a difference feature expression; s4, extracting high-level feature expression by adopting a space-time diagram convolutional network; s5, performing feature fusion on the differential feature expression and the high-level feature expression to obtain multi-modal feature expression; and S6, outputting a maintenance state grade label of each detection section by adopting a classification and state evaluation algorithm. According to the method, the time sequence difference self-encoder and the space-time diagram convolutional network are fused, and intelligent judgment and grading evaluation of the maintenance state of the inner wall of the pipeline are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of image processing and intelligent detection technology, and particularly relates to a pipeline inner wall maintenance state evaluation method based on intelligent image recognition. BACKGROUND

[0002] With the extensive laying of urban infrastructure and industrial pipe network systems, pipelines play a key role in the fields of water supply, gas, chemical industry, energy, etc. The structural health and surface state of the pipeline inner wall are directly related to the safety of fluid transportation and the reliability of system operation. However, during long-term operation, the pipeline inner wall is prone to various surface defects such as corrosion, cracking, fouling, and coating peeling, which can cause pipeline leakage, blockage, and even structural failure in severe cases, posing a significant risk to public safety and enterprise operation. Therefore, regular maintenance state detection and evaluation of the pipeline inner wall have become an important part of pipeline operation and management.

[0003] Existing pipeline inner wall detection methods mainly rely on manual visual inspection, mechanical endoscope shooting, ultrasonic detection, or local sampling analysis. These methods have low detection efficiency, strong subjectivity, and are difficult to achieve large-scale continuous detection, and are highly dependent on environmental lighting and operating experience. Especially in long-distance pipelines, complex bends, concealed installations, or harsh environments, manual detection is difficult to ensure full coverage of the pipeline section, and there is a risk of omission and misjudgment. At the same time, traditional detection methods are difficult to capture subtle structural changes, time series dynamic characteristics, and large-scale distribution patterns of the pipeline inner wall, and have limited ability to fine-tune early defects, development trends, and maintenance effects.

[0004] In recent years, with the rapid development of image recognition, deep learning, graph neural networks, and other artificial intelligence technologies, pipeline detection methods based on machine vision have gradually emerged. For example, a camera acquisition robot or intelligent endoscope is used to realize automatic acquisition of pipeline inner wall surface images, and convolutional neural networks, feature extraction algorithms, etc. are used to detect and identify defects in the images. However, existing pipeline detection methods based on image recognition generally have the following shortcomings:

[0005] First, there is a lack of systematic modeling of spatiotemporal features of dynamic image sequences. Most methods only analyze static single-frame images, ignoring the correlation between defect evolution and surface changes in time and space, making it difficult to accurately capture the details of multiple frames and distinguish dynamic patterns.

[0006] Existing feature fusion and state evaluation algorithms are mostly simple classification or static threshold judgment, which are difficult to adapt to the joint analysis needs of multi-source features, complex environments, and diverse defect types, resulting in limited accuracy and discrimination of state evaluation.

[0007] In addition, the current system still has a large room for improvement in terms of automation, detection speed, intelligent decision-making, and result storage and traceability, and is difficult to meet the actual needs of large-scale, continuous, and intelligent evaluation in urban pipe networks and industrial pipelines.

[0008] How to combine dynamic image spatio-temporal modeling, adaptive feature extraction, and multi-modal intelligent recognition to achieve efficient, fine, and automated evaluation of the maintenance state of the inner wall of the pipeline has become a core problem that needs to be solved by those skilled in the art. SUMMARY

[0009] One object of the present application is to provide a pipeline inner wall maintenance state evaluation method based on intelligent image recognition. The present application realizes automated modeling and multi-dimensional feature analysis of dynamic image sequences of the inner wall of the pipeline by fusing a time series difference autoencoder and a spatio-temporal graph convolution network. In view of the problems existing in the pipeline inner wall detection process, such as multi-source data collaboration, difficulty in capturing detail changes, and insufficient state evaluation accuracy, a complete evaluation process is established from dynamic image acquisition, image enhancement and preprocessing, time series difference feature extraction, spatio-temporal structure feature modeling, multi-modal feature fusion to intelligent grading of the maintenance state. Based on the spatial adjacency of the pixel points and the inter-frame time series relationship, a spatio-temporal graph structure of multiple dynamic images is constructed, and a multi-layer spatio-temporal graph convolution network is used to recursively aggregate node features to obtain high-level structure features reflecting the overall spatial distribution and time series evolution law. Through multi-modal feature fusion of difference features and spatio-temporal features, a unified feature representation suitable for multi-class maintenance state discrimination is generated by using normalization, fully connected fusion network, and feature weighting mechanism. The whole process of structured storage and management of image data, features, and evaluation results is realized, and the result query and traceability according to the time, space, grade, and other dimensions are supported. This method has the advantages of high automation, strong adaptability, wide coverage, fine discrimination, and easy engineering landing, and is suitable for inner wall state inspection and maintenance decision-making scenarios of various typical pipelines such as water supply, gas, petrochemical, and municipal pipelines.

[0010] According to the pipeline inner wall maintenance state evaluation method based on intelligent image recognition, the following steps are included:

[0011] S1, uniformly moving along the pipeline axis to obtain dynamic image data of the inner wall surface of the pipeline, and dividing the dynamic image data according to a preset time window to obtain a plurality of image sub-sequences;

[0012] S2, image enhancement and preprocessing are performed on each frame of image in the image sub-sequence to form a preprocessed image sequence;

[0013] S3, inputting the preprocessed image sequence into a self-encoder model with a time difference mechanism to automatically capture the change information between adjacent frames or local areas, and outputting a differential feature expression distinguishing the surface detail change features of the inner wall of the pipeline through an adaptive feature coding process;

[0014] S4, constructing a space-time graph structure based on the image sub-sequence, taking the pixel points of the multi-frame images as nodes, combining the spatial adjacency and inter-frame time sequence relationship, and extracting a high-level feature expression reflecting the overall space-time distribution pattern and dynamic evolution law by using a space-time graph convolution network;

[0015] S5, performing feature fusion processing on the differential feature expression and the high-level feature expression to obtain a multi-modal feature expression comprehensively reflecting the maintenance state of the inner wall of the pipeline;

[0016] S6, based on the multi-modal feature expression, using a classification and state evaluation algorithm to identify the maintenance state in the image sequence, and combining the feature distribution in the sub-sequence to grade the maintenance state of the inner wall of the pipeline, and finally outputting the maintenance state grade label of each detection section.

[0017] Optionally, the S1 specifically comprises: setting the axial travel speed of the mobile device and the sampling frame rate, synchronously collecting position, ambient light intensity and device posture information in each time window, and time stamping the information and dynamic image data to obtain an image sequence covering the full length of the pipeline and being time-continuous.

[0018] Optionally, the S2 specifically comprises:

[0019] S21, performing adaptive histogram equalization operation on each frame of image, and generating an enhanced image sequence by using the method of random cropping and random angle rotation;

[0020] S22, performing denoising processing and edge preserving filtering on the basis of the enhanced image sequence to form a preprocessed image sequence.

[0021] Optionally, the S3 specifically comprises:

[0022] S31, arranging the preprocessed image sequence in time sequence as the input of the self-encoder model;

[0023] S32, performing pixel-level difference operation on each pair of frames in the image sequence to generate a difference map reflecting the dynamic change of the small structure of the inner wall of the pipeline;

[0024] S33, concatenating the original frame and the difference map corresponding to the original frame in the channel dimension to form a composite feature input;

[0025] S34, input the spliced composite feature sequence into the autoencoder model, the encoder part extracts spatial-temporal joint features through multiple layers of convolution to obtain hidden layer representation, and the decoder part reconstructs the hidden layer features to obtain reconstruction output;

[0026] S35, end-to-end optimization is performed with the reconstruction error of the input and the output as the target, and finally the model encoding feature sequence is output as the differential feature expression for distinguishing the surface detail changes of the inner wall of the pipeline:

[0027]

[0028] wherein, I t (i,j) is the value of the original image at pixel (i,j) in the t-th frame, is the value of the reconstructed image at pixel (i,j) in the t-th frame, H and W are the height and width of the image, T is the total number of frames of the collected dynamic image, L rec is the reconstruction loss function.

[0029] Optionally, the S4 specifically comprises:

[0030] S41, taking all pixel points in the image subsequence as nodes, constructing a space-time graph structure based on spatial adjacency relationship and time sequence corresponding relationship;

[0031] S42, the spatial adjacency relationship adopts a pixel eight-neighborhood structure, the time sequence relationship connects nodes in adjacent frames at the same position, and the node feature is initialized as a differential feature vector of the corresponding pixel;

[0032] S43, applying a space-time graph convolution operation to the space-time graph structure, and the node feature update rule is:

[0033]

[0034] wherein, N S (v) is a spatial neighbor of node v, N T (v) is a time sequence neighbor, is a learnable weight, and σ is an activation function, is a feature vector of node v in the l-th layer of the space-time graph convolution network, is a feature vector of node u in the l-th layer of the space-time graph convolution network, and b (l) is a bias term of the l-th layer;

[0035] S44, in the spatio-temporal graph convolution network, a plurality of spatio-temporal graph convolution units are arranged, each unit is a layer, receives the node features output by the previous layer, and aggregates and updates the features according to the spatial adjacency and time sequence adjacency relationship, each layer is stacked in a predetermined order, the output of the previous layer is used as the input of the next layer, and the layers are recursively propagated until the predetermined network depth, and the node features output by the last layer are subjected to a global aggregation operation to obtain high-level feature expression;

[0036] S45, in each spatio-temporal graph convolution unit, a residual connection structure is arranged, the input features of the spatio-temporal graph convolution unit and the output features after convolution processing are added element by element to form residual output, and the residual output is further subjected to feature normalization processing to standardize the values of each feature component, all layers adopt this structure, and finally normalized spatio-temporal high-level feature expression is obtained as the input of the multi-modal feature fusion module.

[0037] Optionally, the S5 specifically includes:

[0038] S51, the difference feature sequence output by the time sequence difference self-encoder is spliced with the spatio-temporal high-level feature expression output by the spatio-temporal graph convolution network, for each sample to be fused, the feature vector of the sample to be fused in the difference feature expression and the feature vector in the spatio-temporal feature expression are selected respectively, the feature vectors are combined in the feature dimension direction, and the fusion features are processed by using the mean variance normalization method;

[0039] S52, the fusion features are input into a fusion network, the fusion network is a multi-layer fully connected neural network structure, and is composed of a plurality of fully connected layers stacked in sequence, in each fully connected layer, linear transformation and addition operation are performed on the input feature vector, the input vector is multiplied by the layer weight matrix and an offset term is added, then a nonlinear transformation is performed through an activation function to obtain layer output features, and the layers are transmitted layer by layer, and after transformation through all the fully connected layers, a fusion feature vector is output;

[0040] S53, a feature weighting module is arranged in the fusion network, the weight parameters of each feature component are iteratively updated according to historical samples by using a back propagation algorithm, in the forward inference stage, each component of the feature vector output by the fusion network is multiplied by the corresponding weight parameter to obtain weighted fusion features, and all the weighted feature components are combined into fusion features.

[0041] Optionally, the S6 includes:

[0042] S61, input the fusion features into a state recognition module, in the state recognition module, first set two fully connected neural network layers, perform multi-layer feature transformation and extraction on the input fusion features, set an output grading unit in the last layer, map the features to a preset number of maintenance state categories, output a score for each category, compare all output scores, determine the maintenance state category to which each image subsequence belongs according to the maximum value classification rule, and assign the recognized maintenance state category label to the corresponding image subsequence;

[0043] S62, according to the pipeline inner wall maintenance state grading standard, process the feature distribution output by the state recognition module, divide the numerical range of the maintenance state features into several intervals according to the pre-set grading parameters, compare the feature distribution of each image subsequence with each grading interval, and determine the image subsequence as a specific maintenance state level according to the interval rule, the maintenance state level is set to multiple levels, covering the entire detection period section, including the highest level, the higher level, the middle level, the lower level and the lowest level, which correspond to different maintenance state descriptions respectively, the determination process of the section is executed according to the unified grading standard, until all image subsequences complete the assignment of the maintenance state level;

[0044] S63, for samples determined as critical states or located near the adjacent level boundary, a time sequence consistency discrimination mechanism is used based on the arrangement of the image subsequences in time sequence, specifically: comparing the state level of the current detection section with the state levels of the adjacent sections before and after the detection section, correcting the state level determination result, eliminating occasional abnormalities, and finally outputting the maintenance state level label of each detection section.

[0045] Optionally, the state labels output by the maintenance state recognition and grading evaluation module and the corresponding image sequence indexes are stored in the data management module together, and the data management module supports querying and retrieving historical detection results according to time, spatial section and state level.

[0046] The beneficial effects of the present application are:

[0047] The present application configures a dynamic image acquisition module at the pipeline inspection equipment end, continuously obtains pipeline inner wall surface dynamic image data covering the whole period along the whole length of the pipeline, and combines multi-source environmental parameters such as acquisition path, speed, attitude and illumination, to realize time-space synchronization and multi-dimensional data structured management. Compared with traditional manual sampling and single-point static sampling, the present application can obtain more complete and time-continuous image information, and lay a solid data foundation for subsequent automatic evaluation.

[0048] In the data preprocessing stage, the application adopts multiple image processing strategies such as adaptive histogram equalization, random enhancement and denoising to form a unified standard high-quality input image sequence. Based on the feature extraction process of the time difference self-encoder, the dynamic modeling of the adjacent frames and the local area detail changes is realized, and the evolution process of the small structure of the pipeline inner wall surface is effectively reflected. Cooperating with the space adjacency and the inter-frame time sequence relationship established by the space-time graph structure, combined with the multi-layer space-time graph convolution network, the system can progressively mine the spatial distribution law and dynamic evolution mode, and improve the global consistency and discrimination ability of feature expression.

[0049] The application proposes a multi-modal feature fusion and normalization mechanism, which fully integrates the difference features and space-time high-level features through feature-level splicing, fully connected fusion network and feature weighting module, to generate unified and discriminative fusion feature expression. Based on the fusion feature, the automatic grading discrimination module combines the preset maintenance state grade system to perform multi-level label judgment on each detection section of the pipeline inner wall, supporting graded, zoned and time-based state fine management.

[0050] The system integrates all maintenance state evaluation results, image data and related environmental information according to the metadata such as section and time, supports multi-dimensional historical result query, data traceability and automatic report generation. Through the whole process of structured management and intelligent evaluation, the application not only greatly improves the pipeline detection efficiency and information integrity, but also provides reliable technical support for pipeline maintenance decision, operation management and intelligent inspection, effectively promotes the safety of pipeline network, reduces cost and increases efficiency, and upgrades the intelligentization of engineering application. BRIEF DESCRIPTION OF DRAWINGS

[0051] The accompanying drawings are included to provide a further understanding of the application, and constitute a part of the specification, which together with the embodiments of the application, are used to explain the application, and do not constitute a limitation on the application. In the drawings:

[0052] Fig. 1 The application proposes a whole flow chart of a pipeline inner wall maintenance state evaluation method based on intelligent image recognition;

[0053] Fig. 2 The application proposes a feature extraction and fusion structure diagram of a pipeline inner wall maintenance state evaluation method based on intelligent image recognition;

[0054] Fig. 3 The application proposes a maintenance state evaluation grading flow chart of a pipeline inner wall maintenance state evaluation method based on intelligent image recognition. DETAILED DESCRIPTION

[0055] The application will be described in further detail below with reference to the drawings. These drawings are simplified schematic diagrams and only show the basic structure of the application in a schematic manner, and thus only show the components relevant to the application.

[0056] Reference Figs. 1-3 A pipeline inner wall maintenance state evaluation method based on intelligent image recognition, comprising the following steps:

[0057] S1. Moving at a uniform speed along the pipeline axis to obtain dynamic image data of the pipeline inner wall surface, and dividing the dynamic image data according to a preset time window to obtain a plurality of image subsequences;

[0058] S2. Image enhancement and preprocessing are performed on each frame of image in the image subsequence to form a preprocessed image sequence;

[0059] S3. The preprocessed image sequence is input into an autoencoder model with a time difference mechanism, the change information between adjacent frames or local areas is automatically captured, and through an adaptive feature coding process, a difference feature expression distinguishing the detail change characteristics of the pipeline inner wall surface is output;

[0060] S4. Based on the image subsequence, taking the pixel points of multiple images as nodes, constructing a space-time graph structure combining spatial adjacency and inter-frame time sequence relationship, and using a space-time graph convolution network to extract high-level feature expression reflecting the overall space-time distribution pattern and dynamic evolution law;

[0061] S5. Feature fusion processing is performed on the difference feature expression and the high-level feature expression obtained above to obtain a multi-modal feature expression comprehensively reflecting the maintenance state of the pipeline inner wall;

[0062] S6. Based on the multi-modal feature expression, a classification and state evaluation algorithm is used to identify the maintenance state in the image sequence, and the pipeline inner wall maintenance state is graded and evaluated in combination with the feature distribution in the subsequence, and finally the maintenance state grade label of each detection section is output.

[0063] The application forms a closed-loop automatic pipeline inner wall state evaluation chain by constructing an intelligent identification and evaluation process for the dynamic image of the pipeline inner wall surface, from multi-source synchronous acquisition, image enhancement preprocessing, time series difference feature extraction, spatio-temporal structure modeling, feature fusion to maintenance state grading, effectively solving the limitations of the existing detection methods relying on manual sampling inspection, static interpretation and single feature analysis. The time series difference autoencoder and spatio-temporal graph convolution network are innovatively introduced to realize dynamic modeling of multiple frame detail changes and systematic mining of spatial distribution rules, and the precision and reliability of the maintenance state identification are improved by cooperating with the multi-modal feature fusion and grading determination mechanism. The method is suitable for automatic inspection and intelligent maintenance decision of various complex pipeline networks such as urban water supply, gas and petrochemical, and provides strong technical support for the whole life cycle safety management and intelligent operation and maintenance upgrade of pipeline facilities.

[0064] In the embodiment, S1 specifically includes: setting the axial running speed of the mobile device and the sampling frame rate, synchronously acquiring position, environmental light intensity and device posture information in each time window, and timestamp aligning these information with the dynamic image data to obtain an image sequence covering the whole length of the pipeline and being time series continuous.

[0065] In the embodiment, S2 specifically includes:

[0066] S21, adaptive histogram equalization operation is performed on each frame of image, and an enhanced image sequence is generated by using random cropping and random angle rotation;

[0067] S22, on the basis of the enhanced image sequence, denoising processing and edge preserving filtering are performed to form a preprocessed image sequence.

[0068] By optimizing the pipeline inner wall dynamic image acquisition process, first, the axial speed of the mobile device and the sampling frame rate are set to ensure that the image covers the whole length of the pipeline, and multi-source parameters such as position, light and posture are synchronously acquired in each time window. After timestamp alignment, a spatio-temporally continuous and information complete original image sequence is obtained. Based on this basic data, adaptive histogram equalization, random cropping and rotation and other diversified enhancement strategies are used, and denoising and edge preserving filtering processing is added, which effectively improves the contrast and detail performance of the image, and establishes a high-quality data foundation for subsequent dynamic feature extraction and intelligent identification, breaking through the limitations of traditional single acquisition and static processing methods in data coverage, time series consistency and information richness.

[0069] In the embodiment, S3 specifically includes:

[0070] S31, arranging the preprocessed image sequence in time sequence as an input of the autoencoder model;

[0071] S32, for the adjacent two frames in the image sequence, performing pixel-level difference operation on each pair of frames to generate a difference graph reflecting the dynamic change of the small structure of the inner wall surface of the pipeline;

[0072] S33, splicing the original frame and the difference graph corresponding to the original frame in the channel dimension to form a composite feature input;

[0073] S34, input the spliced composite feature sequence into the autoencoder model, the encoder part extracts spatial-temporal joint features through multi-layer convolution to obtain hidden layer representation, and the decoder part reconstructs the hidden layer features to obtain reconstruction output;

[0074] S35, end-to-end optimization is performed with the reconstruction error of the input and the output as the target, and finally the model encoding feature sequence is output as the difference feature expression of the high-sensitive area distinguishing the small structure change of the inner wall surface of the pipeline:

[0075]

[0076] Wherein, I t (i,j) is the value of the original image of the t-th frame at pixel (i,j), is the value of the reconstructed t-th frame image of the autoencoder at pixel (i,j), H and W are the height and width of the image, T is the total number of frames of collecting the dynamic image of the pipeline, L rec is the reconstruction loss function.

[0077] The embodiment introduces the time series difference and the autoencoder deep feature extraction mechanism in the preprocessed image sequence, realizes the high-sensitive capture of the dynamic small structure change of the inner wall surface of the pipeline. The continuous frame images are input into the model in time sequence, pixel-level difference is performed on each pair of adjacent frames, difference graph reflecting the small structure change is generated, and then the original image is spliced to construct the composite feature input. The autoencoder extracts spatial and temporal joint features through multi-layer convolution structure, and takes the reconstruction error between the input and the output as the optimization target to automatically encode and output the difference feature expression. This process breaks through the limitations of traditional static feature extraction and single frame analysis, realizes the deep fusion of multi-frame information and the adaptive representation of dynamic features, and provides accurate and high-dimensional data support for subsequent pipeline state modeling and intelligent identification.

[0078] In the embodiment, the S4 specifically comprises:

[0079] S41, taking all pixel points in the image sub-sequence as nodes, constructing a space-time graph structure based on spatial adjacency relationship and time sequence corresponding relationship;

[0080] S42, the spatial adjacency relationship adopts pixel eight-neighborhood structure, the time sequence relationship connects the nodes in the same position in adjacent frames, and the node feature is initialized as the difference feature vector of the corresponding pixel;

[0081] S43, apply a spatio-temporal graph convolution operation to the spatio-temporal graph structure, and the node feature updating rule is:

[0082]

[0083] wherein N S (v) is a spatial neighbor of node v, N T (v) is a time neighbor, is a learnable weight, and sigma is an activation function, is a feature vector of node v in the lth layer of the spatio-temporal graph convolution network, is a feature vector of node u in the lth layer of the spatio-temporal graph convolution network, and b (l) is a bias term of the lth layer;

[0084] S44, in the spatio-temporal graph convolution network, a plurality of spatio-temporal graph convolution units are set, each unit is a layer, receives the node features output by the previous layer, and aggregates and updates the features according to the spatial adjacency and time adjacency relationship, each layer is stacked in a predetermined order, the output of the previous layer is used as the input of the next layer, and the layers are recursively propagated until the predetermined network depth, and the node features output by the last layer are subjected to a global aggregation operation to obtain a high-level feature expression;

[0085] S45, in each spatio-temporal graph convolution unit, a residual connection structure is set, the input features of the spatio-temporal graph convolution unit and the output features after convolution processing are added element by element to form a residual output, and the residual output is further subjected to feature normalization processing to standardize the values of each feature component, and all layers adopt this structure to finally obtain a normalized spatio-temporal high-level feature expression as the input of the multi-modal feature fusion module.

[0086] The present application aims at the technical problems that subtle structure changes in the dynamic image of the inner wall surface of the pipeline are difficult to accurately capture and spatio-temporal dynamic information fusion is insufficient, and proposes a dynamic feature extraction and fusion method with time series difference auto-encoder and spatio-temporal graph convolution network collaborative modeling as the core.

[0087] In the present application, firstly, the pixel-level difference between any two frames in the image sequence is differentiated to automatically reveal the subtle change information of the pipeline inner wall under continuous time series, and form a time series difference feature. Subsequently, the original image and the difference image are spliced in the channel dimension, input into the auto-encoder deep feature extraction structure, and the spatial distribution and dynamic evolution information are jointly modeled. The auto-encoder uses multi-layer convolution and nonlinear mapping mechanism to automatically encode the spatial texture features and time difference features of the image, and outputs a dynamic structure expression with high sensitivity through end-to-end optimization.

[0088] In the spatio-temporal structure modeling stage, the application takes a pixel point as a node, constructs a spatio-temporal graph structure of multiple frames of images based on spatial adjacency and inter-frame time sequence correspondence, and explicitly represents the spatial connection and dynamic evolution of the inner wall region. A multi-layer spatio-temporal graph convolution network is used to recursively update the node features, realize hierarchical abstraction and deep fusion of the spatial structure mode and the time sequence dynamic law, and effectively alleviate the gradient degradation and feature loss of the deep network through residual connection and feature normalization mechanism, laying a foundation for long-distance dependence and global feature expression.

[0089] In addition, in the feature fusion stage, the application proposes feature-level splicing and normalization of differential features and spatio-temporal features, and introduces a multi-layer fully connected fusion network and a feature weighting module. The fusion network supports end-to-end trainable distribution of the weights of each feature component, so that the model can adaptively adjust the contribution proportion of multi-source features in subsequent state evaluation according to the discriminative information in the training data.

[0090] In the final maintenance state classification evaluation process, the application innovatively combines multi-modal features, unified standard classification parameters and time sequence consistency discrimination mechanism to realize continuous, classified and spatio-temporal consistent determination of the maintenance state in a dynamic pipeline scene, and provides detailed and full-cycle maintenance risk warning and decision reference for actual engineering.

[0091] The above technical path of the application breaks through the technical bottlenecks of existing static feature extraction, single-frame interpretation and manual subjective determination, and for the first time realizes high-dimensional automatic expression of multi-frame dynamic structure information of the inner wall of the pipeline, deep fusion of multi-source spatio-temporal features and intelligent classification, and significantly improves the automation, intelligence and practicability level of pipeline maintenance detection.

[0092] In the embodiment, the S5 specifically includes:

[0093] S51, the differential feature sequence output by the time sequence differential autoencoder is spliced with the spatio-temporal high-level feature expression output by the spatio-temporal graph convolution network at a feature level, for each sample to be fused, the feature vector of the sample to be fused in the differential feature expression and the feature vector in the spatio-temporal feature expression are selected respectively, merged in the feature dimension direction, and the fused features are processed by mean variance normalization;

[0094] S52, the fused features are input into the fusion network, the fusion network is a multi-layer fully connected neural network structure, is composed of a plurality of fully connected layers stacked in sequence, in each fully connected layer, the input feature vector is subjected to linear transformation and addition operation, the input vector is multiplied by the layer weight matrix, and a bias term is added, then the non-linear transformation is performed through the activation function to obtain the layer output feature, which is transmitted layer by layer, and after the transformation of all the fully connected layers, the fused feature vector is output;

[0095] S53, a feature weighting module is arranged inside the fusion network, and the weight parameters of each feature component are iteratively updated according to historical samples through a back propagation algorithm; in the forward inference stage, each component of the feature vector output by the fusion network is multiplied by the corresponding weight parameter to obtain a weighted fusion feature, and all weighted feature components are combined into a fusion feature.

[0096] The application introduces a new design in feature fusion and expression mechanism, and forms a multi-modal feature expression containing space, time and multi-frame dynamic information by splicing the dynamic structure features output by the time difference autoencoder and the global high-level features extracted by the space-time graph convolution network one by one at the feature level. In the fusion process, the difference features and the space-time features are strictly aligned and combined in the feature dimension for each detection sample, and the numerical scale and statistical distribution of the fusion features are unified through mean variance normalization, so as to provide a standardized input basis for subsequent deep fusion.

[0097] In the deep fusion stage, the fusion features are input into a multi-layer fully connected neural network, and linear transformation, bias addition and nonlinear activation operations are used in each layer of the network to complete information reorganization and feature mapping in a high-dimensional space layer by layer. In particular, a trainable feature weighting module is embedded in the fusion network structure, which supports end-to-end learning of the weight of each feature component through back propagation using historical sample data during model training. In the inference stage, the multi-modal information is adaptively adjusted and integrated by weighting each component of the fusion feature. The above mechanism breaks through the limitation of fixed feature splicing or simple weighting, so that the spatial structure, time sequence change and multi-frame difference information in the dynamic image of the pipe inner wall can be fully complementary, and the discriminability and adaptability of the downstream maintenance state evaluation are significantly improved.

[0098] In the embodiment, the S6 body comprises:

[0099] S61, input the fusion features into a state recognition module, in which two fully connected neural network layers are first arranged to perform multi-layer feature transformation and extraction on the input fusion features, an output grading unit is arranged in the last layer to map the features to a preset number of maintenance state categories, a score is output for each category, all output scores are compared, and the maintenance state category to which each image subsequence belongs is determined according to the maximum value classification rule; and the recognized maintenance state category label is assigned to the corresponding image subsequence.

[0100] S62, according to the pipeline inner wall maintenance state classification standard, processing the feature distribution output by the state recognition module, dividing the numerical range of the maintenance state feature into several intervals according to the pre-set classification parameters, comparing the feature distribution of each image sub-sequence with each classification interval, and determining the image sub-sequence as a specific maintenance state level according to the interval rule. The maintenance state level is set to multiple levels, covering the entire detection period section, including the highest level, the higher level, the middle level, the lower level and the lowest level, which correspond to different maintenance state descriptions respectively. The determination process of the section is executed according to the unified classification standard until all image sub-sequences complete the allocation of maintenance state level.

[0101] S63, for samples determined as critical state or located near the adjacent level boundary, based on the arrangement of image sub-sequences in time sequence, a time sequence consistency discrimination mechanism is adopted, specifically: comparing the state level of the current detection section with the state levels of the adjacent sections before and after the detection section, correcting the state level determination result, eliminating accidental anomalies, and finally outputting the maintenance state level label of each detection section.

[0102] The embodiment constructs a pipeline inner wall maintenance state automatic recognition and classification process with multi-modal feature fusion and intelligent classification as the core, covering key links such as feature input, deep classification, interval determination and time sequence consistency correction, forming a closed-loop determination chain covering the whole process of pipeline detection. A multi-layer fully connected neural network is used to map the fused features layer by layer, and combined with the pre-set multi-level state classification standard, each detection section is accurately attributed to the corresponding maintenance state level. For critical sections and boundary samples, a time sequence consistency discrimination mechanism is further introduced to dynamically correct the determination result, ensuring the logical coherence and stability of the classification label in the space-time distribution. It effectively breaks through the inconsistency of traditional static single-point classification and manual interpretation, significantly improves the intelligentization, batch processing and high precision level of pipeline state evaluation, and is suitable for automatic inspection and risk management in complex pipeline network scenarios, helping the transformation and upgrading of pipeline facility maintenance to scientific and intelligent.

[0103] In the embodiment, the state label output by the maintenance state recognition and classification evaluation module is stored together with the corresponding image sequence index to the data management module, and the data management module supports querying and retrieving historical detection results according to time, space section and state level.

[0104] The maintenance state label of each detection section is stored in the data management module in synchronization with the corresponding image sequence index, and efficient query and retrieval of historical detection results are supported in a multi-dimensional manner of time, space section, and state level. The mechanism realizes the whole-process structured archiving and flexible calling of detection results, significantly improves the convenience of data tracing, comparative analysis, and operation and maintenance decision-making, and provides data guarantee and technical foundation for pipeline health management and intelligent operation and maintenance.

[0105] Example 1

[0106] In order to verify the feasibility of the present application in actual engineering, the present application is applied to the annual inspection project of the underground water supply network in a coastal city. The pipe network covers the main urban area and several newly built residential areas, with a total length of more than 22 kilometers, and the distribution area includes old pipe corridors, urban trunk roads and some low-lying areas. The pipe diameter is between 600 mm and 1200 mm, and there are problems such as corrosion, scaling and coating peeling on the inner wall to varying degrees. Some sections are wet all year round and lack of lighting, making it difficult for manual inspection to achieve full coverage.

[0107] In this embodiment, a mobile camera inspection robot is used, which is equipped with a high-definition panoramic camera, lighting and positioning sensors. The robot moves at a constant speed along the axial direction of the pipeline, continuously collects dynamic image data of the inner wall, sets the collection speed to 0.1 m / s, the frame rate to 15 frames / s, and the longest continuous driving distance to 5 km in a single operation. The intensity of environmental light, the posture and position information of the equipment are collected in real time and synchronized with each frame of image data through time stamp alignment, realizing complete recording of multi-source information.

[0108] The system transmits the collected original image sequence to the backend server, and in accordance with the method of the present application, first performs adaptive histogram equalization, random cropping, random rotation and other enhancement processing on the image, and through denoising and edge preservation filtering, forms a standardized preprocessed image sequence. Subsequently, a time series difference autoencoder module is used to automatically capture the pixel-level changes of adjacent frames or local regions, and obtain a difference feature expression that can describe the dynamic changes of the inner wall details.

[0109] In the feature modeling stage, based on the spatial adjacency relationship of the pixel points and the inter-frame time sequence relationship of the image sub-sequence in each time window, a space-time graph structure is constructed, and a multi-layer space-time graph convolution network is used to recursively aggregate multi-layer space and time features, and output global high-level feature expression. The system splices the difference features and the space-time features, uses mean-variance normalization and multi-layer fully connected fusion network, integrates feature weighting mechanism, and obtains unified fusion features for downstream analysis.

[0110] The maintenance state recognition and grading evaluation module determines the maintenance state grade of each detection section according to the fused feature distribution and the pipeline industry grading standard, including multiple grades such as "perfect", "light attention", "moderate maintenance" and "severe treatment". For samples near the boundary value, the system corrects the abnormal grade in combination with the temporal consistency of the front and rear sections. Finally, all detection results and metadata are uniformly stored, supporting multi-dimensional queries according to time, space and grade.

[0111] In practical application, the system completes dynamic image acquisition and automatic evaluation of a 22-kilometer-long pipeline network. After comparison, the consistency rate of the key maintenance sections automatically labeled by the system and the subsequent manual review results reaches 93.5%. The inspection efficiency is significantly improved, and the full-automatic detection and grading of a single 5-kilometer section takes about 2 hours, which is more than 70% shorter than the previous manual interpretation. The pipeline maintenance department adjusts the annual maintenance plan accordingly, prioritizes the repair of high-risk sections, and significantly reduces the occurrence of sudden leaks and pipeline accidents.

[0112] The above embodiments show that the present application can adapt to various pipeline types and complex environments, realize large-scale, continuous and fine inner wall maintenance state evaluation, and provide a solid technical foundation for urban pipeline intelligent operation and risk control.

[0113] The above describes only the preferred embodiments of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can make equivalent replacements or changes to the technical solutions and inventive concepts of the present application within the technical scope disclosed by the present application, which should be covered within the protection scope of the present application.

Claims

1. A method for evaluating the maintenance state of the inner wall of a pipeline based on intelligent image recognition, characterized by, The method comprises the following steps: S1, moving at a uniform speed along the pipeline axis to obtain dynamic image data of the inner wall surface of the pipeline, and dividing the dynamic image data according to a preset time window to obtain a plurality of image subsequences; S2, performing image enhancement and preprocessing on each frame of image in the image subsequence to form a preprocessed image sequence; S3, inputting the preprocessed image sequence into a self-encoder model with a time difference mechanism to automatically capture the change information between adjacent frames or local areas, and outputting a differential feature expression distinguishing the detail change features of the inner wall surface of the pipeline through an adaptive feature coding process; S4, based on the image subsequence, taking the pixel points of multiple images as nodes, constructing a space-time graph structure combining spatial adjacency and inter-frame time sequence relationship, and extracting high-level feature expression reflecting the overall space-time distribution pattern and dynamic evolution law by using a space-time graph convolution network; S5, performing feature fusion processing on the differential feature expression and the high-level feature expression obtained above to obtain a multi-modal feature expression comprehensively reflecting the maintenance state of the inner wall of the pipeline; S6, based on the multi-modal feature expression, using a classification and state evaluation algorithm to identify the maintenance state in the image sequence, and grading the evaluation of the maintenance state of the inner wall of the pipeline in combination with the feature distribution in the subsequence, and finally outputting the maintenance state grade label of each detection section.

2. The method of claim 1, wherein, The S1 specifically comprises: setting the axial travel speed of the moving device and the sampling frame rate, synchronously collecting position, environmental light intensity and device posture information in each time window, and time stamping the information and the dynamic image data to obtain an image sequence covering the full length of the pipeline and being time-continuous.

3. The method of claim 1, wherein the method further comprises: The S2 specifically comprises: S21, performing adaptive histogram equalization operation on each frame of image, and generating an enhanced image sequence by using the method of random cropping and random angle rotation; S22, performing denoising processing and edge preserving filtering on the enhanced image sequence to form a preprocessed image sequence.

4. The method of claim 1, wherein, The S3 specifically comprises: S31, arranging the preprocessed image sequence in time sequence as the input of the self-encoder model; S32, performing pixel-level difference operation on each pair of frames for the adjacent two frames in the image sequence to generate a difference map reflecting the dynamic change of the microstructure of the inner wall surface of the pipeline; S33, concatenating the original frame and the difference map corresponding to the original frame in the channel dimension to form a composite feature input; S34, inputting the concatenated composite feature sequence into the self-encoder model, and the encoder part extracts space-time joint features through multi-layer convolution to obtain hidden layer representation, and the decoder part reconstructs the hidden layer features to obtain reconstruction output; S35, performing end-to-end optimization with the reconstruction error of the input and the output as the target, and finally outputting the model coding feature sequence as the differential feature expression distinguishing the detail change of the inner wall surface of the pipeline with high sensitivity: where I t (i,j) is the value of the original image at pixel (i,j) in the t-th frame, is the value of the reconstructed image at pixel (i,j) in the t-th frame by the autoencoder, H and W are the height and width of the image, T is the total number of frames of the dynamic image acquired in the pipeline, L rec is the reconstruction loss function.

5. The method of claim 1, wherein, The S4 specifically comprises: S41, taking all pixel points in the image subsequence as nodes, and constructing a space-time graph structure based on the spatial adjacency relationship and the time sequence corresponding relationship; S42, the spatial adjacency relationship adopts a pixel eight-neighborhood structure, the time sequence relationship connects the nodes in the same position in adjacent frames, and the node feature is initialized as the differential feature vector of the corresponding pixel; S43, apply a spatio-temporal graph convolution operation to the spatio-temporal graph structure, and a node feature updating rule is: where N S (v) is a spatial neighbor of node v, N T (v) is a temporal neighbor, is a learnable weight, and σ is an activation function, is the feature vector of node v in the lth layer of the spatio-temporal graph convolutional network, is the feature vector of node u in the lth layer of the spatio-temporal graph convolutional network, b (l) is the bias term of the lth layer; S44, in the spatio-temporal graph convolution network, a plurality of spatio-temporal graph convolution units are arranged, each unit is a layer, receives the node features output by the previous layer, and updates the features according to the spatial adjacency and time sequence adjacency relationship, each layer is stacked in a predetermined order, the output of the previous layer is used as the input of the next layer, and the layers are recursively propagated until the predetermined network depth, and the node features output by the last layer are subjected to a global aggregation operation to obtain high-level feature expression; S45, in each spatio-temporal graph convolution unit, a residual connection structure is arranged, the input features of the spatio-temporal graph convolution unit and the output features after convolution are added element by element to form a residual output, and the residual output is further subjected to feature normalization processing to standardize the values of each feature component, and all layers adopt this structure to finally obtain normalized spatio-temporal high-level feature expression as the input of the multi-modal feature fusion module.

6. The method of claim 1, wherein, The S5 specifically includes: S51, the difference feature sequence output by the time sequence difference autoencoder is spliced with the spatio-temporal high-level feature expression output by the spatio-temporal graph convolution network, for each sample to be fused, the feature vector of the sample to be fused in the difference feature expression and the feature vector in the spatio-temporal feature expression are selected respectively, combined in the feature dimension direction, and the fusion features are processed by mean variance normalization; S52, the fusion features are input into the fusion network, the fusion network is a multi-layer fully connected neural network structure, which is composed of a plurality of fully connected layers stacked in sequence, in each fully connected layer, the input feature vector is subjected to linear transformation and addition operation, the input vector is multiplied by the layer weight matrix and an offset item is added, then a nonlinear transformation is performed through an activation function to obtain layer output features, which are transmitted layer by layer, and after transformation through all the fully connected layers, a fusion feature vector is output; S53, a feature weighting module is arranged in the fusion network, the weight parameters of each feature component are iteratively updated according to historical samples by a back propagation algorithm, in the forward inference stage, each component of the feature vector output by the fusion network is multiplied by the corresponding weight parameter to obtain a weighted fusion feature, and all weighted feature components are combined into a fusion feature.

7. The method of claim 1, wherein the method further comprises: determining a location of the pipe based on the image; and determining a location of the pipe based on the image. 7 The S6 includes: S61, the fusion features are input into the state recognition module, in the state recognition module, two fully connected neural network layers are first arranged to perform multi-layer feature transformation and extraction on the input fusion features, an output grading unit is arranged in the last layer to map the features to a predetermined number of maintenance state categories, output a score for each category, compare all output scores, determine the maintenance state category to which each image subsequence belongs according to the maximum value classification rule, and assign the recognized maintenance state category label to the corresponding image subsequence; S62, according to the pipeline inner wall maintenance state classification standard, the feature distribution output by the state recognition module is processed, the numerical range of the maintenance state feature is divided into several intervals according to the pre-set classification parameters, the feature distribution of each image sub-sequence is compared with each classification interval, and the image sub-sequence is determined as a specific maintenance state level according to the interval rule. The maintenance state level is set to multiple levels, covering the entire detection period section, including the highest level, the higher level, the middle level, the lower level and the lowest level, which correspond to different maintenance state descriptions respectively. The determination process of the section is executed according to the unified classification standard until all image sub-sequences complete the allocation of the maintenance state level. S63, for the samples determined as the critical state or located near the adjacent level boundary, a time sequence consistency discrimination mechanism is adopted based on the arrangement of the image sub-sequences in time sequence, specifically: comparing the state level of the current detection section with the state levels of the adjacent sections before and after the detection section, correcting the state level determination result, eliminating the occasional abnormality, and finally outputting the maintenance state level label of each detection section.

8. The method of claim 1, wherein, The state label output by the maintenance state recognition and classification evaluation module is stored together with the corresponding image sequence index to the data management module, and the data management module supports querying and retrieving the historical detection results according to time, space section and state level.