A pipeline inner wall defect detection method and system based on image recognition

By using image decoupling and deep feature fusion, the problem of insufficient separation between structural information and texture details in pipeline inner wall defect detection in existing technologies is solved, enabling accurate detection and quantification of pipeline inner wall defects and generating structured inspection reports.

CN121998984BActive Publication Date: 2026-06-30BAOJI HUALAN NEW MATERIAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BAOJI HUALAN NEW MATERIAL TECH CO LTD
Filing Date
2026-04-09
Publication Date
2026-06-30

AI Technical Summary

Technical Problem

Existing technologies for detecting defects in the inner walls of pipelines fail to effectively separate structural information from texture details. This results in the high-frequency detail features related to defects being submerged or structural information interfering with texture analysis. Furthermore, they fail to accurately quantify the actual geometric parameters and spatial pose of defects, making it difficult to meet the complex image detection requirements of the inner walls of pipelines.

Method used

Image sequences of the structure layer and texture layer are obtained by image decoupling. After field-of-view consistency processing, the enhanced image sequence is mapped to a preset two-dimensional plane. Deep feature fusion is performed by combining multi-scale contextual information, and low-confidence regions are iteratively optimized to finally generate a structured defect detection report.

Benefits of technology

It enables high-quality detection of defects on the inner wall of pipelines, accurately captures defect characteristics, and generates comprehensive and reliable structured inspection reports, providing a basis for pipeline maintenance decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121998984B_ABST
    Figure CN121998984B_ABST
Patent Text Reader

Abstract

This invention relates to the field of image recognition technology, and discloses a method and system for detecting defects in the inner wall of pipes based on image recognition. The method includes: decoupling the original image sequence of the inner wall of the pipe to obtain a structural layer image sequence and a texture layer image sequence; unifying the structural layer image sequence and the texture layer image sequence to obtain an enhanced image sequence; mapping the enhanced image sequence onto a preset two-dimensional plane to obtain a two-dimensional plane unfolded map to obtain a candidate region mask; performing deep feature fusion on multi-scale contextual information to obtain a discriminative deep feature vector; performing preliminary defect classification and confidence assessment on the discriminative deep feature vector, and optimizing low-confidence regions based on the confidence assessment results to obtain defect category labels and pixel-level semantic segmentation contours; determining the actual geometric parameters and spatial pose of the defect to generate a structured defect detection report. This invention can improve the efficiency of defect detection in the inner wall of pipes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image recognition technology, and in particular to a method and system for detecting defects in the inner wall of pipes based on image recognition. Background Technology

[0002] Existing technologies have significant shortcomings in the image preprocessing and enhancement stages of pipeline inner wall defect detection. They fail to decouple the structural and texture layers of the original pipeline inner wall image sequence, relying solely on a single image enhancement method. This fails to effectively separate structural information and texture details from the image, resulting in the submergence of high-frequency detail features related to defects or structural information interfering with texture analysis. Furthermore, the lack of viewpoint consistency processing between the structural and texture layer images, coupled with simple fusion of different modal images, easily leads to image registration bias. Consequently, the enhanced image cannot accurately reflect the true state of the pipeline inner wall, making it difficult to provide a high-quality image foundation for subsequent defect detection and failing to meet the complex image detection requirements of pipeline inner walls.

[0003] Existing technologies fail to combine pipeline geometric parameters with camera imaging parameters to map the enhanced image into a two-dimensional planar unfolded map, performing defect detection directly on the original image. This leads to deformation and missed detection of defect areas due to the curved shape of the pipeline surface. Furthermore, the technologies do not perform deep feature fusion on multi-scale contextual information of candidate regions, extracting only local single features, which makes it difficult to capture the discriminative features of defects, resulting in low defect classification accuracy. Additionally, the technologies do not iteratively optimize low-confidence regions or accurately quantify the actual geometric parameters and spatial pose of defects, simply labeling the defect location without generating a structured inspection report, thus failing to meet the actual needs of accurate assessment and maintenance of pipeline inner wall defects. Summary of the Invention

[0004] This invention provides a method and system for detecting defects in the inner wall of pipes based on image recognition, in order to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides a method for detecting defects in the inner wall of a pipe based on image recognition, comprising:

[0006] S1. Obtain the original image sequence of the inner wall of the pipe, and perform image decoupling on the original image sequence to obtain the structure layer image sequence and texture layer image sequence of the original image sequence;

[0007] S2. Perform view field consistency on the structure layer image sequence and the texture layer image sequence to obtain the enhanced image sequence of the inner wall of the pipe;

[0008] S3. Based on the geometric parameters of the inner wall of the pipe and the camera imaging parameters, the enhanced image sequence is mapped onto a preset two-dimensional plane to obtain a two-dimensional planar unfolded image of the inner wall of the pipe, and the candidate defect region is coarsely extracted from the two-dimensional planar unfolded image to obtain a candidate region mask of the inner wall of the pipe.

[0009] S4. Perform deep feature fusion on the multi-scale context information corresponding to the candidate region mask to obtain the discriminative deep feature vector of the inner wall of the pipe;

[0010] S5. Perform preliminary defect classification and confidence assessment on the discriminative deep feature vector, and iteratively optimize the low confidence region based on the confidence assessment results to obtain the defect category label and pixel-level semantic segmentation contour of the inner wall of the pipe.

[0011] S6. Based on the defect category label, the pixel-level semantic segmentation contour and the image-space mapping relationship, determine the actual geometric parameters and spatial pose of the defects on the inner wall of the pipe, so as to generate a structured defect detection report of the inner wall of the pipe.

[0012] In a preferred embodiment, the step of acquiring the original image sequence of the pipe inner wall and decoupling the original image sequence to obtain the structure layer image sequence and texture layer image sequence of the original image sequence includes:

[0013] Obtain the original image sequence of the inner wall of the pipe;

[0014] The original images in the original image sequence are decomposed in the frequency domain to obtain the low-frequency component image and the high-frequency component image of the original image sequence;

[0015] The low-frequency component image is used as the structural layer image sequence of the original image sequence;

[0016] The high-frequency component image is subjected to a non-downsampled contour wave transform to obtain the sub-band coefficient set of the high-frequency component image;

[0017] The subband coefficient set is subjected to coefficient modulation to obtain the enhanced high-frequency subband coefficient set.

[0018] Based on the enhanced high-frequency subband coefficient set, a non-downsampled contour wave inverse transform is performed to generate a reconstructed high-frequency detail image of the original image sequence, and the reconstructed high-frequency detail image is used as the texture layer image sequence of the original image sequence.

[0019] In a preferred embodiment, the step of performing viewport unification on the structural layer image sequence and the texture layer image sequence to obtain the enhanced image sequence of the pipe inner wall includes:

[0020] Extract feature point sets from the frame images in the structure layer image sequence and the corresponding frame images in the texture layer image sequence;

[0021] Based on the feature point set, multimodal image registration is performed on the structural layer image sequence and the texture layer image sequence to obtain the spatial transformation parameter set of the inner wall of the pipe;

[0022] Based on the spatial transformation parameter set, an affine transformation is performed on the texture layer image sequence to obtain a registered texture layer image sequence.

[0023] The structural layer image sequence and the registration texture layer image sequence are fused to obtain an enhanced image sequence of the inner wall of the pipe.

[0024] In a preferred embodiment, the step of mapping the enhanced image sequence onto a preset two-dimensional plane based on the geometric parameters of the pipe inner wall and the camera imaging parameters to obtain a two-dimensional planar unfolded image of the pipe inner wall includes:

[0025] Based on the focal length, principal point coordinates, and distortion coefficients in the camera imaging parameters, a first mapping relationship is established between the image pixel coordinates in the enhanced image sequence and the three-dimensional spatial coordinates of the inner wall of the pipe, based on the pinhole imaging model.

[0026] Based on the geometric parameters of the inner wall of the pipe and the first mapping relationship, a coordinate transformation rule is established to transform the three-dimensional spatial coordinates to a preset two-dimensional plane.

[0027] Based on the coordinate transformation rules, the enhanced image sequence is subjected to coordinate transformation, and the projection coordinate set of the pixels in the enhanced image sequence on the preset two-dimensional plane is calculated;

[0028] Based on the projection coordinate set, the enhanced image sequence is reprojected to obtain a two-dimensional planar unfolded view of the inner wall of the pipe.

[0029] In a preferred embodiment, the formula for calculating the projected coordinate set is:

[0030] ;

[0031] In the formula, The projection coordinates of the pixel on the preset two-dimensional plane. The camera focal length is one of the camera imaging parameters. The vertical image coordinates of the pixel in the original enhanced image. This represents the vertical coordinate of the principal point in the image from the camera imaging parameters. Here are the horizontal image coordinates of the pixel in the original enhanced image. This represents the lateral coordinates of the principal point in the image from the camera imaging parameters. The pipe radius is one of the geometric parameters of the inner wall of the pipe. The preset constant correction coefficient is used. It is used to fine-tune the equivalent focal length in actual imaging to adapt to changes in the curvature of the pipe.

[0032] In a preferred embodiment, the step of coarsely extracting candidate defect regions from the two-dimensional planar unfolded image to obtain a candidate region mask for the inner wall of the pipe includes:

[0033] According to the coordinate transformation rules of the inner wall of the pipe, the image sequence of the structural layer is mapped to the preset two-dimensional plane to obtain the structural reference diagram of the inner wall of the pipe;

[0034] By comparing the structural reference diagram with the two-dimensional planar unfolded diagram pixel by pixel, a significant difference diagram of the inner wall of the pipe is obtained.

[0035] Adaptive threshold segmentation is performed on the significant difference map to obtain a preliminary binary map of candidate regions of the significant difference map;

[0036] The texture layer image sequence is mapped onto the preset two-dimensional plane to generate the edge structure map of the inner wall of the pipe;

[0037] Based on the edge structure map, the preliminary candidate region binary map is morphologically reconstructed to obtain the optimized candidate region binary map of the pipe inner wall;

[0038] The bounding rectangle of the connected region in the binary graph of the optimized candidate region is used as the candidate region to generate the candidate region mask of the inner wall of the pipe.

[0039] In a preferred embodiment, the step of performing deep feature fusion on the multi-scale context information corresponding to the candidate region mask to obtain the discriminative deep feature vector of the pipe inner wall includes:

[0040] Based on the candidate region mask, extract the corresponding candidate image block set in the two-dimensional planar unfolded image;

[0041] For the candidate image blocks in the candidate image block set, construct the image block-context information pair set of the inner wall of the pipe;

[0042] Feature extraction is performed on the image blocks and corresponding context information in the image block-context information pair set to obtain the local depth feature vector and multi-scale context depth feature vector set of the inner wall of the pipe;

[0043] The local depth feature vector and the multi-scale context depth feature vector set are weighted and fused to obtain the candidate image block depth feature vector of the inner wall of the pipe;

[0044] The depth feature vectors of the candidate image blocks are aggregated to generate the discriminative depth feature vector of the inner wall of the pipe.

[0045] In a preferred embodiment, the step of performing preliminary defect classification and confidence assessment on the discriminative deep feature vector, and iteratively optimizing low-confidence regions based on the confidence assessment results to obtain defect category labels and pixel-level semantic segmentation contours of the pipe inner wall includes:

[0046] The discriminative deep feature vector is input into the classification network of the inner wall of the pipe to obtain the preliminary defect category prediction and prediction confidence of the inner wall of the pipe.

[0047] Regions with predicted confidence levels below a set threshold are marked as low-confidence regions;

[0048] Extract the multi-level features corresponding to the low-confidence region in the preset multi-scale depth feature map, and fuse them to generate an enhanced feature representation;

[0049] Based on the enhanced feature representation, the low-confidence region is reclassified to obtain the updated defect category prediction and the updated prediction confidence of the inner wall of the pipe.

[0050] By integrating the updated defect category predictions, the defect category labels of the inner wall of the pipe are obtained;

[0051] Based on the defect category label, pixel-level semantic segmentation is performed on the two-dimensional planar unfolded image to obtain the pixel-level semantic segmentation contour of the inner wall of the pipe.

[0052] In a preferred embodiment, determining the actual geometric parameters and spatial pose of the defects on the inner wall of the pipe based on the defect category label, the pixel-level semantic segmentation contour, and the image-space mapping relationship, to generate a structured defect detection report for the inner wall of the pipe, includes:

[0053] Based on the image-space mapping relationship, the pixel-level semantic segmentation contour is mapped to the three-dimensional spatial coordinates of the inner wall of the pipe, thereby obtaining a three-dimensional spatial representation of the defect in the inner wall of the pipe.

[0054] Based on the three-dimensional spatial representation of the defect, the spatial morphological features of the defect in the inner wall of the pipe are extracted, the centroid coordinates and principal axis direction of the defect are determined, and the spatial pose of the defect in the inner wall of the pipe is obtained.

[0055] Based on the three-dimensional spatial representation of the defect, the surface shape and volume distribution of the defect are analyzed geometrically to obtain the actual geometric parameters of the defect on the inner wall of the pipe.

[0056] By combining the defect category label, the defect spatial pose, and the actual geometric parameters of the defect, a structured defect detection report for the inner wall of the pipeline is obtained.

[0057] To address the above problems, the present invention also provides an image recognition-based pipe inner wall defect detection system, the system comprising:

[0058] An image decoupling module is used to acquire the original image sequence of the inner wall of the pipe and decouple the original image sequence to obtain a preprocessed image sequence of the original image sequence.

[0059] A visual enhancement module is used to perform field-of-view consistency on the structural layer image sequence and the texture layer image sequence to obtain an enhanced image sequence of the inner wall of the pipe;

[0060] The image unfolding and coarse inspection module is used to map the enhanced image sequence onto a preset two-dimensional plane according to the geometric parameters of the inner wall of the pipe and the camera imaging parameters to obtain a two-dimensional unfolded image of the inner wall of the pipe, and to coarsely extract the candidate region of the defect from the two-dimensional unfolded image to obtain a candidate region mask of the inner wall of the pipe.

[0061] The feature fusion module is used to perform deep feature fusion on the multi-scale context information corresponding to the candidate region mask to obtain the discriminative deep feature vector of the inner wall of the pipe.

[0062] The defect classification module is used to perform preliminary defect classification and confidence assessment on the discriminative deep feature vector, and to iteratively optimize the low-confidence region based on the confidence assessment result, so as to obtain the defect category label and pixel-level semantic segmentation contour of the inner wall of the pipe.

[0063] The defect quantification report module is used to calculate the actual geometric parameters and spatial pose of the defects on the inner wall of the pipe based on the defect category label, the pixel-level semantic segmentation contour and the image-space mapping relationship, so as to generate a structured defect detection report of the inner wall of the pipe.

[0064] Compared with the prior art, the present invention has the following beneficial effects:

[0065] 1. This invention provides a high-quality image foundation for pipeline inner wall defect detection through image layer enhancement and precise unfolding. After acquiring the original image sequence, frequency domain decomposition is performed to decouple the structural and texture layer image sequences. These sequences are then fused through field-of-view consistency registration to generate an enhanced image sequence, preserving defect-related structural and texture details. Combining pipeline geometric parameters and camera imaging parameters, the enhanced image is mapped into a two-dimensional planar unfolded image, eliminating image distortion caused by the pipeline surface. Finally, through difference comparison and morphological reconstruction, candidate defect regions are coarsely extracted to obtain a precise candidate region mask, laying the foundation for subsequent feature extraction.

[0066] 2. This invention significantly improves the accuracy and quantification capabilities of defect detection by leveraging deep feature fusion and iterative optimization. It deeply fuses multi-scale contextual information from candidate region masks to generate discriminative deep feature vectors, accurately capturing the essential features of defects. A classification network is used for initial classification and confidence assessment, followed by iterative optimization of low-confidence regions to obtain accurate defect category labels and pixel-level semantic segmentation contours. Based on image-space mapping, the segmentation contours are mapped to three-dimensional space, quantifying the actual geometric parameters and spatial pose of the defects, generating a structured inspection report, and providing comprehensive and reliable decision-making support for pipeline maintenance. Attached Figure Description

[0067] Figure 1 This is a schematic flowchart of a pipe inner wall defect detection method based on image recognition, provided in an embodiment of the present invention.

[0068] Figure 2 A functional block diagram of a pipeline inner wall defect detection system based on image recognition is provided in an embodiment of the present invention;

[0069] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0070] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0071] This application provides a method for detecting defects in the inner wall of a pipe based on image recognition. The execution entity of this method includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the method for detecting defects in the inner wall of a pipe based on image recognition can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cluster of cloud servers. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0072] Reference Figure 1 The diagram shown is a flowchart illustrating a pipe inner wall defect detection method based on image recognition, according to an embodiment of the present invention. In this embodiment, the pipe inner wall defect detection method based on image recognition includes:

[0073] S1. Obtain the original image sequence of the inner wall of the pipe, and perform image decoupling on the original image sequence to obtain the structure layer image sequence and texture layer image sequence of the original image sequence;

[0074] In this embodiment of the invention, the step of acquiring the original image sequence of the pipe inner wall and performing image decoupling on the original image sequence to obtain the structure layer image sequence and texture layer image sequence of the original image sequence includes:

[0075] Obtain the original image sequence of the inner wall of the pipe;

[0076] The original images in the original image sequence are decomposed in the frequency domain to obtain the low-frequency component image and the high-frequency component image of the original image sequence;

[0077] The low-frequency component image is used as the structural layer image sequence of the original image sequence;

[0078] The high-frequency component image is subjected to a non-downsampled contour wave transform to obtain the sub-band coefficient set of the high-frequency component image;

[0079] The subband coefficient set is subjected to coefficient modulation to obtain the enhanced high-frequency subband coefficient set.

[0080] Based on the enhanced high-frequency subband coefficient set, a non-downsampled contour wave inverse transform is performed to generate a reconstructed high-frequency detail image of the original image sequence, and the reconstructed high-frequency detail image is used as the texture layer image sequence of the original image sequence.

[0081] The detection device, equipped with an image acquisition device, enters the pipeline and moves the device according to the preset acquisition interval and path to continuously capture image information of different positions and angles on the inner wall of the pipeline. All captured images are then organized in chronological order of acquisition time to form the original image sequence of the inner wall of the pipeline.

[0082] The frequency domain decomposition method is used to process each original image in the original image sequence. First, the image is transformed from the spatial domain to the frequency domain. The low-frequency and gently changing parts of the image are filtered out by a low-pass filter. These parts correspond to the overall outline and basic shape of the image. At the same time, the high-frequency and drastically changing parts of the image are filtered out by a high-pass filter. These parts correspond to the details and texture information of the image. The low-frequency component image and high-frequency component image of the original image sequence are obtained respectively.

[0083] The low-frequency component image obtained after frequency domain decomposition is directly determined as the structural layer image sequence of the original image sequence. This sequence completely preserves the core structural information such as the overall outline and basic shape of each original image, and can clearly present the overall condition of the pipe inner wall.

[0084] A non-downsampled contour wave transform is performed on the high-frequency component image. First, the high-frequency component image is decomposed into multiple scales through a non-downsampled pyramid filter bank without downsampling to keep the image pixel size unchanged, resulting in low-frequency coefficients and multiple high-frequency coefficients at different scales. Then, a non-downsampled directional filter bank is used to further decompose the high-frequency coefficients at each scale into sub-band coefficients in multiple directions. These directional sub-band coefficients can accurately capture the edge and texture details in the image. The sub-band coefficients of all scales and directions are integrated to form the sub-band coefficient set of the high-frequency component image.

[0085] For each subband coefficient in the subband coefficient set, the coefficient value is adjusted according to the importance of the corresponding image details using a fixed modulation rule. This enhances the intensity of coefficients that contribute significantly to the image texture and suppresses coefficients corresponding to invalid noise, so that the adjusted subband coefficients can better highlight the texture features of the pipe wall. Finally, the enhanced high-frequency subband coefficient set of the subband coefficient set is obtained.

[0086] Based on the enhanced high-frequency subband coefficient set, an inverse non-subsampled contour wave transform is performed. First, the subband coefficients of each direction are inversely processed by a non-subsampled directional filter bank to restore the high-frequency coefficients of each scale. Then, the low-frequency coefficients of multiple scales and the restored high-frequency coefficients are integrated and reconstructed by a non-subsampled pyramid filter bank to recover the complete high-frequency image. This image is the reconstructed high-frequency detail image of the original image sequence, which is determined as the texture layer image sequence of the original image sequence.

[0087] The beneficial effect is that by using professional image acquisition equipment to capture original images of different positions and angles of the inner wall of the pipeline and organizing them into an original image sequence according to the acquisition time sequence, it provides comprehensive original image data support for subsequent defect detection, ensuring that image information of any area of ​​the inner wall of the pipeline is not missed.

[0088] The original image is decomposed in the frequency domain to accurately separate the low-frequency component image that reflects the overall outline and basic shape of the inner wall of the pipe and the high-frequency component image that carries defect-related information such as texture details and edge features. This achieves effective separation of image structure information and texture details, avoiding mutual interference between the two types of information and affecting the accuracy of defect recognition.

[0089] By directly using low-frequency component images as structural layer image sequences, the core structural features of the pipe's inner wall are fully preserved, providing a stable and reliable structural benchmark for subsequent image registration and defect localization, and ensuring accurate control of the overall pipe morphology during defect detection.

[0090] Non-subsampled contour wave transform is performed on high-frequency component images. Without changing the image pixel size, it captures subtle textures and edge details in high-frequency components through multi-scale and multi-directional decomposition, generating a sub-band coefficient set containing rich defect-related information, laying the foundation for texture detail enhancement.

[0091] The coefficient set of the subnet is modulated to selectively enhance the strength of coefficients that are important for defect identification, while suppressing coefficients corresponding to invalid noise. This makes the texture details related to defects more prominent and improves the recognition of defect information in subsequent processing.

[0092] Based on the enhanced high-frequency subband coefficient set, non-subsampled contour wave inverse transform is performed to reconstruct high-frequency detail images, which are then used as texture layer image sequences to enhance and restore defect-related texture details. This enables the texture layer images to clearly present the subtle defect features of the pipe's inner wall. Combined with the structural layer images, this provides high-quality, high-recognition image data for subsequent field-of-view consistency and defect detection, effectively solving the problem of defect feature overwhelmed or structural information interfered by traditional single image enhancement.

[0093] S2. Perform view field consistency on the structure layer image sequence and the texture layer image sequence to obtain the enhanced image sequence of the inner wall of the pipe;

[0094] In this embodiment of the invention, the step of performing viewport unification on the structural layer image sequence and the texture layer image sequence to obtain the enhanced image sequence of the pipe inner wall includes:

[0095] Extract feature point sets from the frame images in the structure layer image sequence and the corresponding frame images in the texture layer image sequence;

[0096] Based on the feature point set, multimodal image registration is performed on the structural layer image sequence and the texture layer image sequence to obtain the spatial transformation parameter set of the inner wall of the pipe;

[0097] Based on the spatial transformation parameter set, an affine transformation is performed on the texture layer image sequence to obtain a registered texture layer image sequence.

[0098] The structural layer image sequence and the registration texture layer image sequence are fused to obtain an enhanced image sequence of the inner wall of the pipe.

[0099] Feature point detection methods are used to process the structural layer image sequence and the texture layer image sequence respectively. For each frame image in the structural layer image sequence, the image pixel region is traversed to identify pixels with drastic grayscale changes and obvious recognizability. These pixels include image edge intersections, contour turning points, etc. The same operation is performed on the images in the texture layer image sequence that correspond to the structural layer images to accurately capture key pixels in texture details. These identified pixels are then organized to form feature point sets for the structural layer image sequence frames and feature point sets for the texture layer image sequence corresponding to the texture layer image sequence frames.

[0100] Using the feature point set of the structural layer image sequence as a reference, the gray-level similarity and spatial position correlation of each feature point in the feature point set of the corresponding frame image of the texture layer image sequence with the feature points in the reference are calculated. By comparing point by point, matching feature point pairs are found between the two sets of feature point sets. Based on these matching feature point pairs, the differences in spatial position such as offset, rotation, and scaling of the texture layer image relative to the structural layer image are analyzed. Then, the parameters that can eliminate such spatial differences are determined. The parameters corresponding to all frames are summarized to obtain the spatial transformation parameter set of the pipe inner wall.

[0101] According to the parameter requirements of each frame image in the spatial transformation parameter set, affine transformation processing is performed on each frame image in the texture layer image sequence. The spatial position of the image is adjusted by translation operation to align it with the position of the corresponding structural layer image. The angular deviation of the image is corrected by rotation operation, and the size ratio of the image is adjusted by scaling operation to ensure that the texture layer image is completely matched with the structural layer image in spatial dimension. After these transformation processing, the registered texture layer image sequence is obtained.

[0102] An image fusion method is used to process the structural layer image sequence and the registered texture layer image sequence. The core structural information, such as the overall outline and basic shape, of the structural layer image is extracted frame by frame. At the same time, the fine information such as the detailed texture and edge features of the registered texture layer image is extracted. The two types of information are fused at the pixel level according to a fixed ratio. The fused image retains the clear overall shape of the structural layer image and superimposes the rich detailed features of the registered texture layer image, making the image information of the inner wall of the pipe more complete and clearer. After the fusion processing of all images is completed frame by frame, the enhanced image sequence of the inner wall of the pipe is obtained.

[0103] The beneficial effect is that by extracting the feature point sets of the corresponding frame images of the structure layer and texture layer, the key pixels with recognizability in the two types of images can be accurately captured, providing a reliable feature benchmark for subsequent image registration and ensuring that the registration process has a clear corresponding basis.

[0104] Multimodal image registration is carried out based on feature point sets. By comparing the gray-level similarity and spatial correlation of feature points in two types of images, the spatial difference between the texture layer and the structure layer can be accurately analyzed, and then the set of spatial transformation parameters to eliminate the difference can be determined, effectively solving the problem of registration deviation between different modal images.

[0105] Based on the spatial transformation parameter set, an affine transformation is performed on the texture layer image sequence. Through operations such as translation, rotation, and scaling, the texture layer image is fully aligned with the structure layer image in terms of spatial position, angle, and size, ensuring that the field of view of the two types of images remains consistent, thus laying the foundation for spatial matching for subsequent fusion.

[0106] Pixel-level fusion of the structural layer and the registered texture layer image sequence preserves the clear overall outline and basic shape of the structural layer image while overlaying the rich details and edge features of the texture layer image. This allows the enhanced image to have both complete structure and fine details, significantly improving the integrity and clarity of the image information. This provides high-quality image data support for subsequent defect detection and avoids the information loss or interference problems caused by traditional simple fusion.

[0107] S3. Based on the geometric parameters of the inner wall of the pipe and the camera imaging parameters, the enhanced image sequence is mapped onto a preset two-dimensional plane to obtain a two-dimensional planar unfolded image of the inner wall of the pipe, and the candidate defect region is coarsely extracted from the two-dimensional planar unfolded image to obtain a candidate region mask of the inner wall of the pipe.

[0108] In this embodiment of the invention, the step of mapping the enhanced image sequence onto a preset two-dimensional plane based on the geometric parameters of the pipe inner wall and the camera imaging parameters to obtain a two-dimensional planar unfolded image of the pipe inner wall includes:

[0109] Based on the focal length, principal point coordinates, and distortion coefficients in the camera imaging parameters, a first mapping relationship is established between the image pixel coordinates in the enhanced image sequence and the three-dimensional spatial coordinates of the inner wall of the pipe, based on the pinhole imaging model.

[0110] Based on the geometric parameters of the inner wall of the pipe and the first mapping relationship, a coordinate transformation rule is established to transform the three-dimensional spatial coordinates to a preset two-dimensional plane.

[0111] Based on the coordinate transformation rules, the enhanced image sequence is subjected to coordinate transformation, and the projection coordinate set of the pixels in the enhanced image sequence on the preset two-dimensional plane is calculated;

[0112] Based on the projection coordinate set, the enhanced image sequence is reprojected to obtain a two-dimensional planar unfolded view of the inner wall of the pipe.

[0113] The formula for calculating the projected coordinate set is:

[0114] ;

[0115] In the formula, The projection coordinates of the pixel on the preset two-dimensional plane. The camera focal length is one of the camera imaging parameters. The vertical image coordinates of the pixel in the original enhanced image. This represents the vertical coordinate of the principal point in the image from the camera imaging parameters. Here are the horizontal image coordinates of the pixel in the original enhanced image. This represents the lateral coordinates of the principal point in the image from the camera imaging parameters. The pipe radius is one of the geometric parameters of the inner wall of the pipe. The preset constant correction coefficient is used. It is used to fine-tune the equivalent focal length in actual imaging to adapt to changes in the curvature of the pipe.

[0116] The step of coarsely extracting candidate defect regions from the two-dimensional planar unfolded image to obtain a candidate region mask for the inner wall of the pipe includes:

[0117] According to the coordinate transformation rules of the inner wall of the pipe, the image sequence of the structural layer is mapped to the preset two-dimensional plane to obtain the structural reference diagram of the inner wall of the pipe;

[0118] By comparing the structural reference diagram with the two-dimensional planar unfolded diagram pixel by pixel, a significant difference diagram of the inner wall of the pipe is obtained.

[0119] Adaptive threshold segmentation is performed on the significant difference map to obtain a preliminary binary map of candidate regions of the significant difference map;

[0120] The texture layer image sequence is mapped onto the preset two-dimensional plane to generate the edge structure map of the inner wall of the pipe;

[0121] Based on the edge structure map, the preliminary candidate region binary map is morphologically reconstructed to obtain the optimized candidate region binary map of the pipe inner wall;

[0122] The bounding rectangle of the connected region in the binary graph of the optimized candidate region is used as the candidate region to generate the candidate region mask of the inner wall of the pipe.

[0123] By deeply analyzing key information such as lens focal length, imaging sensor size, and camera mounting position included in camera imaging parameters, the correlation logic between this information and image pixel distribution is clarified. By establishing the correspondence between the row and column positions of pixels on the image plane and their actual three-dimensional positions on the pipe inner wall, the three-dimensional spatial position of the pipe inner wall that each pixel coordinate can accurately point to is determined, thereby forming the first mapping relationship between image pixel coordinates and the three-dimensional spatial coordinates of the pipe inner wall in the enhanced image sequence.

[0124] A comprehensive analysis of the geometric parameters of the pipe's inner wall was conducted, covering core data such as the pipe's inner diameter, length, and cross-sectional shape. Based on the established first mapping relationship, the conversion logic between three-dimensional spatial coordinates and preset two-dimensional plane coordinates was analyzed. Considering the curved surface characteristics of the pipe's inner wall, a specific method for unfolding and mapping the curved three-dimensional coordinates to the plane according to fixed rules was formulated. The calculation method for the two-dimensional plane coordinates corresponding to each three-dimensional spatial coordinate was clarified, forming a complete set of coordinate transformation rules.

[0125] Based on the coordinate transformation rules, each image in the enhanced image sequence is processed one by one. For each pixel in the image, the corresponding three-dimensional spatial coordinates are found through the first mapping relationship according to its corresponding pixel coordinates. Then, according to the calculation method in the coordinate transformation rules, the position of the three-dimensional spatial coordinates on the preset two-dimensional plane is determined. The corresponding plane positions of all pixels are sorted and summarized to obtain the set of projection coordinates of the pixels in the enhanced image sequence on the preset two-dimensional plane.

[0126] According to the projection coordinates of each pixel in the projection coordinate set, the pixels in the enhanced image sequence are reprojected one by one onto a preset two-dimensional plane, maintaining the relative positional relationship between pixels and the integrity of image information. This ensures that the reprojected image can fully restore all the detailed features of the pipe's inner wall, while eliminating image distortion caused by the pipe's curved surface, ultimately forming a two-dimensional planar unfolded image that can clearly present the entire inner wall of the pipe.

[0127] The camera's focal length comes from the camera's imaging parameters. It is an inherent hardware parameter of the camera itself and is clearly calibrated when the camera leaves the factory. It can be directly obtained through the camera's technical specification manual or by testing with professional equipment.

[0128] The longitudinal and transverse coordinates of the principal point in the image are derived from the camera imaging parameters. They are the image coordinates corresponding to the center of the camera imaging sensor and are determined through camera calibration experiments. Specifically, multiple sets of calibration board images with known three-dimensional coordinates are captured, and these two coordinate values ​​are calculated by analyzing the imaging positions of the feature points of the calibration board in the images.

[0129] The pipe radius is derived from the geometric parameters of the pipe's inner wall. It is obtained by consulting the pipe's design drawings or by actually measuring the inner wall of the pipe using a laser rangefinder. During the measurement, distances are measured at multiple locations at different cross-sections of the pipe, and the fixed value of the measurement result is taken as the final pipe radius.

[0130] The adaptive adjustment coefficient is a preset value. It is determined through calibration using a large amount of experimental data, taking into account the imaging characteristics of the inner wall image of the pipe, the pipe radius, and the camera installation distance, to ensure that the projection coordinate calculation can be adapted to different pipe inspection scenarios.

[0131] The vertical and horizontal image coordinates of a pixel in the original enhanced image are the inherent positional identifiers of the pixels on the image plane in the enhanced image sequence. By reading the pixel matrix of the enhanced image using an image reading tool, the row and column positions of each pixel in the matrix correspond to the vertical and horizontal image coordinates, respectively.

[0132] The significance of this formula is to establish a precise mapping relationship between the coordinates of pixels in the original enhanced image and their projected coordinates on a preset two-dimensional plane. By integrating camera imaging parameters and pipe inner wall geometric parameters, the three-dimensional spatial position information corresponding to pixels in the enhanced image is converted into coordinates on a preset two-dimensional plane according to a fixed calculation logic.

[0133] During the calculation, the difference between the horizontal coordinates of the pixel in the original enhanced image and the horizontal coordinates of the image principal point is first divided by the sum of the products of the camera focal length, the adaptive adjustment coefficient, the specific function, and the pipeline radius to obtain a ratio. The arctangent operation is then performed on this ratio, and the result is multiplied by the camera focal length to obtain the horizontal projection coordinates on the preset two-dimensional plane.

[0134] Simultaneously, by multiplying the difference between the vertical coordinates of the pixel in the original enhanced image and the vertical coordinates of the image principal point by the pipe radius, a product result is obtained. Then, the square of the difference between the horizontal coordinates of the pixel in the original enhanced image and the horizontal coordinates of the image principal point is calculated, along with the square of the sum of the products of the camera focal length, the adaptive adjustment coefficient, the specific function, and the pipe radius. The square root of this sum is then taken, and finally, the previously obtained product result is divided by this square root to obtain the vertical projection coordinates on the preset two-dimensional plane.

[0135] Through this complete calculation process, each pixel in the enhanced image is accurately mapped to its corresponding position on a preset two-dimensional plane, providing accurate coordinates for the subsequent generation of a two-dimensional planar unfolded diagram of the pipe's inner wall.

[0136] Using the same coordinate transformation rules as mapping the enhanced image sequence to the preset two-dimensional plane, the structural layer image sequence is mapped frame by frame to the preset two-dimensional plane. Then, the mapped structural layer image sequence is spatially aligned with the two-dimensional plane unfolded image so that the pixel positions of the two are completely matched on the preset two-dimensional plane, thus obtaining the structural reference image of the inner wall of the pipe.

[0137] Using the structural reference image as the baseline image, the pixel grayscale values ​​of the corresponding positions in the two-dimensional planar unfolded image are compared pixel by pixel. The grayscale difference of each pixel position is calculated. Pixels with grayscale differences greater than a fixed baseline value are marked as difference pixels, and pixels with grayscale differences less than or equal to the fixed baseline value are marked as non-difference pixels. All marking results are arranged according to the original position of the pixels to obtain the significant difference map of the structural reference image.

[0138] Adaptive threshold segmentation is performed on the significant difference map. First, the overall gray-level distribution characteristics of the significant difference map are statistically analyzed to determine the gray-level threshold that can distinguish between defective and normal regions. Pixels with gray-level values ​​greater than the threshold are identified as candidate defective pixels and their gray-level values ​​are set to the maximum value. Pixels with gray-level values ​​less than or equal to the threshold are identified as normal pixels and their gray-level values ​​are set to the minimum value. In this way, the significant difference map is converted into an image containing only two gray-level values, resulting in a preliminary binary map of candidate regions of the significant difference map.

[0139] Following the previously determined coordinate transformation rules, the texture layer image sequence is mapped frame by frame to a preset two-dimensional plane. Edge detection is performed on the mapped texture layer image to capture pixels with abrupt changes in grayscale value. These pixels correspond to the texture edges and structural edges of the pipe's inner wall. All edge pixels are connected in their original positions and enhanced to generate an edge structure map of the pipe's inner wall.

[0140] Using the edge structure map as a reference for morphological reconstruction, morphological reconstruction operations are performed on the preliminary candidate region binary map. First, the candidate defect region in the preliminary candidate region binary map is expanded by dilation to align its edges with the real edges in the edge structure map. Then, the expanded candidate defect region is shrunk by erosion to eliminate isolated candidate pixels caused by noise. The edge contour of the preliminary candidate region binary map is corrected by a combination of dilation and erosion to obtain the optimized candidate region binary map of the pipe inner wall.

[0141] Connectivity analysis is performed on the binary image of the optimized candidate region to identify the set of all interconnected candidate defect pixels in the image. For each connected region, the smallest rectangle that can completely enclose the region is determined. This rectangle is the bounding rectangle of the connected region. All bounding rectangles are integrated according to their positions on a preset two-dimensional plane to generate an image mask that can mark all suspected defect locations. This image mask is the candidate region mask of the pipe inner wall.

[0142] The beneficial effect is that by combining core information such as lens focal length and image principal point coordinates in camera imaging parameters, a first mapping relationship between the pixel coordinates of the enhanced image sequence and the three-dimensional spatial coordinates of the inner wall of the pipe is established. This achieves a precise correspondence between image pixels and the actual spatial position of the pipe, allowing each pixel to be traced back to its true three-dimensional position on the inner wall of the pipe. This provides a reliable foundation for subsequent coordinate transformation and solves the problem of image decoupling from actual space.

[0143] Based on the geometric parameters of the pipe's inner wall and the first mapping relationship, coordinate transformation rules are formulated. The inherent characteristics of the pipe, such as its inner diameter and cross-sectional shape, are fully considered to ensure that the transformation of three-dimensional spatial coordinates to the preset two-dimensional plane conforms to the actual structure of the pipe. This avoids spatial distortion or positional deviation during the transformation process, allowing the transformation rules to not only fit the physical shape of the pipe but also adapt to the requirements of image unfolding.

[0144] The enhanced image sequence is transformed according to the coordinate transformation rules. The projection coordinates of each pixel on the preset two-dimensional plane are calculated and a projection coordinate set is formed. This achieves accurate positioning of image pixels on the two-dimensional plane and ensures that the projection position of each pixel can accurately reflect its relative position on the inner wall of the pipe. This provides a comprehensive and accurate coordinate basis for subsequent reprojection and ensures the rationality of the pixel distribution of the two-dimensional unfolded image.

[0145] The enhanced image sequence is reprojected using the projection coordinate set to generate a two-dimensional planar unfolded image, which transforms the curved surface image of the pipe inner wall into a planar image. This completely eliminates the image distortion and defect area deformation caused by the curved shape of the pipe, allowing the overall condition of the pipe inner wall and the location of defects to be presented intuitively. At the same time, the structural and texture details in the enhanced image are fully preserved, providing a clear and distortion-free image basis for the subsequent coarse extraction of defect candidate areas, and significantly reducing the probability of missed or false defects.

[0146] The formula comprehensively incorporates core imaging parameters such as camera focal length and image principal point coordinates, as well as the key geometric parameter of pipe radius. It also introduces preset adaptive adjustment coefficients, fully considering the differences in camera hardware characteristics, pipe physical structure, and actual imaging scenarios. This allows the projection coordinate calculation to be adapted to different detection equipment and pipe types, ensuring the consistency between the mapping results and the actual spatial position of the pipe's inner wall from the parameter level, and avoiding mapping deviations caused by missing or single parameters.

[0147] By calculating the difference between the horizontal coordinates of a pixel and the horizontal coordinates of the principal point in the image, and combining the results of correlation calculations involving the camera focal length, adaptive adjustment coefficient, specific function, and pipe radius, the horizontal projection coordinates are obtained through arctangent calculation. This accurately reflects the spatial position of the pixel in the horizontal direction and corrects imaging distortion through nonlinear transformation. The vertical projection coordinates are obtained by combining the difference between the vertical coordinates of the pixel and the vertical coordinates of the principal point with the pipe radius, and dividing by the square root of the sum of the squares of the horizontal difference and the correlation calculation results. This achieves accurate conversion of the vertical spatial position. The overall calculation logic takes into account both spatial geometric relationships and imaging rules, significantly improving the accuracy of coordinate transformation.

[0148] The formula uses scientific coordinate mapping logic to accurately convert distorted pixels in the original enhanced image caused by the curved shape of the pipe into regular coordinates on a preset two-dimensional plane. This effectively eliminates the image distortion problem caused by the curved pipe surface, ensuring that the projection position of each pixel can accurately correspond to its real relative position on the inner wall of the pipe. This provides core technical support for the subsequent generation of a complete and distortion-free two-dimensional planar unfolded image of the inner wall of the pipe, and solves the pain points of deformation and inaccurate positioning of defect areas caused by traditional direct imaging.

[0149] The preset adaptive adjustment coefficients can be flexibly calibrated according to actual scenarios such as the imaging clarity of the pipeline inner wall image, the pipeline radius, and the camera installation distance. This allows the formula to adapt to different detection environments and equipment configurations. Stable and accurate coordinate mapping can be achieved without frequent adjustments to the core calculation logic, which significantly enhances the versatility and practicality of the calculation method in actual pipeline defect detection. It also lays a reliable coordinate foundation for subsequent defect candidate region extraction, feature analysis, and other steps.

[0150] The two-dimensional planar unfolded image and the structural layer image sequence are mapped to a preset two-dimensional plane according to the same coordinate transformation rules. After the two are fully aligned in space, they are fused to generate a structural reference image. This not only preserves the complete information in the two-dimensional planar unfolded image, but also superimposes the core structural features of the structural layer image, providing a stable and reliable benchmark image for subsequent difference comparison and ensuring that the identification of defective areas has a clear structural reference.

[0151] Using a structural reference image as a baseline, a pixel-by-pixel comparison of grayscale values ​​is performed between the image and its two-dimensional unfolded image. Pixels with grayscale differences exceeding a fixed baseline are precisely marked, forming a significant difference map. This process effectively filters out suspected defect areas that deviate from the normal structure, initially separating defects from background information, thus narrowing down the scope and focusing on key areas for subsequent defect area extraction.

[0152] An adaptive threshold segmentation method is applied to the significant difference map. The threshold for distinguishing between defective and normal regions is determined by statistically analyzing the image's grayscale distribution characteristics. Differential pixels and normal pixels are then converted into two fixed grayscale values, generating a preliminary binary map of candidate regions. This processing method transforms ambiguous difference information into a clear black-and-white binary image, intuitively presenting the approximate range of suspected defects and providing a clear target area for subsequent optimization processing.

[0153] Using a unified coordinate transformation rule, the texture layer image sequence is mapped to a preset two-dimensional plane. Edge detection captures pixels with abrupt grayscale changes in the image; these pixels correspond to the texture and structural edges of the pipe's inner wall. This data is then integrated to generate an edge structure map. This map accurately delineates the detailed contours of the pipe's inner wall, providing a precise reference for correcting the edges of candidate regions and avoiding edge recognition errors in defective areas.

[0154] Morphological reconstruction is performed on the binary image of the initial candidate region based on the edge structure map. A dilation operation aligns the edges of the candidate defective regions with the real edges, and an erosion operation eliminates isolated pixels caused by noise. This combined operation corrects the edge contours of the initial candidate region, fills in small gaps within the region, and removes invalid noise interference, resulting in an optimized binary image of the candidate region with accurate boundaries and complete range.

[0155] Connectivity analysis is performed on the binary image of the optimized candidate region to identify all interconnected sets of candidate defect pixels. The bounding rectangle of each connected region is used as the candidate region, and these are integrated to generate a candidate region mask. This mask can accurately mark the location and extent of all suspected defects, providing precise region localization for subsequent deep feature fusion, avoiding interference from irrelevant regions, and significantly improving the efficiency and accuracy of subsequent defect detection.

[0156] S4. Perform deep feature fusion on the multi-scale context information corresponding to the candidate region mask to obtain the discriminative deep feature vector of the inner wall of the pipe;

[0157] In this embodiment of the invention, the step of performing deep feature fusion on the multi-scale context information corresponding to the candidate region mask to obtain the discriminative deep feature vector of the pipe inner wall includes:

[0158] Based on the candidate region mask, extract the corresponding candidate image block set in the two-dimensional planar unfolded image;

[0159] For the candidate image blocks in the candidate image block set, construct the image block-context information pair set of the inner wall of the pipe;

[0160] Feature extraction is performed on the image blocks and corresponding context information in the image block-context information pair set to obtain the local depth feature vector and multi-scale context depth feature vector set of the inner wall of the pipe;

[0161] The local depth feature vector and the multi-scale context depth feature vector set are weighted and fused to obtain the candidate image block depth feature vector of the inner wall of the pipe;

[0162] The depth feature vectors of the candidate image blocks are aggregated to generate the discriminative depth feature vector of the inner wall of the pipe.

[0163] Based on the positions of the bounding rectangles marked in the candidate region mask, the image region covered by each bounding rectangle is precisely extracted in the two-dimensional planar unfolded diagram. Each extracted image region is a candidate image block. All the extracted candidate image blocks are arranged in order of their original positions in the two-dimensional planar unfolded diagram to form a set of candidate image blocks for the inner wall of the pipe.

[0164] For each candidate image block in the candidate image block set, different ranges of image regions are extended outward from the candidate image block as the corresponding context information. The extended range includes the near-distance region, the mid-distance region, and the far-distance region around the candidate image block, ensuring that the context information can cover the local surrounding environment and the broad background environment of the candidate image block. Each candidate image block is paired with its corresponding context information of different ranges to form an image block-context information pair set on the inner wall of the pipe.

[0165] A feature extraction network is used to process each candidate image patch in the image patch-context information set. The network extracts the pixel features of the candidate image patches layer by layer through the convolutional layers of the network, capturing local features such as texture details, grayscale distribution, and edge morphology of the image patches. These local features are then transformed into fixed-length vectors through the fully connected layers of the network to obtain the local depth feature vector of the pipe inner wall. At the same time, the same feature extraction process is performed on the context information of different ranges corresponding to each candidate image patch, extracting the depth features of the context information of each range and transforming them into vectors to form a multi-scale context depth feature vector set of the pipe inner wall.

[0166] Based on the feature importance of each vector in the local depth feature vector and the multi-scale context depth feature vector set, a fixed fusion weight is assigned to each vector. The weight of the local depth feature vector is higher than the weight of each vector in the multi-scale context depth feature vector set. The weight of the context feature vector that is closer to the candidate image patch in the multi-scale context depth feature vector set is higher than the weight of the vector that is farther away. After multiplying each vector with its corresponding weight, the summation operation is performed on all weighted vectors to obtain the depth feature vector of the candidate image patch of the inner wall of the pipe.

[0167] The candidate image block depth feature vectors are arranged in order of their corresponding positions in the candidate region mask. The vector aggregation method is used to integrate all candidate image block depth feature vectors. By calculating the mean of all vectors and retaining the unique feature information of each vector, the multiple scattered candidate image block depth feature vectors are fused into a comprehensive feature vector. This comprehensive feature vector is the discriminative depth feature vector of the pipe inner wall.

[0168] The beneficial effect is that, based on the position of the outer rectangle of the candidate region mask mark, the corresponding image region is accurately extracted in the two-dimensional plane unfolded map to form a candidate image block set, which directly focuses on the suspected defect region, avoids irrelevant background regions from interfering with subsequent feature processing, greatly improves the targeting and efficiency of feature extraction, and ensures that subsequent operations are all carried out around the defect-related region.

[0169] Centered on each candidate image patch, the surrounding areas of different ranges are expanded as contextual information and a pairing set is constructed. This not only preserves the local defect features of the candidate image patch, but also incorporates the related information of its surrounding environment, making up for the lack of isolated feature information of a single image patch and providing rich contextual support for comprehensively capturing defect features.

[0170] By using a feature extraction network to process image patches and their corresponding context information, local depth features such as texture details and grayscale distribution of candidate image patches are accurately extracted. At the same time, multi-scale depth features of different ranges of context are captured, forming a feature system that combines local and global features. This effectively solves the problem that traditional single features are difficult to reflect the full picture of defects, making feature information more recognizable.

[0171] Based on the importance of features, different weights are assigned to local deep feature vectors and multi-scale contextual deep feature vectors to highlight the core role of local defect features, while taking into account the auxiliary value of contextual features of different ranges. By weighted summation, the features are organically fused, and the generated candidate image patch deep feature vectors can comprehensively reflect the local details and global correlation of defects, thereby improving the feature discrimination ability.

[0172] The depth feature vectors of all candidate image blocks are aggregated and processed. While retaining the unique features of each candidate region, they are integrated to form a comprehensive feature vector covering all suspected defect regions, namely the discriminative depth feature vector. This vector can comprehensively and accurately depict the core features of all defects in the inner wall of the pipeline, providing high-quality and highly recognizable feature basis for subsequent defect classification and identification, and significantly improving the accuracy of defect detection.

[0173] S5. Perform preliminary defect classification and confidence assessment on the discriminative deep feature vector, and iteratively optimize the low confidence region based on the confidence assessment results to obtain the defect category label and pixel-level semantic segmentation contour of the inner wall of the pipe.

[0174] In this embodiment of the invention, the step of performing preliminary defect classification and confidence assessment on the discriminative deep feature vector, and iteratively optimizing low-confidence regions based on the confidence assessment results to obtain defect category labels and pixel-level semantic segmentation contours of the pipe inner wall includes:

[0175] The discriminative deep feature vector is input into the classification network of the inner wall of the pipe to obtain the preliminary defect category prediction and prediction confidence of the inner wall of the pipe.

[0176] Regions with predicted confidence levels below a set threshold are marked as low-confidence regions;

[0177] Extract the multi-level features corresponding to the low-confidence region in the preset multi-scale depth feature map, and fuse them to generate an enhanced feature representation;

[0178] Based on the enhanced feature representation, the low-confidence region is reclassified to obtain the updated defect category prediction and the updated prediction confidence of the inner wall of the pipe.

[0179] By integrating the updated defect category predictions, the defect category labels of the inner wall of the pipe are obtained;

[0180] Based on the defect category label, pixel-level semantic segmentation is performed on the two-dimensional planar unfolded image to obtain the pixel-level semantic segmentation contour of the inner wall of the pipe.

[0181] The discriminative deep feature vector is fully input into the pre-trained classification network of the pipe inner wall. The classification network compares the discriminative deep feature vector with various defect feature templates stored in the network through internal feature matching and pattern recognition processes to determine the defect type that best matches the feature vector. At the same time, the reliability of the matching result is calculated, and the preliminary defect category prediction and prediction confidence of the pipe inner wall are directly output.

[0182] Based on the actual accuracy requirements for detecting defects in the inner wall of pipelines, a fixed confidence level judgment standard is set. Areas where the predicted confidence level does not reach the standard are clearly marked as low confidence areas. These areas are those where the reliability of the results in the preliminary classification is insufficient and needs further optimization.

[0183] A preset multi-scale depth feature map is obtained, which contains image feature information at different levels. Based on the position of the low-confidence region in the two-dimensional planar unfolded map, the features corresponding to each level in the preset multi-scale depth feature map are accurately extracted, including shallow texture detail features and deep semantic association features. These features at different levels are superimposed and integrated in a fixed manner to enhance the feature recognition of the low-confidence region and generate an enhanced feature representation.

[0184] The enhanced feature representation is input into the same classification network. The classification network re-matches the defect type based on the richer feature information. By refining the feature comparison process, the bias in the initial classification is corrected, the accurate defect type in the low confidence area is determined, and the reliability of the classification result is recalculated to obtain the updated defect category prediction and the updated prediction confidence of the inner wall of the pipeline.

[0185] Collect the updated defect category prediction results for all regions, including the initial defect category prediction for the original high-confidence regions and the optimized updated defect category prediction for the low-confidence regions. Organize and summarize the results according to the position of each region in the two-dimensional planar unfolded map to form a defect type identifier that comprehensively covers all regions of the pipeline inner wall, and obtain the defect category label of the pipeline inner wall.

[0186] Based on the defect category labels, the two-dimensional planar unfolded image is processed pixel by pixel. According to the defect category label of each pixel, the pixels of different defect categories are clearly distinguished from the pixels of normal areas. The boundaries of various defect areas are marked by contour delineation technology, and the specific range of each defect is clearly defined, finally obtaining the pixel-level semantic segmentation contour of the inner wall of the pipe.

[0187] The beneficial effect is that by inputting the discriminative deep feature vector into the trained classification network, and comparing the defect feature template through feature matching and pattern recognition, a preliminary defect category prediction and prediction confidence can be quickly obtained. This not only achieves efficient preliminary determination of defect category, but also quantifies the reliability of the classification result through confidence measurement, providing a clear basis for subsequent optimization and avoiding the problem of insufficient accuracy caused by direct classification.

[0188] By setting a threshold based on a fixed confidence level, areas that do not meet the standard are marked as low-confidence areas, accurately locating areas with insufficient reliability in the initial classification results. This enables stratified screening of defect classification results, avoids low-quality classification results from affecting the overall detection accuracy, and clarifies the target range for subsequent targeted optimization, thereby improving processing efficiency.

[0189] This process extracts shallow texture details and deep semantic association features from low-confidence regions within a pre-defined multi-scale deep feature map. These features are then superimposed and integrated to enhance feature discriminative power, generating an enhanced feature representation. This process supplements the feature information of low-confidence regions, solving the problem of inaccurate classification with a single feature and providing richer and more discriminative feature support for reclassification.

[0190] Based on enhanced feature representation, low-confidence regions are reclassified. The initial classification bias is corrected by refined feature comparison, resulting in updated defect category prediction and updated prediction confidence. This effectively improves the classification accuracy of low-confidence regions, ensures that the defect category determination of all regions has high reliability, and avoids misclassification due to insufficient features.

[0191] By integrating the preliminary classification results of high-confidence areas with the updated classification results of low-confidence areas, and organizing them according to regional location, a defect category label that fully covers the inner wall of the pipeline is formed. This achieves unified identification of defect categories, ensures that the label information is complete and accurate, and provides a clear and explicit classification basis for subsequent pixel-level semantic segmentation.

[0192] Based on defect category labels, the 2D planar unfolded image is differentiated pixel by pixel. The boundaries of various defect regions are marked by contour delineation, resulting in pixel-level semantic segmentation contours. These contours accurately define the specific range of each defect, clearly distinguishing different defect categories from normal regions. This provides a precise contour foundation for the subsequent quantification of defect geometric parameters and spatial pose, significantly improving the precision of defect detection.

[0193] S6. Based on the defect category label, the pixel-level semantic segmentation contour and the image-space mapping relationship, determine the actual geometric parameters and spatial pose of the defects on the inner wall of the pipe, so as to generate a structured defect detection report of the inner wall of the pipe.

[0194] In this embodiment of the invention, determining the actual geometric parameters and spatial pose of the defects on the inner wall of the pipe based on the defect category label, the pixel-level semantic segmentation contour, and the image-space mapping relationship, in order to generate a structured defect detection report for the inner wall of the pipe, includes:

[0195] Based on the image-space mapping relationship, the pixel-level semantic segmentation contour is mapped to the three-dimensional spatial coordinates of the inner wall of the pipe, thereby obtaining a three-dimensional spatial representation of the defect in the inner wall of the pipe.

[0196] Based on the three-dimensional spatial representation of the defect, the spatial morphological features of the defect in the inner wall of the pipe are extracted, the centroid coordinates and principal axis direction of the defect are determined, and the spatial pose of the defect in the inner wall of the pipe is obtained.

[0197] Based on the three-dimensional spatial representation of the defect, the surface shape and volume distribution of the defect are analyzed geometrically to obtain the actual geometric parameters of the defect on the inner wall of the pipe.

[0198] By combining the defect category label, the defect spatial pose, and the actual geometric parameters of the defect, a structured defect detection report for the inner wall of the pipeline is obtained.

[0199] The image-space mapping relationship is the correspondence rule between pixel coordinates and three-dimensional space coordinates established earlier. Based on this rule, the coordinates of each pixel on the pixel-level semantic segmentation contour are extracted one by one. Through coordinate transformation, these pixel coordinates are accurately mapped to the corresponding three-dimensional space coordinates of the inner wall of the pipe. All the mapped three-dimensional space coordinates are connected in the order of the original contour to completely restore the shape and position of the defect in three-dimensional space, and obtain the three-dimensional space representation of the defect in the inner wall of the pipe.

[0200] Spatial morphology analysis is performed on the three-dimensional spatial representation of the defect. All three-dimensional spatial coordinate points in the representation are traversed, and the average position of all coordinate points is calculated. This average position is the centroid coordinate of the defect. At the same time, by analyzing the distribution trend of the coordinate points, the principal axis that reflects the direction of defect extension is determined. The direction pointed to by the principal axis is the principal axis direction of the defect. Combining the centroid coordinate and the principal axis direction, the specific posture of the defect in three-dimensional space is determined, and the spatial pose of the defect on the inner wall of the pipe is obtained.

[0201] Based on the three-dimensional spatial representation of defects, the surface shape of defects is fitted and analyzed to determine the contour morphology of the defect surface. At the same time, by calculating the three-dimensional spatial range occupied by the defects, the extension dimensions of the defects in the length, width, and height directions are statistically analyzed to determine the volume distribution inside the defects. Through these geometric morphology analysis processes, key data that can quantify the physical morphology of defects are accurately obtained, and the actual geometric parameters of defects on the inner wall of the pipe are obtained.

[0202] According to the pre-set report structure framework, the defect category label, defect spatial pose and actual geometric parameters of the defect are systematically integrated. The defect category label clarifies the specific type of defect, the defect spatial pose clearly describes the three-dimensional position and orientation of the defect on the inner wall of the pipeline, and the actual geometric parameters of the defect quantify the physical characteristics of the defect such as size and shape. This information is arranged in a logical order and supplemented with auxiliary content such as inspection time and basic pipeline information to form a structured defect inspection report of the inner wall of the pipeline with clear structure and complete information.

[0203] The beneficial effect is that, based on the established image-space mapping relationship, the coordinates of each pixel on the pixel-level semantic segmentation contour are accurately converted into the three-dimensional spatial coordinates of the inner wall of the pipe. By connecting these three-dimensional coordinates in the order of the original contour, the true shape and position of the defect in three-dimensional space are completely restored, forming a three-dimensional spatial representation of the defect. This provides a realistic three-dimensional data foundation for subsequent defect quantitative analysis and solves the problem that two-dimensional images cannot reflect the spatial morphology of defects.

[0204] Spatial morphological features are extracted from the three-dimensional representation of the defect. The centroid coordinates of the defect are obtained by traversing all three-dimensional coordinate points and calculating the average position. The distribution trend of the coordinate points is analyzed to determine the direction of the defect's principal axis. The combination of these two methods forms the spatial pose of the defect, which accurately determines the specific location and extension direction of the defect in the three-dimensional space of the pipeline inner wall. This provides a precise spatial positioning basis for pipeline maintenance and avoids the limitation of traditional detection methods that can only mark the planar position.

[0205] Based on the three-dimensional spatial representation of defects, the surface shape of defects is fitted and analyzed to clarify the contour morphology. By calculating the three-dimensional spatial range occupied by the defects and statistically analyzing the dimensions such as length, width, and height, the internal volume distribution is analyzed. The actual geometric parameters of the defects are obtained through geometric morphology analysis, realizing a quantitative description of the physical morphology of defects. This allows the size, shape, and other information of defects to be accurately measured, providing key data support for the assessment of defect severity.

[0206] This system integrates defect category labels, defect spatial pose, and actual geometric parameters, logically organizing core information such as defect type, 3D location, and dimensions. It also supplements this with auxiliary content such as basic pipeline information and inspection time, resulting in a clearly structured and comprehensive structured defect inspection report. This report comprehensively presents the key characteristics of defects, solving the problems of fragmented information and lack of quantitative data in traditional inspection reports. It provides a comprehensive and reliable basis for pipeline maintenance decisions, significantly improving the targeting and efficiency of pipeline maintenance.

[0207] like Figure 2 The diagram shown is a functional block diagram of a pipeline inner wall defect detection system based on image recognition, provided in an embodiment of the present invention.

[0208] The image recognition-based pipe inner wall defect detection system 100 of this invention can be installed in an electronic device. Depending on the functions implemented, the image recognition-based pipe inner wall defect detection system 100 may include an image decoupling module 101, a visual enhancement module 102, an image unfolding and coarse inspection module 103, a feature fusion module 104, a defect fine classification module 105, and a defect quantification report module 106. The module described in this invention can also be called a unit, which refers to a series of computer program segments that can be executed by the processor of an electronic device and can perform a fixed function, and are stored in the memory of the electronic device.

[0209] In this embodiment, the functions of each module / unit are as follows:

[0210] The image decoupling module 101 is used to acquire the original image sequence of the inner wall of the pipe and perform image decoupling on the original image sequence to obtain a preprocessed image sequence of the original image sequence.

[0211] The visual enhancement module 102 is used to perform field-of-view consistency on the structural layer image sequence and the texture layer image sequence to obtain an enhanced image sequence of the inner wall of the pipe.

[0212] The image unfolding and coarse inspection module 103 is used to map the enhanced image sequence onto a preset two-dimensional plane according to the geometric parameters of the inner wall of the pipe and the camera imaging parameters to obtain a two-dimensional unfolded image of the inner wall of the pipe, and to coarsely extract the candidate region of the defect from the two-dimensional unfolded image to obtain a candidate region mask of the inner wall of the pipe.

[0213] The feature fusion module 104 is used to perform deep feature fusion on the multi-scale context information corresponding to the candidate region mask to obtain the discriminative deep feature vector of the inner wall of the pipe.

[0214] The defect classification module 105 is used to perform preliminary defect classification and confidence assessment on the discriminative deep feature vector, and to iteratively optimize the low confidence region based on the confidence assessment result to obtain the defect category label and pixel-level semantic segmentation contour of the inner wall of the pipe.

[0215] The defect quantification report module 106 is used to calculate the actual geometric parameters and spatial pose of the defects on the inner wall of the pipe based on the defect category label, the pixel-level semantic segmentation contour and the image-space mapping relationship, so as to generate a structured defect detection report of the inner wall of the pipe.

[0216] In the several embodiments provided by this invention, it should be understood that the disclosed methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.

[0217] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0218] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.

[0219] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0220] This application embodiment can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0221] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for detecting defects in the inner wall of a pipe based on image recognition, characterized in that, The method includes: S1. Obtain the original image sequence of the inner wall of the pipe, and perform image decoupling on the original image sequence to obtain the structure layer image sequence and texture layer image sequence of the original image sequence; S2. Perform view field consistency on the structure layer image sequence and the texture layer image sequence to obtain the enhanced image sequence of the inner wall of the pipe; S3. Based on the geometric parameters of the inner wall of the pipe and the camera imaging parameters, the enhanced image sequence is mapped onto a preset two-dimensional plane to obtain a two-dimensional planar unfolded image of the inner wall of the pipe, and the candidate defect region is coarsely extracted from the two-dimensional planar unfolded image to obtain a candidate region mask of the inner wall of the pipe. S4. Perform deep feature fusion on the multi-scale context information corresponding to the candidate region mask to obtain the discriminative deep feature vector of the inner wall of the pipe, including: Based on the candidate region mask, extract the corresponding candidate image block set in the two-dimensional planar unfolded image; For candidate image blocks in the candidate image block set, construct a set of image block-context information pairs for the inner wall of the pipe, including: For each candidate image block in the candidate image block set, different ranges of image regions are extended outward from the candidate image block as the corresponding context information. The extended range includes the near-distance region, the mid-distance region, and the far-distance region around the candidate image block, ensuring that the context information can cover the local surrounding environment and the broad background environment of the candidate image block. Each candidate image block is paired with its corresponding context information of different ranges to form the image block-context information pair set of the inner wall of the pipe. Feature extraction is performed on the image blocks and corresponding context information in the image block-context information pair set to obtain the local depth feature vector and multi-scale context depth feature vector set of the inner wall of the pipe; The local depth feature vector and the multi-scale context depth feature vector set are weighted and fused to obtain the candidate image block depth feature vector of the inner wall of the pipe; The depth feature vectors of the candidate image blocks are aggregated to generate the discriminative depth feature vector of the inner wall of the pipe; S5. Perform preliminary defect classification and confidence assessment on the discriminative deep feature vector, and iteratively optimize the low confidence region based on the confidence assessment results to obtain the defect category label and pixel-level semantic segmentation contour of the inner wall of the pipe. S6. Based on the defect category label, the pixel-level semantic segmentation contour and the image-space mapping relationship, determine the actual geometric parameters and spatial pose of the defects on the inner wall of the pipe, so as to generate a structured defect detection report of the inner wall of the pipe.

2. The method for detecting defects in the inner wall of a pipe based on image recognition as described in claim 1, characterized in that, The process of acquiring the original image sequence of the pipe's inner wall and decoupling the original image sequence to obtain the structural layer image sequence and texture layer image sequence includes: Obtain the original image sequence of the inner wall of the pipe; The original images in the original image sequence are decomposed in the frequency domain to obtain the low-frequency component image and the high-frequency component image of the original image sequence; The low-frequency component image is used as the structural layer image sequence of the original image sequence; The high-frequency component image is subjected to a non-downsampled contour wave transform to obtain the sub-band coefficient set of the high-frequency component image; The subband coefficient set is subjected to coefficient modulation to obtain the enhanced high-frequency subband coefficient set. Based on the enhanced high-frequency subband coefficient set, a non-downsampled contour wave inverse transform is performed to generate a reconstructed high-frequency detail image of the original image sequence, and the reconstructed high-frequency detail image is used as the texture layer image sequence of the original image sequence.

3. The method for detecting defects in the inner wall of a pipe based on image recognition as described in claim 1, characterized in that, The step of performing viewport unification on the structural layer image sequence and the texture layer image sequence to obtain the enhanced image sequence of the pipe inner wall includes: Extract feature point sets from the frame images in the structure layer image sequence and the corresponding frame images in the texture layer image sequence; Based on the feature point set, multimodal image registration is performed on the structural layer image sequence and the texture layer image sequence to obtain the spatial transformation parameter set of the inner wall of the pipe; Based on the spatial transformation parameter set, an affine transformation is performed on the texture layer image sequence to obtain a registered texture layer image sequence. The structural layer image sequence and the registration texture layer image sequence are fused to obtain an enhanced image sequence of the inner wall of the pipe.

4. The method for detecting defects in the inner wall of a pipe based on image recognition as described in claim 1, characterized in that, The step of mapping the enhanced image sequence onto a preset two-dimensional plane based on the geometric parameters of the pipe's inner wall and the camera's imaging parameters to obtain a two-dimensional planar unfolded image of the pipe's inner wall includes: Based on the camera imaging parameters, a first mapping relationship is established between the image pixel coordinates in the enhanced image sequence and the three-dimensional spatial coordinates of the inner wall of the pipe; Based on the geometric parameters of the inner wall of the pipe and the first mapping relationship, a coordinate transformation rule is established to transform the three-dimensional spatial coordinates to a preset two-dimensional plane. Based on the coordinate transformation rules, the enhanced image sequence is subjected to coordinate transformation, and the projection coordinate set of the pixels in the enhanced image sequence on the preset two-dimensional plane is calculated; Based on the projection coordinate set, the enhanced image sequence is reprojected to obtain a two-dimensional planar unfolded view of the inner wall of the pipe.

5. The method for detecting defects in the inner wall of a pipe based on image recognition as described in claim 4, characterized in that, The formula for calculating the projected coordinate set is: ; In the formula, Let f be the projection coordinates of the pixel on the preset two-dimensional plane, and let f be the camera focal length in the camera imaging parameters. Let y be the vertical image coordinate of the pixel in the original enhanced image. This represents the vertical coordinate of the principal point in the image from the camera imaging parameters. Here are the horizontal image coordinates of the pixel in the original enhanced image. R represents the lateral coordinate of the principal point in the camera imaging parameters, and R represents the pipe radius in the geometric parameters of the pipe's inner wall. This is the preset adaptive adjustment coefficient.

6. The method for detecting defects in the inner wall of a pipe based on image recognition as described in claim 1, characterized in that, The step of coarsely extracting candidate defect regions from the two-dimensional planar unfolded image to obtain a candidate region mask for the inner wall of the pipe includes: The two-dimensional planar unfolded diagram and the structural layer image sequence are mapped onto the preset two-dimensional plane to obtain a structural reference diagram of the inner wall of the pipe. A pixel-by-pixel difference comparison is performed on the structural reference image to obtain a significant difference map of the structural reference image; Adaptive threshold segmentation is performed on the significant difference map to obtain a preliminary binary map of candidate regions of the significant difference map; The texture layer image sequence is mapped onto the preset two-dimensional plane to generate the edge structure map of the inner wall of the pipe; Based on the edge structure map, the preliminary candidate region binary map is morphologically reconstructed to obtain the optimized candidate region binary map of the pipe inner wall; The bounding rectangle of the connected region in the binary graph of the optimized candidate region is used as the candidate region to generate the candidate region mask of the inner wall of the pipe.

7. The method for detecting defects in the inner wall of a pipe based on image recognition as described in claim 1, characterized in that, The process of performing preliminary defect classification and confidence assessment on the discriminative deep feature vector, and iteratively optimizing low-confidence regions based on the confidence assessment results, yields defect category labels and pixel-level semantic segmentation contours for the inner wall of the pipe, including: The discriminative deep feature vector is input into the classification network of the inner wall of the pipe to obtain the preliminary defect category prediction and prediction confidence of the inner wall of the pipe. Regions with predicted confidence levels below a set threshold are marked as low-confidence regions; Extract the multi-level features corresponding to the low-confidence region in the preset multi-scale depth feature map, and fuse them to generate an enhanced feature representation; Based on the enhanced feature representation, the low-confidence region is reclassified to obtain the updated defect category prediction and the updated prediction confidence of the inner wall of the pipe. By integrating the updated defect category predictions, the defect category labels of the inner wall of the pipe are obtained; Based on the defect category label, pixel-level semantic segmentation is performed on the two-dimensional planar unfolded image to obtain the pixel-level semantic segmentation contour of the inner wall of the pipe.

8. The method for detecting defects in the inner wall of a pipe based on image recognition as described in claim 1, characterized in that, The process of determining the actual geometric parameters and spatial pose of the defects on the inner wall of the pipe based on the defect category label, the pixel-level semantic segmentation contour, and the image-space mapping relationship, in order to generate a structured defect detection report for the inner wall of the pipe, includes: Based on the image-space mapping relationship, the pixel-level semantic segmentation contour is mapped to the three-dimensional spatial coordinates of the inner wall of the pipe, thereby obtaining a three-dimensional spatial representation of the defect in the inner wall of the pipe. Based on the three-dimensional spatial representation of the defect, the spatial morphological features of the defect in the inner wall of the pipe are extracted, the centroid coordinates and principal axis directions of the defect are determined, and the spatial pose of the defect in the inner wall of the pipe is obtained. Based on the three-dimensional spatial representation of the defect, the surface shape and volume distribution of the defect are analyzed geometrically to obtain the actual geometric parameters of the defect on the inner wall of the pipe. By combining the defect category label, the defect spatial pose, and the actual geometric parameters of the defect, a structured defect detection report for the inner wall of the pipeline is obtained.

9. A pipeline inner wall defect detection system based on image recognition, characterized in that, The system for implementing the image recognition-based pipe inner wall defect detection method of claim 1 includes: An image decoupling module is used to acquire the original image sequence of the inner wall of the pipe and decouple the original image sequence to obtain a preprocessed image sequence of the original image sequence. A visual enhancement module is used to perform field-of-view consistency on the structural layer image sequence and the texture layer image sequence to obtain an enhanced image sequence of the inner wall of the pipe; The image unfolding and coarse inspection module is used to map the enhanced image sequence onto a preset two-dimensional plane according to the geometric parameters of the inner wall of the pipe and the camera imaging parameters to obtain a two-dimensional unfolded image of the inner wall of the pipe, and to coarsely extract the candidate region of the defect from the two-dimensional unfolded image to obtain a candidate region mask of the inner wall of the pipe. The feature fusion module is used to perform deep feature fusion on the multi-scale context information corresponding to the candidate region mask to obtain the discriminative deep feature vector of the inner wall of the pipe, specifically for: Based on the candidate region mask, extract the corresponding candidate image block set in the two-dimensional planar unfolded image; For candidate image blocks in the candidate image block set, construct a set of image block-context information pairs for the inner wall of the pipe, including: For each candidate image block in the candidate image block set, different ranges of image regions are extended outward from the candidate image block as the corresponding context information. The extended range includes the near-distance region, the mid-distance region, and the far-distance region around the candidate image block, ensuring that the context information can cover the local surrounding environment and the broad background environment of the candidate image block. Each candidate image block is paired with its corresponding context information of different ranges to form the image block-context information pair set of the inner wall of the pipe. Feature extraction is performed on the image blocks and corresponding context information in the image block-context information pair set to obtain the local depth feature vector and multi-scale context depth feature vector set of the inner wall of the pipe; The local depth feature vector and the multi-scale context depth feature vector set are weighted and fused to obtain the candidate image block depth feature vector of the inner wall of the pipe; The depth feature vectors of the candidate image blocks are aggregated to generate the discriminative depth feature vector of the inner wall of the pipe; The defect classification module is used to perform preliminary defect classification and confidence assessment on the discriminative deep feature vector, and to iteratively optimize the low-confidence region based on the confidence assessment result, so as to obtain the defect category label and pixel-level semantic segmentation contour of the inner wall of the pipe. The defect quantification report module is used to calculate the actual geometric parameters and spatial pose of the defects on the inner wall of the pipe based on the defect category label, the pixel-level semantic segmentation contour and the image-space mapping relationship, so as to generate a structured defect detection report of the inner wall of the pipe.