Drainage pipeline shielding segmentation method based on computer vision
Through the CLIPSEG model and deformable convolution combined with morphological fusion module, the problem of image occlusion and segmentation of drag-and-drop drainage pipe robots is solved, and high-precision segmentation of irregular connectors and ropes is achieved, which improves the accuracy of pipeline disease detection.
Patent Information
- Application Number
- CN202510486562.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art is difficult to effectively divide the blocking area caused by mechanical structure in the image of the drag-and-drop drainage pipe robot, which affects the accuracy of disease detection, especially the lack of segmentation accuracy for irregularly shaped connectors and rope parts.
Using a method based on CLIPSEG model and deformable convolution, combined with the morphological fusion module, the center position and shape of the convolution kernel are adaptively adjusted, the fuzzy edge of the connector is strengthened, and a complete occlusion segmentation mask is generated.
It improves the segmentation accuracy of irregular objects, enhances the ability to identify complex targets, outputs more complete segmentation results, and improves the accuracy of pipeline disease recognition.
Smart Images

Figure SMS_10 
Figure SMS_13 
Figure SMS_14
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of deep learning and image processing, and particularly to a method for occluded segmentation of drainage pipes based on computer vision. Background Art
[0002] In the urban pipeline environment, the drag-type drainage pipe robot plays an important role in obtaining internal pipeline images for pipeline disease recognition. When the drainage pipe robot collects images in a real pipeline environment (mostly concrete pipelines), it will deploy mechanical components to grasp the inner wall of the pipeline and move forward continuously. However, the characteristics of its own mechanical structure will cause problems that some areas of the collected images are blocked by robot components, which will affect the analysis and judgment of the images by the later disease detection system. Therefore, it is of great significance to perform pipeline occlusion segmentation on the collected images.
[0003] The parts of the pipeline inner wall blocked by the pipeline robot mainly include two components: the rope part and the connecting part. Since the robot in the pipeline environment has unique semantic features and the available labeled data volume is relatively small, if these two parts are not split, a large amount of manpower and material resources need to be invested in labeling. In addition, the recognition accuracy of the general segmentation model for targets with complex structures is relatively low, while for targets with simple, clear, and consistent structures, it is relatively high. Based on this, a decoupling and merging method is used to process the segmentation problem of pipeline images to cover the positioning area as much as possible.
[0004] In traditional convolution operations, each convolution kernel can only fixedly cover a region with a fixed shape, so it is difficult to obtain relatively accurate features of irregular objects. This causes the network to be difficult to accurately judge the same object under different scenarios and perspectives when identifying and segmenting irregular objects (such as the connecting parts in this article). To solve this problem of traditional convolution, the current conventional solutions include data augmentation and setting features or algorithms invariant to geometric transformations (such as SIFT and sliding windows). However, these methods all have some limitations. The sample limitations of data augmentation will lead to poor generalization ability of the model, and the manually designed features or algorithms cannot handle overly complex transformations. Moreover, the scale of the labeled connecting part dataset is relatively small, the shape of the robot connecting part itself is irregular, and the texture of the connecting part in the occluded image is highly similar to the background (inner wall of the pipeline), resulting in problems of blurred edges and difficult segmentation.
[0005] The present invention proposes a method for occluded segmentation of drainage pipes based on computer vision, which can effectively solve the problem of segmentation when encountering images with irregular shapes, and is applicable to scientific fields such as intelligent detection and recognition of underground pipeline diseases of drag-type drainage pipe robots. Summary of the Invention
[0006] In view of the above problems in the prior art, the present invention proposes a method for segmenting occlusions in drainage pipes based on computer vision. This method uses the CLIPSEG (Image Segmentation Using Text and Image Prompts) model to segment the rope components causing pipe occlusions, uses deformable convolutions to adaptively adjust the center position and shape of the convolution kernel, better adapting to irregularly shaped connectors, and uses a morphological fusion module to orderly mine the detailed information in the fuzzy region, enabling the network to continuously strengthen the fuzzy edges of the connectors and output a more complete segmentation result. This method can effectively segment the connectors, contribute to the subsequent repair work of pipe images, is applicable to scientific fields such as pipe disease identification, and can be applied to urban pipe inspection and maintenance work.
[0007] The technical solution adopted by the present invention is as follows:
[0008] Step (1): Use a drag-type drainage pipe robot to collect pipe images in an experimental pipe model. Annotate the rope segmentation dataset and semantic segmentation dataset of the pipe images, and divide them into a training dataset and a test dataset.
[0009] Step (2): Construct a rope segmentation model based on CLIPSEG and a robot connector segmentation model based on deformable convolutions.
[0010] Step (3): Send the datasets divided in step (1) into the corresponding models constructed in step (2) to obtain the masks for these two parts, and use the morphological fusion module to process the obtained masks to generate the mask for the final full-image area to be repaired.
[0011] Step (4): Use the test dataset constructed in step (1) to test the trained network model, implement drainage pipe occlusion segmentation, and conduct an objective evaluation of the segmentation effect.
[0012] The beneficial effect of the present invention is that, compared with some traditional segmentation methods, the present invention proposes a method for segmenting occlusions in drainage pipes based on computer vision, which is well optimized for the problems of connector and rope segmentation. Combining deformable convolutions improves the segmentation accuracy of irregular objects. Based on the morphological aggregation module, the network can continuously strengthen the fuzzy edges of the connectors, and finally output a more complete segmentation result, improving the accuracy of target recognition for complex structures. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The present invention will be further described below with reference to the drawings and embodiments:
[0014] Figure 1 FIG. is a flowchart of a method for segmenting occlusions in drainage pipes based on computer vision according to an embodiment of the present invention;
[0015] Figure 2 The figure of the rope segmentation algorithm based on CLIPSEG according to an embodiment of the present invention
[0016] Figure 3 The figure of the deformable convolution calculation process according to an embodiment of the present invention;
[0017] Figure 4 The figure of the connector segmentation network according to an embodiment of the present invention;
[0018] Figure 5 The figure of the morphological processing result according to an embodiment of the present invention;
[0019] Figure 6 The figure of the objective evaluation of the segmentation effect according to an embodiment of the present invention. Specific implementation manners
[0020] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. The following embodiments do not limit the present invention. The embodiments described with reference to the accompanying drawings are exemplary and are intended to explain the present invention, but should not be construed as limiting the present invention.
[0021] Figure 1 The flowchart of the drainage pipe occlusion segmentation method based on computer vision according to an embodiment of the present invention; as Figure 1 shown, the drainage pipe occlusion segmentation method based on computer vision includes the following steps:
[0022] S1010: Use a drag-type drainage pipe robot to collect pipe images in an experimental pipe model. Label the rope segmentation dataset and the semantic segmentation dataset of the pipe images, and divide the training dataset and the test dataset.
[0023] S1020: Construct a rope segmentation model based on CLIPSEG and a robot connector segmentation model based on deformable convolution.
[0024] Figure 2 The structural diagram of the rope segmentation based on CLIPSEG. Use the CLIPSEG model to segment the rope components that cause pipe occlusion, and these components have common semantic information. The model input includes the query image to be segmented and the text query prompt for informing the model of the segmentation target. The CLIPSEG model consists of three main parts: the frozen CLIP Visual Transformer network T text 、the frozen CLIP Text Transformer network T vision and De seg (CLIPSeg Decoder decoder). Ttext For extracting query image features, T vision For extracting text prompt features, De seg Generate a segmentation mask map by combining the text prompt and query image features through the skip connections of a U-Net-like architecture. T text and T vision The backbone network adopted is ViT-B / 16 containing 12 Transformer blocks, T vision Specify N activation blocks among them, V i(i=1,2,…N) , for exchanging information with De seg , so De seg also only contains N Transformer blocks denoted as D i(i=1,2,…N) . In the present invention, the actual N used is 3, which are respectively T vision the 3rd, 5th, and 7th Transformer blocks in
[0025] Specifically, the input text prompt t is input to T text After obtaining the features as mentioned, it is mapped to the space with the same dimension as the De seg tokens, denoted as T. And T vision After obtaining the input occluded image x ∈ R W×H×3 , first, the output of V N and T are input together to the conditional network layer in De seg to combine the text prompt information and query image information, and output d containing the conditional restrictions of the segmentation target. Subsequently, d is input to D1, and its output is then jump-connected with V N-1 , and the result after connection is used as the input of d2 until the output of the last D N is output through the mapping layer to obtain the final segmentation result m1 ∈ {0, 1} H×W×1 (1 indicates that the pixel here belongs to the rope class, otherwise it is 0). Its formula is expressed as follows:
[0026] T = T text (t)(1)
[0027] m1 = De seg (T vision (x), T)(2)
[0028] Figure 3 Figure is the process diagram of deformable convolution calculation. The key to the implementation of deformable convolution is divided into four stages: reference sampling point positioning, offset sampling point positioning, linear interpolation, and convolution calculation.
[0029] Specifically, first, a set of reference sampling points in the area covered by the convolution kernel W with size k × k needs to be determined. Similar to ordinary convolution, their relative positions on the feature map O ∈ R H×W×C are still rectangular, such asFigure 3 (a) The red dots are the reference sampling points covered by the convolutional kernel W each time when k = 3. For the convolutional kernel W, W i,j represents the value at the position of the i-th row and j-th column of the convolutional kernel, where i ∈ [1, k] and j ∈ [1, k].
[0030] Subsequently, for each reference sampling point r i,j , its offset Δr is predicted through an offset regression network i,j H×W×N , N = k × k × 2, that is, two offsets in the horizontal and vertical directions are predicted for each reference point in the area covered by the convolutional kernel. This process is expressed by the formula:
[0031] Δr i,j = Δp reg (O, r i,j ; Θ Δ ) (3)
[0032] In formula (3), r i,j , i ∈ [1, k], j ∈ [1, k] means that each time the convolutional kernel K traverses the feature map, it covers the i-th row and j-th column in the convolutional kernel of the reference sampling point on the feature map. Δp reg represents the offset regression function, and Θ Δ represents the network parameters. Based on this, the position coordinates of the theoretically offset sampling point are calculated which are Figure 3 the blue dots in (b), and its formula is:
[0033]
[0034] However, the offset Δr obtained by convolution i,j is often a decimal , and there may not be a real corresponding pixel point in the feature map O. Therefore, the method of bilinear interpolation is used to calculate the feature value at the theoretically offset point position as shown in formula (5):
[0035]
[0036] where q is the reference sampling point with all spatial coordinates being integers in the feature map O (that is, the point where there are real pixels in the feature map O), q x , q y are the horizontal and vertical coordinates of q respectively, x ∈ [1, W], y ∈ [1, H], and G bi is a two-dimensional bilinear interpolation kernel, which can be decomposed in the row and column directions as:
[0037]
[0038] In formula (6), gbi is a simple linear interpolation function, defined as:
[0039] g bi (a,b)=max(0,1-|ab|) (7)
[0040] In formula (7), Calculate and The row and column directions of the points are interpolated.
[0041] Finally Convolution operation is performed with the weight convolution kernel K to obtain the corresponding value Y(r i,j ), as shown in formula (8):
[0042]
[0043] like Figure 3 Taking the three blue offset points in the rightmost orange box in the left image of (c) as an example, their eigenvalues are calculated by linear interpolation of the four yellow points closest to the offset points in the three red squares in the right image of 3(c). The final calculation result of the eigenvalue of the offset sampling point of the deformable convolution is 3(d).
[0044] S1030: The self-built paired pipeline training data set needs to be fed into the drainage pipeline occlusion segmentation network based on computer vision to obtain the masks of the two parts. The masks obtained by processing the morphological fusion module are used to generate the final mask of the area to be repaired in the whole image.
[0045] like Figure 4 As shown, the connector segmentation network structure designed by the present invention includes the following steps: first, the input occluded image is feature extracted through a ResNet module based on deformable convolution to obtain a set of feature representations; then, these features are sent to a partial decoder module to output a rough position positioning map of the connector. Next, through three parallel RA modules, adaptive reverse attention learning of high-level features and rough global attention maps of connectors is realized. On this basis, the segmentation network gradually refines the preliminary estimate into an accurate and complete fine activation map. The network structure eliminates the estimated connector salient areas from the high-level features and orderly mines the detailed information of the fuzzy areas, so that the network can continuously strengthen the fuzzy edges of the connectors and finally output a more complete segmentation result.
[0046] Finally, if Figure 5 As shown in Figure (a), the mask graph m3∈{0,1} after the simple superposition of the rope segmentation mask m1 and the connector segmentation mask m2 H×W×1(0 and 1 represent whether the pixels here are occluded respectively), there will inevitably be gaps at the edge joints, resulting in the failure to locate the actually occluded areas in the whole image. Therefore, in order to make the mask image required for restoration cover as much as possible the entire occluded area, morphological opening operation is used in this book to expand the edges in a small range to try to segment out all the occluded areas. After two morphological opening operations After that, as shown in Figure 5 (b), the segmentation mask image m ∈ {0, 1} of the whole image occluded area is obtained H×W×1 (0 and 1 represent whether the pixels here are occluded respectively). Figure 5 (c) shows the edge changes of the covered area before and after the morphological operation, Figure 5 (d) is the standard segmentation map of the whole image occluded area.
[0047] The whole process formula is as follows:
[0048] m3 = m1 + m2 (9)
[0049]
[0050] S1040: Use the self-built pipeline test dataset to verify the effect of the segmentation network model, realize the segmentation of pipeline videos, and conduct objective evaluation.
[0051] To objectively verify the segmentation accuracy of the segmentation model of the present invention and other related methods. The objective evaluation indexes used in the present invention are Mean IoU, Mean Dice, MAE, Fβ, Sα, Eφmax. And record the predicted segmentation mask image as Prediction, and the true segmentation mask image as Groud. The sizes of both are H×W, and both have M pixel points of H×W. The categories to be recognized and segmented in the image are K. Predictionk represents the pixel set of the kth category in the predicted segmentation mask image, and Groudk represents the pixel set of the kth category in the true segmentation mask image, where k is an integer in the range from 1 to K.
[0052] Mean IoU is an index used to evaluate the overlap degree between the predicted result and the true result. It calculates the intersection and union at the pixel level between the predicted result and the true result. The larger its value, the closer the predicted result is to the true result. Although it can measure the overlap degree between the predicted regions and the true regions of all categories and reflect the model's ability to locate the target, however, it does not consider the accuracy of the segmentation boundary. The calculation formula is as follows:
[0053]
[0054] Among them, Intersection(Prediction k ,Groundk ) and Union(Prediction k , Ground k ) represent the intersection and union of the pixels of the k-th class in the predicted segmentation mask image Prediction and the ground truth segmentation mask image Groud, respectively.
[0055] Mean Dice is an index used to evaluate the similarity between the predicted result and the ground truth result. It calculates the pixel-level intersection for each class in the predicted segmentation mask image Prediction and the ground truth segmentation mask image Groud, and has good robustness to the class imbalance problem. The larger its value, the closer the predicted result is to the ground truth result. The formula for Mean Dice is:
[0056]
[0057] MAE is used to measure the difference between the predicted pixel class and the ground truth pixel class, which is obtained by averaging the absolute values of the errors of each pixel point in the ground truth segmentation image and the predicted segmentation image. The smaller its value, the closer the predicted result is to the ground truth result. The formula for MAE is as follows:
[0058]
[0059] Among them, Prediction i represents the ground truth class of the i-th pixel, and Groud i represents the predicted class of the i-th pixel. Although the above three pixel-level metrics can reflect the degree of difference between the predicted result and the ground truth result, they can only reflect the global-level constraints and ignore the similarity of the segmentation target structure, spatial position information, and high-level semantic features of the image. Therefore, this book further strengthens the model's attention to the segmentation target using the following three metrics.
[0060] F β Comprehensively consider the precision and recall rate of the segmentation algorithm to measure the quality of the model's style result. In image segmentation, the precision of the algorithm represents the ratio of the number of correctly segmented pixels to the number of predicted segmented pixels, while the recall rate represents the ratio of the number of correctly segmented pixels to the number of ground truth segmented pixels. It comprehensively considers the model's recognition ability for negative and positive samples and is the harmonic mean of the two. The higher its value, the more robust the model is. The formula is as follows:
[0061]
[0062] Among them, TP represents true positives, FP represents false positives, and FN represents false negatives. β is a weight parameter, and in this book, its value is taken as 0.2.
[0063] Eφmax is an index for evaluating the quality of image segmentation based on the edge matching situation, which comprehensively considers the global mean of the image and the degree of local pixel matching, and pays more attention to the similarity between the edge of the segmentation result and the edge of the true segmentation image. The larger its value, the higher the degree of edge alignment between the segmentation result and the true result, that is, the closer the segmentation result is to the true result.
[0064]
[0065] where φ S (i,j) is the alignment matrix of the i-th row and j-th column in the image. The larger the value of Eφmax, the higher the degree of edge alignment between the segmentation result and the true result, that is, the closer the segmentation result is to the true result.
[0066] Sα more accurately reflects the similarity between the image segmentation result and the true result by comprehensively evaluating the object structural similarity S object and the regional structural similarity S region . The larger its value, the better the predicted segmentation result. Its formula is as follows:
[0067] S α =α·S object +(1-α)·S region (19)
[0068] S object focuses on the similarity of each pixel position between Prediction and Groud, and S region focuses on the similarity of local regions between the two, which makes S α have higher accuracy and robustness. α is a weight used to adjust the object structural similarity and the regional structural similarity, and its value in this section is 0.5.
[0069] To objectively verify the segmentation accuracy of the segmentation model in this chapter and other related methods. The objective evaluation indexes used in this book are Mean IoU, Mean Dice, MAE, Fβ, Sα, Eφmax. Denote the predicted segmentation mask image as Prediction and the true segmentation mask image as Ground. Both of them have the size of H×W, and there are M pixel points in both, where H×W = M. The categories to be recognized and segmented in the image are K. Predictionk represents the set of pixels of the k-th category in the predicted segmentation mask image, and Groud k represents the set of pixels of the k-th category in the true segmentation mask image, and k is an integer in the range from 1 to K.
[0070] Mean IoU is an indicator used to evaluate the degree of overlap between the predicted result and the ground truth. It calculates the intersection and union at the pixel level between the predicted result and the ground truth. The larger its value, the closer the predicted result is to the ground truth. Although it can measure the overlap degree of the predicted regions and the ground truth for all classes and reflect the model's ability to locate targets, however, it does not consider the accuracy of the segmentation boundary. The calculation formula is as follows:
[0071]
[0072] Among them, Intersection(Prediction k ,Groud k ) and Union(Prediction k ,Ground k ) respectively represent the intersection and union of the pixels of the k-th class in the predicted segmentation mask image Prediction and the ground truth segmentation mask image Groud.
[0073] Mean Dice is an indicator used to evaluate the similarity between the predicted result and the ground truth. It calculates the intersection at the pixel level for each class in the predicted segmentation mask image Prediction and the ground truth segmentation mask image Groud, and has good robustness to the problem of class imbalance. The larger its value, the closer the predicted result is to the ground truth. The formula of Mean Dice is:
[0074]
[0075] MAE is used to measure the difference between the predicted pixel class and the ground truth pixel class, which is obtained by averaging the absolute values of the errors of each pixel point in the ground truth segmentation image and the predicted segmentation image. The smaller its value, the closer the predicted result is to the ground truth. The formula of MAE is as follows:
[0076]
[0077] Among them, Prediction i represents the ground truth class of the i-th pixel, and Groud i represents the predicted class of the i-th pixel. Although the above three pixel-level indicators can reflect the difference degree between the predicted result and the ground truth, they can only reflect the global-level constraints and ignore the similarity of the segmentation target structure, spatial position information, and high-level semantic features of the image. Therefore, this book further strengthens the model's attention to the segmentation target with the following three indicators.
[0078] F βComprehensively consider the precision and recall rate of the segmentation algorithm to measure the quality of the model's style results. In image segmentation, the precision of the algorithm represents the ratio of the number of correctly segmented pixels to the number of predicted segmented pixels, while the recall rate represents the ratio of the number of correctly segmented pixels to the number of truly segmented pixels. It comprehensively considers the recognition ability of the model for negative and positive samples, and is the harmonic mean of the two. The higher the value, the more robust the model. The formula is as follows:
[0079]
[0080] Among them, TP represents true positives, FP represents false positives, and FN represents false negatives. β is a weight parameter, and its value in this book is 0.2.
[0081] Eφmax is an index for evaluating the quality of image segmentation based on the edge matching situation. It comprehensively considers the global mean of the image and the degree of local pixel matching, and pays more attention to the similarity between the edge of the segmentation result and the edge of the true segmentation image. The larger its value, the higher the degree of edge alignment between the segmentation result and the true result, that is, the closer the segmentation result is to the true result.
[0082]
[0083] Among them, φ s (i, j) is the alignment matrix of the i-th row and j-th column in the image. The larger the value of Eφmax, the higher the degree of edge alignment between the segmentation result and the true result, that is, the closer the segmentation result is to the true result.
[0084] Sα more accurately reflects the similarity between the image segmentation result and the true result by comprehensively evaluating the target structural similarity S object and the regional structural similarity S region , and the larger its value, the better the predicted segmentation result. The formula is as follows:
[0085] S α =α·S object +(1-α)·S region (28)
[0086] S object focuses on the similarity of each pixel position between Prediction and Groud, and S region focuses on the similarity of local regions between the two, which makes s α have high accuracy and robustness. α is a weight for adjusting the target structural similarity and regional structural similarity, and its value in this section is 0.5.
[0087] Figure 6 Data shows that the segmentation effect of the drainage pipe occlusion segmentation method based on computer vision is good and the details are rich.
[0088] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Those skilled in the art to which the present invention pertains may make various modifications or supplements to the described specific embodiments or use similar means for substitution, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.
Claims
1. A method for occluded segmentation of drainage pipes based on computer vision, comprising the following steps: Step (1): Use a drag-type drainage pipe robot to collect pipe images in an experimental pipe model. Annotate the rope segmentation dataset and semantic segmentation dataset of the pipe images, and divide them into training datasets and test datasets. Step (2): Based on the rope segmentation structure diagram of CLIPSeg, use the CLIPSeg model to segment the rope components that cause pipeline occlusion, and these components have common semantic information. The model input includes the query image to be segmented and the text query prompt used to inform the model of the segmentation target. The CLIPSeg model consists of three main parts: the frozen CLIP Visual Transformer network T text , the frozen CLIP Text Transformer network T vision and De seg (CLIPSeg Decoder decoder). T text is used to extract the query image features, T vision is used to extract the text prompt features, and De seg generates a segmentation mask map by combining the text prompt and the query image features through skip connections of a U-Net-like structure. T text and T vision adopt the ViT-B / 16 backbone network containing 12 Transformer blocks. T vision designates N of the activation blocks V i(i=1,2,...N) for exchanging information with De seg . Therefore, it also only contains N Transformer blocks denoted as D i(i=1,2,...N) . In the actual use of the present invention, N is 3, which are the 3rd, 5th, and 7th Transformer blocks in T vision respectively. Specifically, the input text prompt is from t to T text After the features are mentioned and taken, they are mapped to the space with the same dimension as the De seg token and denoted as T. And T vision After obtaining the input occluded image x ∈ R W×H×3 , first, the output of V N and T are input into the conditional network layer in De seg to combine the text prompt information and the query image information, and output d with the conditional restrictions of the segmentation target. Subsequently, d is input into D1, and its output is then jump-connected with V N-1 . The result after the connection is used as the input of d2 until the output of the last D N passes through the mapping layer to output the final segmentation result m1 ∈ {0, 1} H×W×1 (1 indicates that the pixel here belongs to the rope class, otherwise it is 0). Its formula is expressed as follows: T = T text (t) (1) m1 = De seg (T vision (x), T)(2) The key to the implementation of deformable convolution is divided into four stages: benchmark sampling point positioning, offset sampling point positioning, linear interpolation, and convolution calculation. Specifically, first, a set of reference sampling points in the area covered by the convolution kernel W of size k×k needs to be determined. Just like in ordinary convolution, their relative positions on the feature map O∈R H×W×C are still rectangular. The red points in Figure 3(a) are the reference sampling points covered by the convolution kernel W each time when k = 3. For the convolution kernel W, W i,j represents the value of the convolution kernel at the position of the i-th row and the j-th column, where i∈[1, k] and j∈[1, k]. Subsequently, for each reference sampling point r i,j , its offset Δr is predicted by an offset regression network i,j H×W×N , N = k×k×2, that is, for each reference point in the convolution kernel coverage area, the offsets in both the horizontal and vertical directions are predicted. This process is expressed by the formula: Δr i,j = Δp reg (O, r i,j ; Θ Δ ) (3) In formula (3), r i,j , where \(i\in[1,k]\) and \(j\in[1,k]\) indicate that when the convolutional kernel \(K\) traverses the feature map each time, it covers the \(i\)-th row and \(j\)-th column within the convolutional kernel of the reference sampling point on the feature map. \(\Delta p reg represents the offset regression function, and \(\Theta Δ represents the network parameters. Based on this, the theoretically offset sampling point position coordinates are the blue points in Figure 3(b), and its formula is: However, the offset Δr obtained by convolution i,j is often a decimal number, and there may not be a real corresponding pixel point in the feature map O. Therefore, the bilinear interpolation method is used to calculate the eigenvalue theoretically at the offset point position as shown in formula (5): where q is the reference sampling point with integer spatial coordinates in the feature map O (i.e., the point where there are real pixels in the feature map O), q x , q y are the horizontal and vertical coordinates of q respectively, x ∈ [1, W], y ∈ [1, H], G bi is a two-dimensional bilinear interpolation kernel, which can be decomposed in the row and column directions as: In formula (6), g bi is a simple linear interpolation function, defined as: g bi (a, b) = max(0, 1 - |a - b|) (7) In formula (7), calculate respectively the row-direction and column-direction interpolations corresponding to the points. Finally, perform a convolution operation with the weighted convolution kernel K to obtain the corresponding value Y(r i,j ) in the output feature map Y, as shown in Equation (8): Step (3): It is necessary to send the self-built paired pipe training dataset into the drainage pipe occlusion segmentation network based on computer vision to obtain the masks of these two parts. Use the morphological fusion module to process the obtained masks to generate the mask of the final full-image area to be repaired. The connector segmentation network structure designed in the present invention includes the following steps: First, the input occluded image is subjected to feature extraction through a ResNet module based on deformable convolution to obtain a set of feature representations; then, these features are sent to a partial decoder module to output a rough position localization map of the connector. Immediately afterwards, through three parallel RA modules, adaptive inverse attention learning of high-level features and a rough global attention map of the connector is realized. On this basis, the segmentation network gradually refines the initial estimate into an accurate and complete fine activation map. This network structure eliminates the significantly estimated connector regions from the high-level features, and orderly mines the detailed information of the fuzzy regions, so that the network can continuously strengthen the fuzzy edges of the connector and finally output a more complete segmentation result. Finally, since there are inevitably gaps at the edge joints in the mask image m3 ∈ {0, 1} (where 0 and 1 represent whether the pixels here are occluded or not) after simply superimposing the rope segmentation mask m1 and the connecting part segmentation mask m2, the actually occluded areas in the whole image are not located. Therefore, in order to make the mask image required for repair cover as much as possible the entire occluded area, morphological opening operation is used in this book to expand the edge in a small range to try to segment out all the occluded areas. After two morphological opening operations H×W×1 , the segmentation mask image m ∈ {0, 1} of the whole image occluded area as shown in Fig. 5(b) (where 0 and 1 represent whether the pixels here are occluded or not) is obtained. Fig. 5(c) shows the edge changes of the covered area before and after the morphological operation, and Fig. 5(d) is the standard segmentation map of the whole image occluded area. After that, the segmentation mask image m ∈ {0, 1} of the whole image occluded area as shown in Fig. 5(b) (where 0 and 1 represent whether the pixels here are occluded or not) is obtained. H×W×1 (0, 1 represent whether the pixels here are occluded or not respectively). Fig. 5(c) shows the edge changes of the covered area before and after the morphological operation, and Fig. 5(d) is the standard segmentation map of the whole image occluded area. The formula for the whole process is as follows: m3 = m1 + m2 (9) Step (4): Use the self-built pipe test dataset to verify the effect of the segmentation network model, realize the segmentation of the pipe video, and conduct an objective evaluation. To objectively verify the segmentation accuracy of the segmentation model in this chapter and other related methods. The objective evaluation metrics used in this book are Mean IoU, MeanDice, MAE, Fβ, Sα, and Eφmax. Denote the predicted segmentation mask image as Prediction and the ground truth segmentation mask image as Groud. Both have a size of H×W, and there are M pixel points in both, where H×W = M. The number of classes to be recognized and segmented in the image is K. prediction k represents the set of pixels of the k-th class in the predicted segmentation mask image, Groud k represents the set of pixels of the k-th class in the ground truth segmentation mask image, where k is an integer in the range from 1 to K. Mean IoU is an index used to evaluate the overlap degree between the prediction result and the true result. It calculates the intersection and union at the pixel level between the prediction result and the true result. The larger its value, the closer the prediction result is to the true result. Although it can measure the overlap degree between the prediction regions and the true regions of all categories and reflect the model's localization ability for the target, however, it does not consider the accuracy of the segmentation boundary. The calculation formula is as follows: Among them, Inter section(Prediction k ,Groud k ) and Union(Predicion k ,Groud k ) represent the intersection and union of the pixels of the k-th class in the predicted segmentation mask image Prediction and the ground truth segmentation mask image Groud, respectively. Mean Dice is an index used to evaluate the similarity between the prediction result and the true result. It calculates the intersection at the pixel level for each category in the predicted segmentation mask image Prediction and the true segmentation mask image Groud, and has good robustness to the problem of class imbalance. The larger its value, the closer the prediction result is to the true result. The formula for Mean Dice is: MAE is used to measure the difference between the predicted pixel category and the true pixel category, and is obtained by taking the average of the absolute values of the errors of each pixel point in the true segmentation image and the predicted segmentation image. The smaller its value, the closer the prediction result is to the true result. The formula for MAE is as follows: Among them, Prediction i represents the true category of the i-th pixel, and Groud i represents the predicted category of the i-th pixel. Although the above three pixel-level metrics can reflect the degree of difference between the prediction result and the true result, they can only reflect the global-level constraints and ignore the similarity of the segmentation target structure, spatial position information, and high-level semantic features of the image. Therefore, this book still uses the following three metrics to further strengthen the model's attention to the segmentation target. F β Comprehensively consider the precision and recall rate of the segmentation algorithm to measure the quality of the model's style results. In image segmentation, the precision of the algorithm represents the ratio of the number of correctly segmented pixels to the number of predicted segmented pixels, while the recall rate represents the ratio of the number of correctly segmented pixels to the number of actual segmented pixels. It comprehensively considers the model's recognition ability for negative and positive samples and is the harmonic mean of the two. The higher the value, the more robust the model. The formula is as follows: Among them, TP represents the true positive example, FP represents the false positive example, and FN represents the false negative example. β is a weight parameter, and the value in this book is 0.
2. Eφmax is an index for evaluating the quality of image segmentation based on edge matching conditions. It comprehensively considers the global mean of the image and the degree of local pixel matching, and pays more attention to the similarity between the edges of the segmentation result and the edges of the true segmentation image. The larger its value, the higher the degree of edge alignment between the segmentation result and the true result, that is, the closer the segmentation result is to the true result. where φ s (i,j) is the alignment matrix of the i-th row and j-th column in the image. The larger the value of Eφmax, the higher the degree of edge alignment between the segmentation result and the true result, that is, the closer the segmentation result is to the true result. Sα more accurately reflects the similarity between the image segmentation result and the ground truth by comprehensively evaluating the target structure similarity S object and the regional structure similarity S region . The larger the value of Sα, the better the predicted segmentation result. The formula is as follows: S α = α·S object + (1 - α)·S region (19) S object focuses on the similarity of each pixel position between Prediction and Ground, S region while focuses on the similarity of local regions between the two, which makes S α have high accuracy and robustness. α is a weight used to adjust the object structural similarity and the regional structural similarity, and the value in this section is 0.
5. To objectively verify the segmentation accuracy of the segmentation model in this chapter and other related methods. The objective evaluation metrics used in this book are Mean IoU, Mean Dice, MAE, Fβ, Sα, and Eφmax. Denote the predicted segmentation mask image as Prediction and the ground truth segmentation mask image as Groud. Both have a size of H×W, and there are M pixel points in both H×W. The number of classes to be recognized and segmented in the image is K. Prediction k represents the set of pixels of the k-th class in the predicted segmentation mask image, and Groud k represents the set of pixels of the k-th class in the ground truth segmentation mask image, where k is an integer in the range from 1 to K. Mean IoU is an index used to evaluate the degree of overlap between the prediction result and the true result. It calculates the intersection and union at the pixel level between the prediction result and the true result. The larger its value, the closer the prediction result is to the true result. Although it can measure the degree of overlap between the predicted regions and the true regions of all classes and reflect the model's ability to locate targets, however, it does not consider the accuracy of the segmentation boundary. The calculation formula is as follows: Among them, Intersection(Prediction k , Groud k ) and Union(Prediction k , Groud k ) respectively represent the intersection and union of the pixels of the k-th class in the predicted segmentation mask image Prediction and the ground truth segmentation mask image Ground. Mean Dice is an index used to evaluate the similarity between the prediction result and the true result. It calculates the intersection at the pixel level for each class in the predicted segmentation mask image Prediction and the true segmentation mask image Ground, and has good robustness to the problem of class imbalance. The larger its value, the closer the prediction result is to the true result. The formula for Mean Dice is: MAE is used to measure the difference between the predicted pixel class and the true pixel class, and is obtained by averaging the absolute values of the errors of each pixel point in the true segmentation image and the predicted segmentation image. The smaller its value, the closer the prediction result is to the true result. The formula for MAE is as follows: Among them, Prediction i represents the true category of the i-th pixel, and Ground i represents the predicted category of the i-th pixel. Although the above three pixel-level metrics can reflect the degree of difference between the prediction result and the true result, they can only reflect the global-level constraints, ignoring the similarity of the segmentation target structure, spatial position information, and high-level semantic features of the image. Therefore, this book still uses the following three metrics to further strengthen the model's attention to the segmentation target. F β Comprehensively consider the precision and recall rate of the segmentation algorithm to measure the quality of the model's style results. In image segmentation, the precision of the algorithm represents the ratio of the number of correctly segmented pixels to the number of predicted segmented pixels, while the recall rate represents the ratio of the number of correctly segmented pixels to the number of actual segmented pixels. It comprehensively considers the model's recognition ability for negative and positive samples and is the harmonic mean of the two. The higher the value, the more robust the model. The formula is as follows: Among them, TP represents true positive, FP represents false positive, and FN represents false negative. β is a weight parameter, and its value in this book is 0.
2. Eφmax is an index for evaluating the quality of image segmentation based on edge matching conditions. It comprehensively considers the global mean of the image and the degree of local pixel matching, and pays more attention to the similarity between the edges of the segmentation result and the edges of the true segmentation image. The larger its value, the higher the degree of edge alignment between the segmentation result and the true result, that is, the closer the segmentation result is to the true result. where φ s (i, j) is the alignment matrix of the i-th row and j-th column in the image. The larger the value of Eφmax, the higher the degree of edge alignment between the segmentation result and the true result, that is, the closer the segmentation result is to the true result. Sα more accurately reflects the similarity between the image segmentation result and the ground truth by comprehensively evaluating the target structure similarity S object and the regional structure similarity S region , and the larger its value, the better the predicted segmentation result. The formula is as follows: S α = α·S object +(1 - α)·S region (28) S object focuses on the similarity between each pixel position in Prediction and Groud, S region while focuses on the similarity of local regions between the two, which makes S α have high accuracy and robustness. α is a weight used to adjust the object structural similarity and regional structural similarity, and the value in this section is 0.5.
Citation Information
Patent Citations
Segmenting objects in digital images utilizing a multi-object segmentation model framework
US20220237799A1