Pipe body defect detection method, device and equipment based on machine vision

Through the machine vision-based tube defect detection method, the image embedding and mask matrix is ​​generated using a pre-trained network, which solves the problems of inaccurate detection and high cost in the prior art, and achieves efficient and accurate tube defect detection.

CN119991660AActive Publication Date: 2025-05-13XIDIAN UNIV HANGZHOU RES INST +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510459106.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-13
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

The prior art has problems such as high material surface quality requirements, inability to detect deep defects, high equipment costs, complex maintenance and insufficient data in the detection of pipe body defects, resulting in inaccurate and high cost.

Method used

Using a tube body defect detection method based on machine vision, by acquiring the tube body images, using a pretrained image encoder and a trained prompt generation network to generate image embedding and prompt embedding, combined with the pretrained mask decoder to generate a mask matrix, determine the tube body boundary coordinates and perform defect detection.

Benefits of technology

It realizes detection that is not affected by the surface quality of the material, can accurately detect deep defects, reduce equipment and maintenance costs, improve detection efficiency and accuracy, and overcome the problem of insufficient data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991660A_ABST
    Figure CN119991660A_ABST
Patent Text Reader

Abstract

The invention discloses a pipe body defect detection method, device and equipment based on machine vision. The method comprises the following steps: acquiring an image of a to-be-detected pipe body to obtain a to-be-detected image; a pre-trained image encoder is adopted to generate image embedding of the to-be-detected image; adopting the trained prompt generation network to generate prompt embedding of the to-be-detected image; the prompt embedding is used for prompting a pipe body area in the to-be-detected image; adopting a pre-trained mask decoder to generate a mask matrix corresponding to the to-be-detected image according to the image embedding and the prompt embedding; determining multiple groups of pipe body boundary coordinates according to the mask matrix corresponding to the to-be-detected image; and determining a defect detection result of the to-be-detected pipe body based on the plurality of groups of pipe body boundary coordinates. The method does not need a large amount of training data, can improve the defect detection capability, and greatly saves the labor cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of deep learning technology, and specifically relates to a pipe body defect detection method, device and equipment based on machine vision. Background Art

[0002] In the development of the current medical system, accurate drug delivery can not only improve the treatment effect, but also significantly improve the overall efficiency of the medical system. Infusion accuracy is an indispensable part of efficient treatment. For the peristaltic pump-driven infusion method, the infusion accuracy is determined by the quality of the peristaltic pump rubber tube. The peristaltic pump rubber tube directly affects the infusion accuracy, infusion pressure and its service life, and thus becomes a key control factor. During the infusion process of the peristaltic pump tube, the liquid is squeezed forward by the pump tube. In order to meet the requirements of infusion accuracy, the volume of each section of the liquid delivered must be consistent. The determination of the volume of each section of the liquid depends on the length and diameter of the pipeline, and the length of the pipeline is determined by the spacing of the peristaltic pump wheels. The pump wheel spacing of each device is fixed, and the delivery volume can be accurately determined by measuring the pipeline diameter. Therefore, accurate measurement of the pump tube diameter is the key to ensuring infusion accuracy. Due to the soft nature of the peristaltic pump tube, traditional caliper or contact measurement methods are difficult, and these methods are prone to cause pipeline deformation, thereby affecting the accuracy of the measurement. In addition, traditional tools such as calipers cannot directly measure the inner diameter, and usually require the pipeline to be cut open to destroy the integrity of the sample.

[0003] At present, there are some non-contact defect detection methods, such as ultrasonic defect detection methods, eddy current defect detection technology, and some defect detection methods based on deep learning. Although ultrasonic defect detection methods are widely used in industrial detection, they have some limitations. First, this method has high requirements on the surface quality of the detection object. Rough or contaminated surfaces will affect the propagation of sound waves, resulting in signal distortion. Secondly, the detection depth of ultrasound is limited by the material properties and frequency. Although higher-frequency ultrasound can provide higher resolution, the penetration depth is shallower; while low-frequency ultrasound has strong penetration but lower resolution. In addition, ultrasonic testing is mainly used to identify defects such as cracks and pores, and has low detection sensitivity for some small or complex defects. In addition, this method also requires high professional skills for operators, and those with insufficient experience may cause misjudgment or missed judgment. In addition, ultrasonic testing equipment is expensive and complex to maintain, and its application in complex environments may be limited by the portability of the equipment. Finally, the reflection and refraction of ultrasonic signals are affected by the location and morphology of the defects, and data interpretation also requires high-level technical support, which increases labor costs and detection time. Although eddy current defect detection technology has the advantages of non-destructive, fast and efficient, it also has some limitations. First of all, this technology also has high requirements for the surface of the material, otherwise it will affect the accuracy of the signal. Secondly, eddy current testing is mainly suitable for the detection of surface or near-surface defects. It has low sensitivity to deep defects and is difficult to effectively detect problems inside thick materials. It also has certain requirements for users. In addition, the equipment cost is high and requires regular calibration and maintenance, which further increases the cost of use. Although the defect detection method based on deep learning has the advantages of high efficiency and automation, it also has some significant defects. First, the deep learning model requires a large amount of high-quality labeled data, while defect samples are often scarce and difficult to obtain, and the labeling cost is high. Secondly, the model training time is long and the computing resource requirements are large, especially in the case of large-scale data sets and complex models; poor data quality or labeling errors will affect the accuracy of the model, and overfitting problems may also reduce the generalization ability of the model. Finally, the model has insufficient recognition ability for complex or rare defects, especially in some special working conditions, and its adaptability is poor. Summary of the invention

[0004] In order to solve the above problems existing in the prior art, the present invention provides a method, device and equipment for detecting pipe defects based on machine vision.

[0005] The technical problem to be solved by the present invention is achieved through the following technical solutions: The present invention provides a method for detecting tube defects based on machine vision, comprising: Acquire an image of the tube body to be inspected to obtain an image to be inspected; Using a pre-trained image encoder to generate an image embedding of the image to be inspected; Using a trained hint generation network to generate a hint embedding of the image to be inspected; the hint embedding is used to hint the tube body area in the image to be inspected; Using a pre-trained mask decoder to generate a mask matrix corresponding to the image to be inspected according to the image embedding and the prompt embedding; Determining multiple groups of tube body boundary coordinates according to the mask matrix corresponding to the image to be inspected; Based on the multiple groups of pipe body boundary coordinates, a defect detection result of the pipe body to be detected is determined.

[0006] The present invention also provides a tube body defect detection device based on machine vision, comprising: An acquisition module is used to acquire an image of the tube body to be inspected, thereby obtaining an image to be inspected; The boundary coordinate determination module is used to generate the image embedding of the image to be inspected by using a pre-trained image encoder; generate the hint embedding of the image to be inspected by using a trained hint generation network; the hint embedding is used to indicate the tube body area in the image to be inspected; generate the mask matrix corresponding to the image to be inspected by using a pre-trained mask decoder according to the image embedding and the hint embedding; determine multiple groups of tube body boundary coordinates according to the mask matrix corresponding to the image to be inspected; The defect detection module is used to determine the defect detection result of the pipe body to be detected based on the multiple groups of pipe body boundary coordinates.

[0007] The present invention also provides a pipe body defect detection device based on machine vision, comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; The memory is used to store computer programs; The processor is used to implement the steps of the above-mentioned tube body defect detection method based on machine vision when executing the program stored in the memory.

[0008] Compared with the prior art, the present invention has the following beneficial effects: The pipe body defect detection method of the present invention uses a trained and pre-trained network to perform pipe body defect detection based on the image of the pipe body to be detected. Therefore, it is not affected by the surface quality of the material of the detected object, and the rough or contaminated surface will not affect the final detection result. Defect detection can be performed without the operation of professionals, and the trained model does not need regular maintenance, thereby greatly reducing costs, overcoming the limitation of being unable to detect deep defects, and being able to provide more accurate detection results. In addition, when performing pipe body defect detection, some models used in the present invention are pre-trained image encoders and pre-trained mask decoders, so a large amount of training data is not required, solving the problem of insufficient data, and because the pre-trained model has been trained on a large data set, it has good generalization ability, thereby improving the defect detection ability, and there is no need to train the pre-trained model in subsequent training, thereby reducing the training time and the required computing resources. In addition, when performing tube body defect detection, the present invention uses a trained prompt generation network to automatically generate prompt embedding for the image to be inspected, and can automatically generate high-quality prompts for the region of interest to assist the pre-trained mask decoder to accurately generate a mask for the target area. This not only eliminates the need for manual intervention and significantly saves labor costs, but also improves the accuracy of the generated mask matrix, thereby improving defect detection capabilities.

[0009] The present invention will be further described in detail below with reference to the accompanying drawings and specific implementation methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 It is a flow chart of a pipe body defect detection method based on machine vision provided by an embodiment of the present invention; Figure 2 is an input and output connection relationship diagram between an image encoder, a hint generation network, and a mask decoder provided by an embodiment of the present invention; Figure 3 is a schematic diagram of a structure of a prompt generation network provided by an embodiment of the present invention; Figure 4 is a structural diagram of a mask decoder provided by an embodiment of the present invention; Figure 5 It is a schematic diagram of a set of tube body boundary coordinates of a tube body to be detected provided by an embodiment of the present invention, and a principle of calculating a corresponding set of index values ​​according to the set of tube body boundary coordinates. DETAILED DESCRIPTION

[0011] The present invention is further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.

[0012] Figure 1 FIG. 1 is a flow chart of a tube defect detection method based on machine vision provided by an embodiment of the present invention. Figure 1 As shown, the method includes: S101, acquiring an image of a tube body to be inspected, and obtaining an image to be inspected.

[0013] Here, the tube body to be tested can be any tube, for example, a soft tube, a hard tube, etc., which is not limited in the present invention. Exemplarily, the tube body to be tested can be a soft tube on a peristaltic pump.

[0014] In some embodiments, the image to be inspected may be a global image of the tube body. In some embodiments, the image to be inspected may also be each of the multiple local images of the tube body, and the multiple local images are spliced ​​along the axial direction of the tube body to form a global image of the tube body. When the image to be inspected is each of the multiple local images of the tube body, the following S102 to S106 are the processing process of each of the multiple local images.

[0015] S102: Generate an image embedding of the image to be inspected using a pre-trained image encoder.

[0016] Here, the image embedding of the image to be inspected is the image feature of the image to be inspected. Exemplarily, the pre-trained image encoder can be a MAE pre-trained visual transformer (ViT), where MAE is a Masked Autoencoder and ViT is a Vision Transformer.

[0017] S103, using the trained hint generation network to generate hint embedding of the image to be inspected; the hint embedding is used to hint the tube body area in the image to be inspected.

[0018] S104, using a pre-trained mask decoder to generate a mask matrix corresponding to the image to be inspected according to the image embedding and the prompt embedding.

[0019] For example, Figure 2 is a graph of the input and output connections between the image encoder, the hint generation network, and the mask decoder.

[0020] Here, the number of rows and columns of the mask matrix corresponding to the image to be inspected is the same as the resolution of the image to be inspected, and each element in the mask matrix represents whether a pixel in the image to be inspected that is at the same position as the element belongs to the tube body or the background. Specifically, the mask matrix is ​​a matrix composed of 0 and 1. In some embodiments, 0 represents that a pixel in the image to be inspected that is at the same position as itself belongs to the tube body, and 1 represents that a pixel in the image to be inspected that is at the same position as itself belongs to the background. In other embodiments, 1 represents that a pixel in the image to be inspected that is at the same position as itself belongs to the tube body, and 0 represents that a pixel in the image to be inspected that is at the same position as itself belongs to the background.

[0021] S105 , determining multiple groups of tube body boundary coordinates according to the mask matrix corresponding to the image to be inspected.

[0022] S106. Determine defect detection results of the pipe to be inspected based on the multiple sets of pipe boundary coordinates.

[0023] In some embodiments, the prompt generation network includes: an encoder and a decoder; the input of the encoder is used as the input of the prompt generation network, and the output of the decoder is used as the output of the prompt generation network; the decoder includes at least a first upsampling module and a second upsampling module, the output of the first upsampling module and the output of the encoder are both connected to the input of the second upsampling module, and the output of the second upsampling module is used as the output of the decoder. Exemplarily, the encoder can be a HarmonicDenseNet; the first upsampling module includes: a first convolutional layer, a second convolutional layer, a first batch of normalization layers and a first activation layer in sequence; the second upsampling module includes: a third convolutional layer, a fourth convolutional layer, a second batch of normalization layers and a second activation layer in sequence, and the parameters of the first convolutional layer, the second convolutional layer, the third convolutional layer and the fourth convolutional layer are the same, the parameters of the first batch of normalization layers and the second batch of normalization layers are the same, and the activation functions used by the first activation layer and the second activation layer are different. Figure 3 It is a schematic diagram of the structure of the prompt generation network. Figure 3 The prompt generation network is composed of an encoder Encoder and a decoder Decoder, and the decoder Decoder is composed of a first upsampling module and a second upsampling module connected in series, the first upsampling module is composed of two convolutional layers Cov with a kernel size of 3×3 and a zero padding Pad of 1, a batch normalization layer Batch Norm and a ReLu activation layer, and the second upsampling module is composed of two convolutional layers Cov with a kernel size of 3×3 and a zero padding Pad of 1, a batch normalization layer Batch Norm and a Tanh activation layer. The output of the encoder Encoder and the output of the ReLu activation layer are both connected to the input of the first convolutional layer Cov in the second upsampling module. Both the first upsampling module and the second upsampling module can generate features with a resolution of 64×64 and an output channel of 256. It should be noted that, in some embodiments, the decoder in the prompt generation network may also include more than two of the above-mentioned upsampling modules, and the structures of the upsampling modules are the same.

[0024] In the present invention, the trained hint generation network is obtained by using a binary cross entropy loss function, a Dice (differentiable image entropy compression) loss function and a training set, wherein the training set contains a plurality of labeled training samples, and the label of each training sample is the true mask matrix of the training sample, and the true mask matrix represents the tube body area in the training sample. The Dice loss can measure the overlap between the predicted mask matrix and the true mask matrix.

[0025] Here, the expressions of the binary cross entropy loss function and the Dice loss function are as follows: ; ; in, represents the binary cross entropy loss for each training sample, represents the Dice loss of each training sample, Represents each training sample, represents the hint embedding for each training sample, represents the true mask matrix for each training sample, The mask matrix representing the prediction for each training sample, represents the true positive example between the real mask matrix and the predicted mask matrix for each training sample, represents the false positives between the true mask matrix and the predicted mask matrix for each training sample, Represents the false negatives between the true mask matrix and the predicted mask matrix for each training sample.

[0026] Specifically, the trained prompt generation network is trained using the method described in the following steps S001 to S004 to obtain: S001. Obtain a training set and an initial model respectively; each training sample in the training set is a tube body image; the initial model includes a pre-trained image encoder, an initial prompt generation network and a pre-trained mask decoder.

[0027] Exemplarily, the training set can be prepared by the following method: the camera is mounted on a stepper motor, the stepper motor is fixed on a straight axis, the tube sample is placed under the camera, and the extension direction of the two ends of the tube sample is kept parallel to the axis; the camera starts to collect a local image of the tube sample from one end of the tube sample, and each collected image is used as an original sample, and each time an original sample is collected, the camera moves forward a distance, such as 2 cm, until all positions of the tube sample are collected, and this method is used to collect images of one or more tube samples, so as to obtain multiple original samples. After obtaining multiple original samples, each original sample is preprocessed, for example, denoising, deblurring, and size and format unification are performed in sequence, so as to obtain multiple training samples, and then a mask matrix of each training sample is generated, and the mask matrix is ​​used as the true mask matrix of the training sample.

[0028] Here, since the prompt generation network is composed of an encoder and a decoder, the network parameters of the initial prompt generation network are composed of two parts: the initial parameters of the encoder and the initial parameters of the decoder. Exemplarily, the initial parameters of the encoder can be generated using a pre-trained ImageNet network, and the initial parameters of the decoder can be obtained by random generation.

[0029] S002. During each training, select the training samples from the training set, and input the training samples into the model obtained in the previous training to obtain the predicted mask matrix of each training sample in the training set; the model obtained in the previous training includes a pre-trained image encoder, a prompt generation network obtained in the previous training, and a pre-trained mask decoder.

[0030] Here, the number of training samples selected in each training can be set according to actual needs, and the present invention does not limit this.

[0031] S003. Using the binary cross entropy loss function and the Dice loss function, according to the predicted mask matrix of each training sample in this training sample and the true mask matrix of each training sample in this training sample, the binary cross entropy loss and the Dice loss are calculated respectively.

[0032] S004. According to the binary cross entropy loss and the Dice loss, the network parameters of the prompt generation network obtained in the previous training are adjusted to obtain the prompt generation network obtained in this training. According to the prompt generation network obtained in this training, the pre-trained image encoder and the pre-trained mask decoder, the model obtained in this training is obtained. The training is iterated in this way until the training termination condition is reached to obtain the trained prompt generation network.

[0033] Here, after adding the calculated binary cross entropy loss and the Dice loss, the total loss of this training can be obtained. According to the total loss of this training, the network parameters of the prompt generation network obtained in the previous training are updated through reverse gradient propagation to obtain the prompt generation network obtained in this training. According to the prompt generation network obtained in this training, the pre-trained image encoder and the pre-trained mask decoder, the model obtained in this training is obtained. Iterative training is performed according to the above principle until the loss obtained is less than or equal to the preset loss or the number of training times reaches the preset number, indicating that the training termination condition is met, so that the prompt generation network obtained in the last training can be used as the trained prompt generation network. It should be noted that the principle of updating network parameters through reverse gradient propagation according to the total loss is an existing principle, and the present invention will not elaborate on this.

[0034] In some embodiments, the mask decoder is a decoder designed using a pre-trained Transformer decoder. The main task of the mask decoder is to map the image embedding, the prompt embedding, and the output token to the mask space, where the output token is a learnable token specifically used for mask prediction, which is used to dynamically generate parameters related to the mask. Specifically, Figure 4As shown in the figure, the mask decoder includes: two decoding layers (Decoding Layer 1, Decoding Layer 2), an image embedding upsampling module (Upsampling), a multilayer perceptron (Multilayer Perceptron, MLP), a dynamic linear classifier and a dot product unit (Dot). The decoding layer (Decoder Layer) is a Transformer decoder. The decoding layer updates the interaction between the prompt embedding (Prompt Encoder Input) through prompt self-attention (Prompt Self-Attention), and performs bidirectional information interaction between the image embedding (Image Encoder Input) and the prompt embedding through cross-attention (Cross-Attention in TwoDirections) to generate matrix features (Mask Feature). In this way, the model can maintain a comprehensive understanding of the image and prompt information during the decoding process, thereby improving the accuracy of mask prediction. Decoding Layer 1 and Decoding Layer 2 update all embeddings and output updated image embeddings and prompt embeddings, as well as output tokens. After running decoding layer 1 and decoding layer 2, the matrix features output by decoding layer 2 are spatially upsampled through the image embedding upsampling module to restore it to the resolution of the input image. At the same time, the output tokens are processed by a multi-layer perceptron, and the multi-layer perceptron inputs the output to a dynamic linear classifier. The dynamic linear classifier calculates the mask foreground probability of each pixel of the image and outputs the mask foreground probability of each pixel to a dot product operation unit. The dot product operation unit performs a dot product operation on the upsampled matrix features and the mask foreground probability of each pixel of the image, and finally generates a segmentation mask (i.e., a mask matrix Mask).

[0035] In some embodiments, the above S105 is implemented by the following steps: S1051, from the mask matrix corresponding to the image to be inspected, select a row of mask values ​​every Q rows, where Q is a preset positive integer, for example, Q may be 200.

[0036] S1052. For each selected row of mask values, analyze whether every two adjacent mask values ​​in a row of mask values ​​are the same based on the arrangement order of the mask values ​​contained in the row of mask values, and when two adjacent mask values ​​are different, record the position of the first mask value of the two adjacent mask values ​​in the mask matrix corresponding to the image to be inspected as a changed position, thereby obtaining a plurality of changed positions arranged in order.

[0037] Here, when the number of rows of the mask matrix is ​​X and the number of columns is Y, the Y mask values ​​of the 1st row can be selected for the first time, the Y mask values ​​of the 1st+Qth row can be selected for the second time, and so on, until the last Y mask values ​​are taken out.

[0038] Specifically, when the mask value changes from 0 to 1 or from 1 to 0, the position corresponding to this 0 or the position corresponding to this 1 is recorded. The recorded position is a change position, and these change points are boundary positions, marking the dividing line between the target area and the background area. Multiple change positions arranged in sequence can be obtained from each row taken out.

[0039] S1053. Obtain a set of tube body boundary coordinates according to a plurality of sequentially arranged change positions.

[0040] The multiple sequentially arranged change positions obtained from each row taken out are the important boundary positions of the tube body. Theoretically, four sequentially arranged change positions can be obtained from each row taken out, and these four sequentially arranged change positions are the four boundary positions of the tube body. Therefore, in some embodiments, the first change position among the multiple sequentially arranged change positions can be used as the outer boundary point coordinates of the left tube wall of the tube body; the second change position among the multiple sequentially arranged change positions can be used as the inner boundary point coordinates of the left tube wall of the tube body; the third change position among the multiple sequentially arranged change positions can be used as the inner boundary point coordinates of the right tube wall of the tube body; the fourth change position among the multiple sequentially arranged change positions can be used as the outer boundary point coordinates of the right tube wall of the tube body; the outer boundary point coordinates and inner boundary point coordinates of the left tube wall of the tube body, as well as the outer boundary point coordinates and inner boundary point coordinates of the right tube wall of the tube body, can be used as a set of tube body boundary coordinates. In this way, a set of tube body boundary coordinates can be obtained according to a row of mask values ​​selected each time.

[0041] S1054: After traversing the mask matrix corresponding to the image to be inspected, multiple groups of tube body boundary coordinates are obtained.

[0042] In some embodiments, the above S106 is implemented by the following steps: S1061. Determine a set of index values ​​according to each set of tube body boundary coordinates; wherein the set of index values ​​includes: the length of the outer side of the left tube wall from the boundary, the length of the inner side of the left tube wall from the boundary, the length of the outer side of the right tube wall from the boundary, the length of the inner side of the right tube wall from the boundary, the thickness of the left tube wall, the thickness of the right tube wall, the inner diameter length, and the outer diameter length.

[0043] In some embodiments, a set of index values ​​may also only include: inner diameter length and outer diameter length.

[0044] Specifically, for a set of tube body boundary coordinates, the horizontal coordinate of the coordinate of the outer boundary point of the left tube wall is used as the length of the outer side of the left tube wall from the boundary; the horizontal coordinate of the coordinate of the inner boundary point of the left tube wall is used as the length of the inner side of the left tube wall from the boundary; the horizontal coordinate of the coordinate of the outer boundary point of the right tube wall is used as the length of the outer side of the right tube wall from the boundary; the horizontal coordinate of the coordinate of the inner boundary point of the right tube wall is used as the length of the inner side of the right tube wall from the boundary; the horizontal coordinate of the coordinate of the inner boundary point of the left tube wall is subtracted from the horizontal coordinate of the coordinate of the inner boundary point of the right tube wall to obtain an inner diameter length; the horizontal coordinate of the coordinate of the outer boundary point of the left tube wall is subtracted from the horizontal coordinate of the coordinate of the outer boundary point of the right tube wall to obtain an outer diameter length. The outer diameter length reflects the outer width of the entire tube body, and the inner diameter length reflects the width of the internal space of the tube body, that is, the width of the area where liquid or material flows.

[0045] For example, Figure 5 It is a schematic diagram of a set of tube body boundary coordinates of the tube body to be detected, and the principle of calculating the length of the outer side of the left tube wall from the boundary, the length of the inner side of the left tube wall from the boundary, the length of the outer side of the right tube wall from the boundary, the length of the inner side of the right tube wall from the boundary, the left tube wall thickness, the right tube wall thickness, the inner diameter length and the outer diameter length according to the set of tube body boundary coordinates, wherein point A is the outer boundary point of the left tube wall of the tube body, point B is the inner boundary point of the left tube wall of the tube body, point C is the inner boundary point of the right tube wall of the tube body, and point D is the outer boundary point of the right tube wall of the tube body.

[0046] S1062. Determine defect detection results of the pipe to be inspected based on multiple groups of index values ​​corresponding to multiple groups of pipe boundary coordinates.

[0047] Specifically, S1062 can be implemented through steps S1 to S5: S1. Calculate the mean and standard deviation of the inner diameter lengths in multiple groups of index values ​​corresponding to multiple groups of pipe body boundary coordinates, and obtain a first mean and a first standard deviation.

[0048] Here, the mean of the inner diameter length among the multiple groups of index values ​​corresponding to the multiple groups of tube body boundary coordinates is calculated, and the obtained mean is recorded as the first mean, and the standard deviation of the inner diameter length among the multiple groups of index values ​​corresponding to the multiple groups of tube body boundary coordinates is calculated, and the obtained standard deviation is recorded as the first standard deviation.

[0049] S2. Calculate the mean and standard deviation of the outer diameter length in multiple groups of index values ​​to obtain a second mean and a second standard deviation.

[0050] Here, the mean of the outer diameter length among the multiple groups of index values ​​corresponding to the multiple groups of tube body boundary coordinates is calculated, and the obtained mean is recorded as the second mean, and the standard deviation of the outer diameter length among the multiple groups of index values ​​corresponding to the multiple groups of tube body boundary coordinates is calculated, and the obtained standard deviation is recorded as the second standard deviation.

[0051] S3. Calculate a first standard deviation percentage based on the first mean and the first standard deviation, and calculate a second standard deviation percentage based on the second mean and the second standard deviation.

[0052] S4. When the first mean satisfies the first preset condition, the second mean satisfies the second preset condition, the first standard deviation satisfies the third preset condition, the second standard deviation satisfies the fourth preset condition, the first standard deviation percentage is within the first preset standard deviation percentage range, the second standard deviation percentage is within the second preset standard deviation percentage range, the length of the outer side of the left tube wall from the boundary in multiple groups of indicators are all within the first preset boundary value range, the length of the inner side of the left tube wall from the boundary in multiple groups of indicators are all within the second preset boundary value range, the length of the outer side of the right tube wall from the boundary in multiple groups of indicators are all within the third preset boundary value range, and the length of the inner side of the right tube wall from the boundary in multiple groups of indicators are all within the fourth preset boundary value range, a defect detection result is obtained that characterizes that the tube diameter size of the tube body to be inspected is free of defects.

[0053] S5. When at least one of the first mean, the first standard deviation, the second mean and the second standard deviation does not meet the corresponding preset conditions, or at least one of the first standard deviation percentage and the second standard deviation percentage is not within the corresponding preset standard deviation percentage range, or the lengths of the outer side of the left tube wall from the boundary in multiple groups of indicators are not all within the first preset boundary value range, or the lengths of the inner side of the left tube wall from the boundary in multiple groups of indicators are not all within the second preset boundary value range, or the lengths of the outer side of the right tube wall from the boundary in multiple groups of indicators are not all within the third preset boundary value range, or the lengths of the inner side of the right tube wall from the boundary in multiple groups of indicators are not all within the fourth preset boundary value range, a defect detection result is obtained, indicating that the diameter size of the tube body to be inspected is defective.

[0054] Here, through these indicators, the dimensional consistency and stability of the pipe to be inspected at different positions can be evaluated, thereby providing a more accurate basis for defect detection.

[0055] Here, the first preset condition, the second preset condition, the third preset condition, the fourth preset condition, the first preset standard deviation percentage range, the second preset standard deviation percentage range, the first preset boundary value range, the second preset boundary value range, the third preset boundary value range and the fourth preset boundary value range can all be set according to actual needs, and the present invention is not limited to this.

[0056] In some embodiments, the first preset condition is a preset first mean threshold, the second preset condition is a preset second mean threshold, the third preset condition is a preset first standard deviation threshold, and the fourth preset condition is a preset second standard deviation threshold. Accordingly, when the first mean is less than or equal to the preset first mean threshold, the second mean is less than or equal to the preset second mean threshold, the first standard deviation is less than or equal to the preset first standard deviation threshold, and the second standard deviation is less than or equal to the preset second standard deviation threshold, a defect detection result is obtained, indicating that the tube diameter size of the tube body to be inspected is free of defects; otherwise, a defect detection result is obtained, indicating that the tube diameter size of the tube body to be inspected is defective.

[0057] In some embodiments, the first preset condition is a preset first numerical range, the second preset condition is a preset second numerical range, the third preset condition is a preset third numerical range, and the fourth preset condition is a preset fourth numerical range. Accordingly, when the first mean belongs to the preset first numerical range, the second mean belongs to the preset second numerical range, the first standard deviation belongs to the preset third numerical range, and the second standard deviation belongs to the preset fourth numerical range, it means that the inner diameter and outer diameter of the tube body to be tested are qualified.

[0058] In some embodiments, after obtaining the defect detection results of each image to be inspected, multiple groups of indicators of the image to be inspected, as well as the first mean, first standard deviation, second mean and second standard deviation can be organized into a table, and the defect detection results of the image to be inspected can be output together with the table.

[0059] The pipe body defect detection method of the present invention uses a trained and pre-trained network to perform pipe body defect detection based on the image of the pipe body to be detected. Therefore, it is not affected by the surface quality of the material of the detected object, and the rough or contaminated surface will not affect the final detection result. Defect detection can be performed without the operation of professionals, and the trained model does not need regular maintenance, thereby greatly reducing costs, overcoming the limitation of being unable to detect deep defects, and being able to provide more accurate detection results. In addition, when performing pipe body defect detection, some models used in the present invention are pre-trained image encoders and pre-trained mask decoders, so a large amount of training data is not required, solving the problem of insufficient data, and because the pre-trained model has been trained on a large data set, it has good generalization ability, thereby improving the defect detection ability, and there is no need to train the pre-trained model in subsequent training, thereby reducing the training time and the required computing resources. In addition, when performing tube body defect detection, the present invention uses a trained prompt generation network to automatically generate prompt embedding for the image to be inspected, and can automatically generate high-quality prompts for the region of interest to assist the pre-trained mask decoder in accurately generating the mask of the target area. This not only eliminates the need for manual prompt input, significantly saving labor costs, but also improves the accuracy of the generated mask matrix, thereby improving the defect detection capability.

[0060] The present invention also provides a tube body defect detection device based on machine vision, comprising: An acquisition module is used to acquire an image of the tube body to be inspected, thereby obtaining an image to be inspected; The boundary coordinate determination module is used to generate an image embedding of the image to be inspected by using a pre-trained image encoder; generate a hint embedding of the image to be inspected by using a trained hint generation network; the hint embedding is used to indicate the tube body area in the image to be inspected; generate a mask matrix corresponding to the image to be inspected by using a pre-trained mask decoder according to the image embedding and the hint embedding; determine multiple sets of tube body boundary coordinates according to the mask matrix corresponding to the image to be inspected; The defect detection module is used to determine the defect detection result of the pipe body to be detected based on multiple sets of pipe body boundary coordinates.

[0061] It should be noted that the specific execution principle of each of the above modules has been described in detail in the above-mentioned tube defect detection method based on machine vision, and will not be repeated here.

[0062] The present invention also provides a tube defect detection device based on machine vision, comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; the memory is used to store computer programs; the processor is used to implement the steps of the above-mentioned tube defect detection method based on machine vision when executing the program stored in the memory.

[0063] It should be noted that the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.

[0064] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification.

[0065] In the specification, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality of components. Certain measures are recorded in different embodiments, but this does not mean that these measures cannot be combined to produce good effects.

[0066] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the scope of protection of the present invention.

Claims

1. A tube defect detection method based on machine vision, characterized in that: include: Acquire an image of the tube body to be inspected to obtain an image to be inspected; Using a pre-trained image encoder to generate an image embedding of the image to be inspected; Using a trained hint generation network to generate a hint embedding of the image to be inspected; the hint embedding is used to hint the tube body area in the image to be inspected; Using a pre-trained mask decoder to generate a mask matrix corresponding to the image to be inspected according to the image embedding and the prompt embedding; Determining multiple groups of tube body boundary coordinates according to the mask matrix corresponding to the image to be inspected; Based on the multiple groups of pipe body boundary coordinates, a defect detection result of the pipe body to be detected is determined.

2. The tube defect detection method based on machine vision according to claim 1, characterized in that: The trained prompt generation network is obtained by training with a binary cross entropy loss function, a Dice loss function and a training set, wherein the training set contains a plurality of labeled training samples, the label of each training sample is a true mask matrix of the training sample, and the true mask matrix represents the tube body area in the training sample.

3. The tube defect detection method based on machine vision according to claim 2, characterized in that: The trained prompt generation network is trained using the following method: The training set and the initial model are obtained respectively; each training sample in the training set is a tube body image; the initial model includes the pre-trained image encoder, the initial hint generation network and the pre-trained mask decoder; In each training, a training sample is selected from the training set, and the training sample is input into the model obtained in the previous training to obtain a predicted mask matrix of each training sample in the training sample; the model obtained in the previous training includes the pre-trained image encoder, the cue generation network obtained in the previous training, and the pre-trained mask decoder; Using the binary cross entropy loss function and the Dice loss function, according to the predicted mask matrix of each training sample in the current training samples and the real mask matrix of each training sample in the current training samples, respectively calculate the current binary cross entropy loss and the current Dice loss; According to the binary cross entropy loss and the Dice loss, the network parameters of the prompt generation network obtained in the previous training are adjusted to obtain the prompt generation network obtained in this training. According to the prompt generation network obtained in this training, the pre-trained image encoder and the pre-trained mask decoder, the model obtained in this training is obtained. The training is iterated in this way until the training termination condition is reached to obtain the trained prompt generation network.

4. The tube defect detection method based on machine vision according to claim 2, characterized in that: The expressions of the binary cross entropy loss function and the Dice loss function are as follows: ; ; in, represents the binary cross entropy loss for each training sample, represents the Dice loss of each training sample, Represents each training sample, represents the hint embedding for each training sample, represents the true mask matrix for each training sample, The mask matrix representing the prediction for each training sample, represents the true positive example between the real mask matrix and the predicted mask matrix for each training sample, represents the false positives between the true mask matrix and the predicted mask matrix for each training sample, Represents the false negatives between the true mask matrix and the predicted mask matrix for each training sample.

5. The tube defect detection method based on machine vision according to claim 1, characterized in that: The prompt generation network includes: an encoder and a decoder; the input of the encoder is used as the input of the prompt generation network, and the output of the decoder is used as the output of the prompt generation network; the decoder includes a first upsampling module and a second upsampling module, the output of the first upsampling module and the output of the encoder are both connected to the input of the second upsampling module, and the output of the second upsampling module is used as the output of the decoder.

6. The method for detecting tube defects based on machine vision according to claim 5, characterized in that: The encoder is Harmonic DenseNet; The first upsampling module includes, in sequence: a first convolutional layer, a second convolutional layer, a first batch of normalization layers and a first activation layer; the second upsampling module includes, in sequence: a third convolutional layer, a fourth convolutional layer, a second batch of normalization layers and a second activation layer, and the parameters of the first convolutional layer, the second convolutional layer, the third convolutional layer and the fourth convolutional layer are the same, the parameters of the first batch of normalization layers and the second batch of normalization layers are the same, and the first activation layer and the second activation layer use different activation functions.

7. The method for detecting tube defects based on machine vision according to claim 1, characterized in that: The number of rows and columns of the mask matrix corresponding to the image to be inspected is the same as the resolution of the image to be inspected, and each element in the mask matrix corresponding to the image to be inspected is used to represent whether a pixel in the image to be inspected that has the same position as the element belongs to the tube body or the background; determining multiple groups of tube body boundary coordinates according to the mask matrix corresponding to the image to be inspected includes: From the mask matrix corresponding to the image to be inspected, a row of mask values ​​is selected every Q rows, where Q is a preset positive integer; For each selected row of mask values, according to the arrangement order of the mask values ​​contained in the row of mask values, analyzing whether every two adjacent mask values ​​in the row of mask values ​​are the same, and when two adjacent mask values ​​are different, recording the position of the first mask value of the two adjacent mask values ​​in the mask matrix corresponding to the image to be inspected as a change position, to obtain a plurality of change positions arranged in sequence; Obtaining a set of tube body boundary coordinates according to the plurality of sequentially arranged change positions; After traversing the mask matrix corresponding to the image to be inspected, the multiple groups of tube body boundary coordinates are obtained.

8. The method for detecting tube defects based on machine vision according to claim 7, characterized in that: The step of determining the defect detection result of the pipe to be detected based on the multiple sets of pipe boundary coordinates includes: According to each set of pipe body boundary coordinates, a set of index values ​​is determined; wherein the set of index values ​​includes: the length of the outer side of the left pipe wall from the boundary, the length of the inner side of the left pipe wall from the boundary, the length of the outer side of the right pipe wall from the boundary, the length of the inner side of the right pipe wall from the boundary, the thickness of the left pipe wall, the thickness of the right pipe wall, the inner diameter length, and the outer diameter length; Based on the multiple groups of index values ​​corresponding to the multiple groups of pipe body boundary coordinates, the defect detection result of the pipe body to be detected is determined.

9. A tube defect detection device based on machine vision, characterized in that: include: An acquisition module is used to acquire an image of the tube body to be inspected, thereby obtaining an image to be inspected; The boundary coordinate determination module is used to generate the image embedding of the image to be inspected by using a pre-trained image encoder; generate the hint embedding of the image to be inspected by using a trained hint generation network; the hint embedding is used to indicate the tube body area in the image to be inspected; generate the mask matrix corresponding to the image to be inspected by using a pre-trained mask decoder according to the image embedding and the hint embedding; determine multiple groups of tube body boundary coordinates according to the mask matrix corresponding to the image to be inspected; The defect detection module is used to determine the defect detection result of the pipe body to be detected based on the multiple groups of pipe body boundary coordinates.

10. A pipe defect detection device based on machine vision, comprising a processor, a communication interface, a memory and a communication bus, characterized in that: The processor, the communication interface and the memory communicate with each other via the communication bus; The memory is used to store computer programs; The processor is used to implement the method steps described in any one of claims 1-8 when executing the program stored in the memory.

Citation Information

Patent Citations

  • Deep learning-based remote sensing image cultivated land non-agricultural change detection method

    CN115690591A

  • Multi-scale mask-based training-free defect detection method and defect detection equipment

    CN119168955A

  • Product defect unsupervised semantic segmentation method and device based on pre-training model

    CN119600296A

  • Few-sample defect detection method based on prototype prompt fine-tuning visual basic model SAM

    CN119624927A

  • Method for detecting defects on images of composite articles

    WO2024112226A1