Deep learning-based weld inspection methods and devices
By employing a deep learning-based weld inspection method, which utilizes multi-scale feature extraction and global feature processing, the limitations of traditional weld inspection methods in terms of accuracy and efficiency are overcome. This enables efficient and accurate inspection of pipeline welds, and it is particularly adaptable to micro-cracks and complex weld structures.
Patent Information
- Application Number
- CN202511098539.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-08-06
AI Technical Summary
Traditional weld inspection methods have limitations in terms of accuracy, efficiency, and applicability, especially when dealing with complex welded structures or minor defects, making it difficult to meet the demands of modern industry for speed, accuracy, and automation.
A deep learning-based weld detection method is adopted. By acquiring images from different regions of the pipeline, multi-scale feature extraction, feature fusion, and global feature processing are performed to determine the key regions of the pipeline and detect the weld regions. This includes denoising, feature fusion, image segmentation and embedding, and global feature processing. By combining FPN, ViT, and self-attention networks, the detection capability for complex backgrounds and minor defects is improved.
It achieves high-precision and high-speed inspection of pipeline welds, effectively capturing weld features at different scales. It is particularly adaptable to micro-cracks and complex weld structures, improving inspection efficiency and accuracy.
Smart Images

Figure CN120598949B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robotics, and in particular to a method and apparatus for weld inspection based on deep learning. Background Technology
[0002] With industrial development, pipelines, as a crucial infrastructure, are widely used in various fields such as petroleum, natural gas, chemical, and power. The welding quality of pipelines directly affects the safety and reliability of the entire system; therefore, weld inspection plays a vital role in pipeline construction and maintenance. Traditional pipeline weld inspection methods mostly employ manual inspection or physical testing techniques, such as X-ray inspection, ultrasonic testing, and magnetic particle testing. While these methods can detect weld defects to some extent, they have some significant limitations, especially in terms of inspection accuracy, efficiency, and applicability.
[0003] Traditional inspection methods have significant limitations. Manual inspection requires technicians to interpret welds based on experience, which not only carries the risk of subjectivity and misjudgment but also requires substantial manpower and time. While X-ray inspection can reveal internal defects in pipe welds, the equipment is expensive, the operation is complex, and it has high requirements for the operating environment. Furthermore, it cannot perform real-time inspections, impacting pipeline production efficiency. Although ultrasonic testing can accurately assess weld quality, its sensitivity to surface defects is low, especially for detecting minute cracks or hidden defects. While these traditional methods can identify some common weld defects, they still suffer from insufficient accuracy and low efficiency when dealing with complex weld structures or minute defects, failing to meet the modern industrial demand for rapid, accurate, and automated pipeline inspection.
[0004] In recent years, with the rapid development of deep learning technology, especially in the field of computer vision, artificial intelligence has made breakthrough progress in image processing. Deep learning-based image analysis methods can automatically learn features from large amounts of data and perform accurate classification and recognition, and have been widely applied in medical imaging, autonomous driving, industrial inspection, and many other fields. In pipeline weld inspection, deep learning algorithms can overcome the limitations of traditional methods, automatically extracting complex features from weld images and identifying subtle defects in the weld, achieving high accuracy and efficiency. However, existing deep learning models mostly rely on local feature extraction and lack effective integration of global information, making it difficult to handle complex backgrounds and minute defects in images. Furthermore, traditional convolutional neural networks (CNNs) may face problems such as high computational resource consumption and limited accuracy when processing large-scale, high-resolution images. Although some advanced image processing techniques such as U-Net and Faster R-CNN have been applied to weld inspection, they still suffer from low processing efficiency, poor model stability, and insufficient adaptability to complex welding structures. Summary of the Invention
[0005] This invention provides a weld inspection method and apparatus based on deep learning, which addresses the technical problem that traditional methods only focus on local information.
[0006] This invention provides a deep learning-based weld inspection method for pipeline robots; the method includes:
[0007] First images are acquired in different areas of the pipeline; the different areas include straight pipe sections, bends, and joints.
[0008] The first image is preprocessed to obtain a first texture feature and a first morphological feature; the preprocessing includes at least multi-scale feature extraction.
[0009] The key regions of the pipeline are determined based on the first texture features and the first morphological features;
[0010] The weld area of the pipeline is inspected based on the key area.
[0011] In some embodiments, the method further includes:
[0012] Obtain the acquisition time, pipe location, and image number corresponding to the first image; the acquisition time, pipe location, and image number are used to locate the source of the first image.
[0013] In some embodiments, the preprocessing further includes denoising, feature fusion, image segmentation and embedding, and global feature processing; the preprocessing of the first image yields first texture features and first morphological features; the preprocessing at least includes multi-scale feature extraction, including:
[0014] The first image is denoised to obtain the second image;
[0015] Multi-scale feature extraction is performed on the second image to obtain first features at different scales;
[0016] The first feature is subjected to feature fusion processing to obtain the second feature;
[0017] Based on the second feature, image segmentation and embedding are performed to obtain the third image;
[0018] Global feature processing is performed on the third image to obtain the first texture feature and the first morphological feature.
[0019] In some embodiments, the step of denoising the first image to obtain the second image includes:
[0020] By removing environmental noise or image interference from the first image, a fourth image is obtained;
[0021] The fourth image is then normalized in terms of brightness and contrast to obtain the second image.
[0022] In some implementations, the step of performing multi-scale feature extraction on the second image to obtain first features at different scales includes:
[0023] The second image is subjected to multi-scale feature extraction using a pre-defined feature pyramid network (FPN) to obtain the first features at different scales.
[0024] In some implementations, the step of performing image segmentation and embedding based on the second feature to obtain a third image includes:
[0025] The second feature is processed into image blocks using a preset visual transformer (ViT) to obtain the fifth image;
[0026] Determine the position encoding of the fifth image;
[0027] The third image is obtained based on the location encoding embedding process.
[0028] In some implementations, performing global feature processing on the third image to obtain the first texture feature and the first morphological feature includes:
[0029] The third image is subjected to global feature processing through a pre-defined multi-layered self-attention network to obtain the first texture feature and the first morphological feature.
[0030] In some implementations, determining the key regions of the pipe based on the first texture feature and the first morphological feature includes:
[0031] The similarity between various positions in the pipeline is determined based on the first texture feature and the first morphological feature;
[0032] The similarity is normalized to obtain the center attention weights for each region in the pipeline.
[0033] The critical areas of the pipeline are determined based on the attention weights.
[0034] In some embodiments, the method further includes:
[0035] The weld area is classified into defects to obtain the corresponding defect types; the defect types include at least cracks, porosity, incomplete penetration, and slag inclusions.
[0036] Determine the confidence score for the defect type;
[0037] The defect type and confidence score are fed back to the user corresponding to the pipeline robot; the defect type and confidence score are used by the user to query and retrieve historical data.
[0038] This invention also provides a deep learning-based weld inspection device for use in pipeline robots; the device includes:
[0039] An acquisition unit is used to acquire a first image in different regions of the pipeline; the different regions include straight pipe sections, bends, and joints.
[0040] A processing unit is configured to preprocess the first image to obtain first texture features and first morphological features; the preprocessing includes at least multi-scale feature extraction.
[0041] A determining unit is configured to determine the key regions of the pipe based on the first texture features and the first morphological features;
[0042] The detection unit is used to detect the weld area of the pipeline based on the key area.
[0043] This invention provides a deep learning-based weld inspection device, the device comprising: a processor and a memory for storing a computer program capable of running on the processor, wherein the processor, when running the computer program, executes the steps of any of the methods described above.
[0044] This invention provides a storage medium storing a computer program; when the computer program is executed by a processor, it implements the steps of any of the methods described above.
[0045] This invention provides a deep learning-based weld detection method applied to a pipeline robot. The method includes: acquiring first images of different regions of a pipeline; the different regions include straight pipe sections, bends, and joints; preprocessing the first images to obtain first texture features and first morphological features; the preprocessing includes at least multi-scale feature extraction; determining key regions of the pipeline based on the first texture features and the first morphological features; and detecting the weld region of the pipeline based on the key regions. By employing the technical solution of this application, preprocessing the first images of the pipeline, including straight pipe sections, bends, and joints, including multi-scale feature extraction, yields first texture features and first morphological features, thereby determining the key regions of the pipeline and detecting the weld region. That is, through multi-scale feature extraction, weld features at different scales are effectively captured, especially showing better adaptability for detecting micro-cracks and complex weld structures. Attached Figure Description
[0046] Figure 1 A schematic flowchart of a deep learning-based weld inspection method provided in an embodiment of the present invention;
[0047] Figure 2 A schematic diagram of a weld inspection device based on deep learning provided in an embodiment of the present invention;
[0048] Figure 3 This is a schematic diagram of the hardware structure of a weld inspection device based on deep learning according to an embodiment of the present invention. Detailed Implementation
[0049] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0050] The specific technical features described in the various embodiments in the detailed implementation can be combined in various ways without contradiction. For example, different implementation methods can be formed by combining different specific technical features. In order to avoid unnecessary repetition, the various possible combinations of the specific technical features in this invention will not be described separately.
[0051] It should also be noted that, in order to avoid obscuring the present invention with unnecessary details, only the structures and / or processing steps closely related to the present invention are shown in the accompanying drawings, while other details that are not closely related to the present invention are omitted.
[0052] Additionally, it should be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. In the following description, the terms "first," "second," etc., are used merely to distinguish different objects and do not indicate any similarity or connection between them. It should be understood that the directional descriptions such as "above," "below," "inside," and "outside" refer to the orientation under normal use conditions.
[0053] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the specific technical solutions of the invention will be further described in detail below with reference to the accompanying drawings of the embodiments of the present invention. The following embodiments are used to illustrate the present invention, but are not intended to limit the scope of the present invention.
[0054] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.
[0055] This invention provides a deep learning-based weld inspection method for use in pipeline robots; such as... Figure 1 As shown, Figure 1 This invention provides a schematic flowchart of a deep learning-based weld inspection method; the method includes:
[0056] Step S101: Obtain first images in different areas of the pipeline; the different areas include straight pipe sections, bends, and joints.
[0057] Step S102: Preprocess the first image to obtain the first texture feature and the first morphological feature; the preprocessing includes at least multi-scale feature extraction.
[0058] Step S103: Determine the key regions of the pipeline based on the first texture features and the first morphological features.
[0059] Step S104: Detect the weld area of the pipeline according to the key area.
[0060] In this embodiment, the deep learning-based weld detection method can be determined according to the actual situation and is not limited here. As an example, the deep learning-based weld detection method can be a deep learning-based pipeline robot weld detection method.
[0061] The pipeline robot can be determined based on actual conditions and is not limited here. As an example, the pipeline robot can be equipped with a high-definition camera and a powerful lighting system, enabling it to acquire clear visible light images inside the pipeline.
[0062] In step S101, the specific acquisition process for obtaining the first image in different areas of the pipe can be determined according to the actual situation and is not limited here. As an example, the robot is equipped with a high-definition camera and a powerful lighting system, which can acquire clear visible light images inside the pipe. The camera and lighting system are precisely mounted on the robot to ensure that complete images can be acquired in different areas of the pipe (e.g., straight pipe sections, bends, joints, etc.).
[0063] In step S102, the preprocessing can be determined according to the actual situation and is not limited here. As an example, the preprocessing includes denoising, multi-scale feature extraction, feature fusion, image segmentation and embedding, and global feature processing; the specific processing steps in preprocessing the first image to obtain the first texture feature and the first morphological feature can be determined according to the actual situation and are not limited here. As an example, the preprocessing of the first image to obtain the first texture feature and the first morphological feature; the preprocessing at least includes multi-scale feature extraction, which may include denoising the first image to obtain a second image; performing multi-scale feature extraction on the second image to obtain first features at different scales; performing feature fusion on the first features to obtain second features; performing image segmentation and embedding based on the second features to obtain a third image; and performing global feature processing on the third image to obtain the first texture feature and the first morphological feature.
[0064] In step S103, the specific process of determining the key regions of the pipeline based on the first texture features and the first morphological features can be determined according to the actual situation and is not limited here. As an example, determining the key regions of the pipeline based on the first texture features and the first morphological features may include determining the similarity between various positions in the pipeline based on the first texture features and the first morphological features; performing normalization processing based on the similarity to obtain the attention weight of each region in the pipeline; and determining the key regions of the pipeline based on the attention weight.
[0065] In step S104, the weld area of the pipeline can be determined according to the actual situation, and is not limited here. As an example, the weld area of the pipeline may include defective areas such as weld joints, cracks, and porosity.
[0066] The specific detection process for detecting the weld area of the pipeline based on the key region can be determined according to the actual situation and is not limited here. As an example, by combining the multi-scale features extracted by FPN, the global features captured by ViT, and the key regions focused by the self-attention mechanism, the model can accurately detect the weld area in the image. These areas may include defect areas such as weld joints, cracks, and porosity.
[0067] This invention provides a deep learning-based weld detection method. By preprocessing a first image of a pipeline, including straight sections, bends, and joints, with multi-scale feature extraction, first texture features and first morphological features are obtained. These are then used to determine the key areas of the pipeline and detect the weld areas. In other words, through multi-scale feature extraction, weld features at different scales are effectively captured, exhibiting better adaptability, particularly for detecting micro-cracks and complex weld structures.
[0068] In some embodiments, the method further includes:
[0069] Obtain the acquisition time, pipe location, and image number corresponding to the first image; the acquisition time, pipe location, and image number are used to locate the source of the first image.
[0070] In this embodiment, the acquisition time, pipeline location, and image number corresponding to the first image can be stored in the robot's local storage device.
[0071] In practical applications, the robot stores the acquired visible light images on local storage devices. Each image is accompanied by relevant metadata, such as acquisition time, pipeline location, and image number, ensuring that subsequent data processing and analysis can trace and accurately locate the source of each image.
[0072] In some embodiments, the preprocessing further includes denoising, feature fusion, image segmentation and embedding, and global feature processing; the preprocessing of the first image yields a first texture feature and a first morphological feature; the preprocessing at least includes multi-scale feature extraction, including:
[0073] The first image is denoised to obtain the second image;
[0074] Multi-scale feature extraction is performed on the second image to obtain first features at different scales;
[0075] The first feature is subjected to feature fusion processing to obtain the second feature;
[0076] Based on the second feature, image segmentation and embedding are performed to obtain the third image;
[0077] Global feature processing is performed on the third image to obtain the first texture feature and the first morphological feature.
[0078] In this embodiment, the preprocessing further includes denoising, feature fusion, image segmentation and embedding, and global feature processing. The denoising can be understood as noise removal. The feature fusion can be understood as the fusion of features at different scales, or feature fusion and output. The image segmentation and embedding can be simply referred to as image segmentation and embedding. The global feature processing can be understood as noise removal. The denoising can be understood as capturing the dependencies between different regions in the image globally using a self-attention mechanism.
[0079] The specific processing steps for denoising the first image to obtain the second image can be determined based on actual conditions and are not limited here. As an example, denoising the first image to obtain the second image may include removing environmental noise or image interference from the first image to obtain a fourth image; and then performing brightness and contrast normalization processing on the fourth image to obtain the second image. For example, removing environmental noise or image interference from the first image to obtain the fourth image can be achieved through a preset algorithm. The preset algorithm can be determined based on actual conditions and is not limited here. As an example, the preset algorithm can be a Gaussian filtering method. In practical applications, Gaussian filtering is used to remove noise from the acquired images, eliminating any environmental noise or image interference that may exist in the image, ensuring clearer details in the weld area. Next, brightness and contrast normalization processing is performed on the image to ensure consistent brightness and clarity in images taken from different pipe sections, making the weld area more prominent. This process can also be called data preprocessing.
[0080] The specific processing steps for multi-scale feature extraction of the second image to obtain first features at different scales can be determined according to the actual situation and are not limited here. As an example, multi-scale feature extraction of the second image to obtain first features at different scales may include using a preset Feature Pyramid Network (FPN) to extract features from the second image at multiple scales to obtain the first features at different scales. The specific extraction process for using the preset FPN to extract features from the second image at multiple scales to obtain the first features at different scales can be determined according to the actual situation and is not limited here. As an example, multi-scale feature extraction of the second image using the preset FPN to obtain the first features at different scales may be using a preset dynamic weighting algorithm to extract features from the second image at multiple scales to obtain the first features at different scales. The preset dynamic weighting algorithm can be determined according to the actual situation and is not limited here. As an example, the preset dynamic weighting algorithm can be understood as a dynamic weighting mechanism, including bottom path feature extraction, top path feature extraction, and dynamic weighting. In practical applications, a Feature Pyramid Network (FPN) is used to extract features from the input image at multiple scales. FPN extracts features at different levels through bottom-up convolutional paths and performs feature fusion using top-down paths, ensuring the capture of multi-level information from local details to global features. To improve the model's ability to detect complex weld features, this invention introduces a dynamic scale weighting mechanism, which automatically adjusts the feature fusion weights at different scales based on the complexity of the image region, making the model focus more on the key areas of weld defects. Dynamic weighting mechanism formula: Bottom-down path (feature extraction): Top-down path (feature fusion): Dynamic weighting mechanism: ;in, These are the dynamic weighting coefficients for the feature maps at each scale. This is the final feature map after weighted fusion. The coefficients are determined by the image content and enable the model to adaptively adjust its focus on the weld area.
[0081] The specific processing steps for feature fusion of the first feature to obtain the second feature can be determined based on actual circumstances and are not limited here. As an example, Feature Fusion Network (FPN) can be used to perform feature fusion on the first feature to obtain the second feature. In practical applications, FPN ensures the effective extraction of multi-scale information, from micro-cracks to larger weld areas, through the fusion of features at different scales. After feature fusion, the feature map generated by the model will contain the overall structure and subtle defects of the weld, which is helpful for subsequent detection and classification.
[0082] The specific processing steps for image segmentation and embedding based on the second feature to obtain the third image can be determined according to actual conditions and are not limited here. As an example, the image segmentation and embedding based on the second feature to obtain the third image may include using a preset visual transformer (ViT) to segment the second feature into images to obtain a fifth image; determining the position code of the fifth image; and embedding based on the position code to obtain the third image. The preset visual transformer can be determined according to actual conditions and is not limited here. The preset visual transformer can be simply referred to as a visual transformer. Using the preset visual transformer (ViT) to segment the second feature into images to obtain the fifth image may involve using the preset visual transformer (ViT) to perform adaptive tile partitioning processing on the second feature to obtain the fifth image; wherein the adaptive tile partitioning processing can be an adaptive tile partitioning algorithm, for example... ;in It is an adaptive selection. The specific determination process for determining the position code of the fifth image can be determined according to the actual situation and is not limited here. As an example, determining the position code of the fifth image can be done by determining the position code of the fifth image through a preset position coding algorithm. The preset position coding algorithm can be determined according to the actual situation and is not limited here. As an example, the preset position coding algorithm can be... ;here, For adaptive selection of tiles, The image is a flattened patch. In practical applications, the Visual Transformer (ViT) divides the feature map output by the FPN into multiple patches. The size of the patches is adaptively adjusted according to the complexity of the image region to ensure that important weld features such as fine cracks are fully represented. Each patch is flattened and has a positional encoding added before being fed into the ViT model as an input sequence.
[0083] Adaptive tile partitioning formula: ;in It is adaptively selected. Position encoding: ;here, For adaptive selection of tiles, This is the flattened image tile. This process can also be called image tiling and embedding.
[0084] The specific processing steps for performing global feature processing on the third image to obtain the first texture feature and the first morphological feature can be determined according to the actual situation and are not limited here. As an example, performing global feature processing on the third image to obtain the first texture feature and the first morphological feature may include performing global feature processing on the third image through a preset multi-layer self-attention network to obtain the first texture feature and the first morphological feature. The preset multi-layer self-attention network can be determined according to the actual situation and is not limited here.
[0085] As an example, the pre-defined multi-layered self-attention network may include a hierarchical self-attention algorithm, for example, ;in, i Indicates the level number, , , These are the query, key, and value vectors of the current layer. The hierarchical self-attention algorithm can also be understood as a hierarchical self-attention mechanism. The first texture feature and the first morphological feature can be determined according to the actual situation and are not limited here. As an example, the first texture feature can be a low-level texture feature; the first morphological feature can be a high-level morphological feature. In practical applications, ViT captures the dependencies between various regions in an image globally through a self-attention mechanism. Especially for the analysis of complex patterns in weld seam regions, ViT can effectively identify regions with blurred weld seam boundaries or irregular shapes. To further enhance the ability to capture global features, this invention introduces a hierarchical self-attention mechanism, which processes the low-level texture features and high-level morphological features of the weld seam region through a multi-layered self-attention network.
[0086] Hierarchical self-attention formula: ;in, i Indicates the level number, , , These are the query, key, and value vectors for the current layer. The hierarchical self-attention mechanism can learn low-level texture features and high-level morphological features in weld seam images at different levels.
[0087] In some embodiments, the denoising process on the first image to obtain the second image includes:
[0088] By removing environmental noise or image interference from the first image, a fourth image is obtained;
[0089] The fourth image is then normalized in terms of brightness and contrast to obtain the second image.
[0090] In this embodiment, the process of removing environmental noise or image interference from the first image to obtain a fourth image can be achieved by using a preset algorithm. The preset algorithm can be determined based on actual conditions and is not limited here. As an example, the preset algorithm can be a Gaussian filtering method.
[0091] In practical applications, Gaussian filtering is used to remove noise from the acquired images, eliminating any environmental noise or image interference that may be present, ensuring clearer details in the weld area. Next, brightness and contrast are normalized to ensure consistent brightness and clarity in images taken from different pipe sections, making the weld area stand out more. This process can also be called data preprocessing.
[0092] In some embodiments, the multi-scale feature extraction of the second image to obtain first features at different scales includes:
[0093] The second image is subjected to multi-scale feature extraction using a pre-defined feature pyramid network (FPN) to obtain the first features at different scales.
[0094] In this embodiment, the preset feature pyramid network can be determined according to the actual situation, and is not limited here. In practical applications, the preset feature pyramid network can be simply referred to as the feature pyramid network.
[0095] The process of extracting multi-scale features from the second image using a preset Feature Pyramid Network (FPN) to obtain the first features at different scales can be determined based on actual conditions and is not limited here. As an example, the process of extracting multi-scale features from the second image using a preset Feature Pyramid Network (FPN) to obtain the first features at different scales can be achieved by using a preset dynamic weighting algorithm to extract multi-scale features from the second image using a preset Feature Pyramid Network (FPN). The preset dynamic weighting algorithm can be determined based on actual conditions and is not limited here. As an example, the preset dynamic weighting algorithm can be understood as a dynamic weighting mechanism, including bottom-path feature extraction, top-path feature extraction, and dynamic weighting.
[0096] In practical applications, a Feature Pyramid Network (FPN) is used to extract features at multiple scales from the input image. The FPN extracts features at different levels through bottom-up convolutional paths while simultaneously using top-down paths for feature fusion, ensuring the capture of multi-level information from local details to global features. To improve the model's ability to detect complex weld features, this invention introduces a dynamic scale weighting mechanism, automatically adjusting the feature fusion weights at different scales based on the complexity of the image region, allowing the model to focus more on the critical areas of weld defects.
[0097] Dynamic weighting mechanism formula: Bottom path (feature extraction): Top-down path (feature fusion): Dynamic weighting mechanism: ;in, These are the dynamic weighting coefficients for the feature maps at each scale. This is the final feature map after weighted fusion. The coefficients are determined by the image content and enable the model to adaptively adjust its focus on the weld area.
[0098] In some embodiments, the image segmentation and embedding process based on the second feature to obtain a third image includes:
[0099] The second feature is processed into image blocks using a preset visual transformer (ViT) to obtain the fifth image;
[0100] Determine the position encoding of the fifth image;
[0101] The third image is obtained based on the location encoding embedding process.
[0102] In this embodiment, the preset visual transformer can be determined according to the actual situation, and is not limited here. The preset visual transformer can be simply referred to as the visual transformer.
[0103] The second feature is processed into image blocks using a preset visual transformer (ViT) to obtain the fifth image. Alternatively, the second feature can be processed into adaptive patch partitioning using the preset visual transformer (ViT) to obtain the fifth image. The adaptive patch partitioning can be an adaptive patch partitioning algorithm, for example... ;in It is an adaptive selection.
[0104] The specific determination process for determining the position code of the fifth image can be determined according to actual circumstances and is not limited here. As an example, determining the position code of the fifth image can be done through a preset position coding algorithm. The preset position coding algorithm can be determined according to actual circumstances and is not limited here. As an example, the preset position coding algorithm can be... ;here, For adaptive selection of tiles, This is the flattened tile.
[0105] In practical applications, the Visual Transformer (ViT) divides the feature map output by the FPN into multiple patches. The size of each patch is adaptively adjusted according to the complexity of the image region to ensure that important weld features, such as fine cracks, are fully represented. Each patch is flattened and encoded with a position, then fed into the ViT model as an input sequence. The adaptive patch partitioning formula is as follows: ;in It is an adaptive selection.
[0106] Location coding: ;here, For adaptive selection of tiles, This is the flattened tile.
[0107] In some embodiments, performing global feature processing on the third image to obtain the first texture feature and the first morphological feature includes:
[0108] The third image is subjected to global feature processing through a pre-defined multi-layered self-attention network to obtain the first texture feature and the first morphological feature.
[0109] In this embodiment, the preset multi-layered self-attention network can be determined according to actual conditions and is not limited here. As an example, the preset multi-layered self-attention network may include a hierarchical self-attention algorithm, for example, ;in, i Indicates the level number, , , These are the query, key, and value vectors for the current layer. The hierarchical self-attention algorithm can also be understood as a hierarchical self-attention mechanism.
[0110] The first texture feature and the first morphological feature can be determined according to the actual situation, and are not limited here. As an example, the first texture feature can be a low-level texture feature; the first morphological feature can be a high-level morphological feature.
[0111] In practical applications, ViT captures the dependencies between regions in an image globally through a self-attention mechanism. Particularly in the analysis of complex patterns in weld seams, ViT can effectively identify regions with blurred weld boundaries or irregular shapes. To further enhance the ability to capture global features, this invention introduces a hierarchical self-attention mechanism, using multi-layered self-attention networks to process low-level texture features and high-level morphological features of the weld seam region separately. The hierarchical self-attention formula is as follows: ;in, i Indicates the level number, , , These are the query, key, and value vectors for the current layer. The hierarchical self-attention mechanism can learn low-level texture features and high-level morphological features in weld images at different levels.
[0112] In some embodiments, determining the key regions of the pipeline based on the first texture feature and the first morphological feature includes:
[0113] The similarity between various positions in the pipeline is determined based on the first texture feature and the first morphological feature;
[0114] The similarity is normalized to obtain the center attention weights for each region in the pipeline.
[0115] The critical areas of the pipeline are determined based on the attention weights.
[0116] In this embodiment, the process of determining the similarity between various positions in the pipe based on the first texture feature and the first morphological feature can be determined according to actual circumstances and is not limited here. As an example, determining the similarity between various positions in the pipe based on the first texture feature and the first morphological feature can be done by using a preset self-attention algorithm based on the first texture feature and the first morphological feature. The preset self-attention algorithm can be determined according to actual circumstances and is not limited here. As an example, the preset self-attention algorithm can be simply referred to as a self-attention mechanism, for example... ;in, These are region weighting coefficients calculated based on local features (such as cracks and pores), used to enhance the focus on key areas.
[0117] The specific processing steps for normalizing based on the similarity to obtain the central attention weights of each region in the pipeline can be determined according to the actual situation and are not limited here. As an example, the normalization process for obtaining the central attention weights of each region in the pipeline based on the similarity can be performed using Softmax to normalize based on the similarity.
[0118] For example, the attention weights can be referenced .
[0119] The specific process for determining the critical regions of the pipeline based on the attention weights can be determined according to the actual situation and is not limited here. As an example, the self-attention mechanism focuses the model's attention on key parts of the weld region, such as defects like cracks and porosity, through calculated weighting coefficients. By dynamically adjusting the weights, the model can balance local details and global structure, improving the detection capability of weld defects, especially in situations with complex backgrounds or uneven lighting, ensuring that these key features are not ignored.
[0120] Adaptive weight update formula: ;in, It is an additional region weighting function used to guide the model to focus more on key parts of the weld region, such as cracks or defects.
[0121] In practical applications, the self-attention mechanism determines which regions in an image should be focused on by calculating the similarity between different locations. The feature vector of each pixel is multiplied by the feature vectors of other locations in the image to calculate a similarity score, which is then normalized using Softmax to obtain the attention weights. The self-attention calculation formula is as follows: ;in, These are region-weighted coefficients calculated based on local features (such as cracks and pores), used to enhance attention to key areas. Attention weight calculation: This process allows the model to weight and focus based on the features of each region in the image, ensuring the detection of weld areas, especially minute defects. The self-attention mechanism uses calculated weighting coefficients to focus the model's attention on key parts of the weld area, such as cracks and porosity. By dynamically adjusting the weights, the model can balance local details and global structure, improving its ability to detect weld defects, especially in complex backgrounds or uneven lighting conditions, ensuring that these key features are not overlooked. Adaptive weight update formula: ;in, It is an additional region weighting function used to guide the model to focus more on key parts of the weld region, such as cracks or defects.
[0122] In some embodiments, the method further includes:
[0123] The weld area is classified into defects to obtain the corresponding defect types; the defect types include at least cracks, porosity, incomplete penetration, and slag inclusions.
[0124] Determine the confidence score for the defect type;
[0125] The defect type and confidence score are fed back to the user corresponding to the pipeline robot; the defect type and confidence score are used by the user to query and retrieve historical data.
[0126] In this embodiment, the defect types include at least cracks, porosity, incomplete penetration, and slag inclusions. In practical applications, by combining multi-scale features extracted by FPN, global features captured by ViT, and key regions focused by the self-attention mechanism, the model can accurately detect weld areas in images. These areas may include weld joints, cracks, porosity, and other defect areas.
[0127] In practical applications, after detecting a weld area, the system uses a classification layer in a deep learning model to further analyze each weld area. The classification layer's role is to make decisions based on previously extracted features (multi-scale features extracted by FPN and global features captured by ViT), determining whether a defect exists in the area and further identifying the type of defect. The types of weld defects collected include... After sufficient training, the deep learning model can automatically distinguish different types of defects based on these features and perform annotation and classification.
[0128] Each detected weld area is classified based on the output of the deep learning model. If the model determines that the area is a crack, it outputs the crack type and further indicates the crack's location, length, depth, and other information. The model's output includes not only the defect type but also a related confidence score, representing the reliability of the defect classification. This score helps in subsequent assessments of the defect's severity and whether repair is necessary.
[0129] After classification, the system will label the identified defect types and corresponding areas and overlay them onto the image for easy viewing and confirmation by the operator. The specific location, type, and other relevant information of the defects will be displayed on the image interface, providing clear detection results.
[0130] In practical applications, as an example, a deep learning-based weld detection method can be specifically described as a deep learning-based pipeline robot weld detection method. First, the pipeline robot is equipped with a high-definition camera and a powerful lighting system, enabling it to acquire clear visible light images inside the pipeline. The robot moves along a preset path within the pipeline, ensuring that each weld area is fully captured. Images are stored in real-time with associated metadata for subsequent analysis and processing. Next, Gaussian filtering is used to denoise the images, removing environmental noise or image interference. Brightness and contrast normalization ensures image consistency across different pipeline sections, highlighting weld area features. The preprocessed images are then fed into a Feature Pyramid Network (FPN) for multi-scale feature extraction. The FPN extracts local features through a bottom-up convolutional path and performs feature fusion through a top-down path, further enhancing the capture of multi-level information from details to the global picture. To improve the detection capability for complex weld features, this study introduces a dynamic scale weighting mechanism, automatically adjusting the feature fusion weights based on the complexity of the image region, thus giving greater attention to weld defect areas. After feature extraction, a Visual Transformer (ViT) is used for global feature analysis. ViT (Vibes in Welding) divides images into blocks and inputs them into the model through adaptive patching and positional encoding. It captures the dependencies between regions through a self-attention mechanism, demonstrating its powerful ability, especially in handling areas with blurred weld boundaries or irregular shapes. To further optimize ViT's performance in weld detection, this study also constructs a hierarchical self-attention mechanism. This mechanism uses a multi-layered network to process low-level texture features and high-level morphological features of the weld, improving the model's ability to resolve complex weld images. The self-attention mechanism adjusts the focus area by calculating the similarity between different locations in the image, particularly enhancing details such as microcracks and porosity. Through adaptive weighted calculation, the model can effectively focus attention on weld defect areas, ensuring accurate defect detection even in complex backgrounds and uneven lighting conditions. Finally, by combining multi-scale features extracted by FPN, global features captured by ViT, and key areas focused by the self-attention mechanism, the algorithm can accurately detect weld areas and classify defects. The system provides real-time feedback on the detection results, offering the type and location of weld defects, allowing operators to perform repair operations based on the detection information. All data and image results are stored and can be used for subsequent analysis and model optimization. This algorithm improves the accuracy and efficiency of pipeline weld inspection by introducing dynamically weighted FPN, innovative ViT, and a self-attention mechanism, providing a highly efficient automated inspection solution for pipeline maintenance and safety assurance. The specific steps include:
[0131] 1. Robot image acquisition.
[0132] Equipped with high-definition cameras and a powerful lighting system, the robot can capture clear visible light images inside pipelines. The camera and lighting system are precisely mounted on the robot to ensure complete images are acquired across different areas of the pipeline (straight sections, bends, joints, etc.). The acquired images primarily focus on the weld areas of the pipeline, which are typically the focus of inspection. The robot moves along a pre-set path inside the pipeline, ensuring that each weld location is fully covered. Because the robot has an independent lighting system, it can continue to operate in low-light or complex environments, ensuring that the image quality of the welds remains at an optimal level. Furthermore, the robot can adjust the intensity and angle of the lighting in real time to adapt to the lighting requirements of different locations within the pipeline, further improving image clarity and visibility.
[0133] 2. Image saving.
[0134] The robot stores the acquired visible light images on a local storage device. Each image is accompanied by relevant metadata, such as acquisition time, pipeline location, and image number, ensuring that subsequent data processing and analysis can trace and accurately locate the source of each image.
[0135] 3. Data preprocessing.
[0136] The acquired images were processed using Gaussian filtering to remove any environmental noise or image interference, ensuring clearer details in the weld area. Next, the images were normalized for brightness and contrast to ensure consistent brightness and clarity across different pipe sections, further emphasizing the weld area.
[0137] 4. Multi-scale feature extraction.
[0138] The Feature Pyramid Network (FPN) is used to extract multi-scale features from the input image. The FPN extracts features at different levels through bottom-up convolutional paths, while simultaneously using top-down paths for feature fusion, ensuring the capture of multi-level information from local details to global features. To improve the model's ability to detect complex weld features, this invention introduces a dynamic scale weighting mechanism, automatically adjusting the feature fusion weights at different scales based on the complexity of the image region, allowing the model to focus more on the critical areas of weld defects.
[0139] Dynamic weighting mechanism formula:
[0140] Bottom path (feature extraction):
[0141] ;
[0142] Top-down path (feature fusion):
[0143] ;
[0144] Dynamic weighting mechanism:
[0145] ;
[0146] in, These are the dynamic weighting coefficients for the feature maps at each scale. This is the final feature map after weighted fusion. The coefficients are determined by the image content and enable the model to adaptively adjust its focus on the weld area.
[0147] 5. Feature fusion and output.
[0148] FPN ensures the effective extraction of multi-scale information, from micro-cracks to larger weld areas, through the fusion of features at different scales. After feature fusion, the feature map generated by the model will contain the overall structure of the weld and subtle defects, which is helpful for subsequent detection and classification.
[0149] 6. Image segmentation and embedding.
[0150] The Visual Transformer (ViT) divides the feature map output by the FPN into multiple patches. The size of the patches is adaptively adjusted according to the complexity of the image region to ensure that important weld features such as fine cracks are fully represented. Each patch is flattened and encoded with a position, and then fed into the ViT model as an input sequence.
[0151] Adaptive tile partitioning formula:
[0152] ;
[0153] in It is an adaptive selection.
[0154] Location coding:
[0155] ;
[0156] here, For adaptive selection of tiles, This is the flattened tile.
[0157] 7. Global feature processing.
[0158] ViT captures the dependencies between regions in an image globally through a self-attention mechanism. Particularly in the analysis of complex patterns in weld seams, ViT can effectively identify regions with blurred weld boundaries or irregular shapes. To further enhance the ability to capture global features, this invention introduces a hierarchical self-attention mechanism, using multi-layered self-attention networks to process low-level texture features and high-level morphological features of the weld seam region separately.
[0159] Hierarchical self-attention formula:
[0160] ;
[0161] in, i Indicates the level number, , , These are the query, key, and value vectors for the current layer. The hierarchical self-attention mechanism can learn low-level texture features and high-level morphological features in weld images at different levels.
[0162] 8. Calculate similarity and weight.
[0163] The self-attention mechanism determines which regions in an image should be focused on by calculating the similarity between different locations in the image. The feature vector of each pixel is multiplied by the feature vectors of other locations in the image to calculate a similarity score, which is then normalized using Softmax to obtain the attention weights.
[0164] Self-attention calculation formula:
[0165] ;
[0166] in, These are region weighting coefficients calculated based on local features (such as cracks and pores), used to enhance the focus on key areas.
[0167] Attention weight calculation:
[0168] ;
[0169] This process allows the model to perform weighted focusing based on the features of each region in the image, ensuring the detection of weld areas, especially minute defects.
[0170] 9. Focus on key areas.
[0171] The self-attention mechanism uses calculated weighting coefficients to focus the model's attention on key parts of the weld area, such as cracks and porosity. By dynamically adjusting the weights, the model can balance local details and global structure, improving its ability to detect weld defects, especially in complex backgrounds or uneven lighting conditions, ensuring that these key features are not overlooked.
[0172] Adaptive weight update formula:
[0173] ;
[0174] in, It is an additional region weighting function used to guide the model to focus more on key parts of the weld region, such as cracks or defects.
[0175] 10. Weld area inspection.
[0176] By combining multi-scale features extracted by FPN, global features captured by ViT, and key regions focused by the self-attention mechanism, the model can accurately detect weld areas in images. These areas may include defective regions such as weld joints, cracks, and porosity.
[0177] 11. Defect classification.
[0178] After detecting weld seam areas, the system uses a classification layer in a deep learning model to further analyze each weld seam area. The classification layer's role is to make decisions based on previously extracted features (multi-scale features extracted by FPN and global features captured by ViT), determining whether a defect exists in the area and further identifying the type of defect. The types of weld seam defects collected include [list of types]. After sufficient training, the deep learning model can automatically distinguish different types of defects based on these features and perform annotation and classification.
[0179] Each detected weld area is classified based on the output of the deep learning model. If the model determines that the area is a crack, it outputs the crack type and further indicates the crack's location, length, depth, and other information. The model's output includes not only the defect type but also a related confidence score, representing the reliability of the defect classification. This score helps in subsequent assessments of the defect's severity and whether repair is necessary.
[0180] After classification, the system will label the identified defect types and corresponding areas and overlay them onto the image for easy viewing and confirmation by the operator. The specific location, type, and other relevant information of the defects will be displayed on the image interface, providing clear detection results.
[0181] 12. Real-time feedback.
[0182] The inspection results will be fed back to the robot control system in real time, and the operator can view the results on the interface, including the location, type, and severity of weld defects. All inspection data and image results will be recorded in detail and stored in the cloud or local database. The inspection data for each weld will be categorized and stored according to time sequence, inspection location, welding process, etc., for easy retrieval and historical data query.
[0183] After completing the assigned task, the robot automatically returns to its initial position or preset sleep point along the planned path and uploads the task data. Based on the high-precision 3D point cloud constructed in real time by cameras and LiDAR, it simultaneously records diverse information such as changes in environmental status, path trajectory, and obstacle characteristics. This information is stored in batches in local memory or transmitted to the cloud server via 5G network, forming a structured database for subsequent applications such as progress monitoring, defect analysis, and safety assessment. The control device generates a visual report based on the task execution log (including path optimization suggestions, abnormal event markers, and energy consumption statistics). Operators can retrieve data in real time and issue new instructions through the human-machine interface, realizing closed-loop optimization of "perception-decision-execution" and effectively improving the factory's operation and maintenance efficiency and intelligence level.
[0184] This application proposes a pipeline weld detection method based on Feature Pyramid Network (FPN), Vision Transformer (ViT), and self-attention mechanism. The core advantage of this method lies in the combination of multi-scale feature extraction, global information processing, and self-attention mechanism. FPN can extract features from images at multiple scales, effectively capturing weld features at different scales, and is particularly adaptable to detecting micro-cracks and complex weld structures. By introducing ViT, this application can capture global features in the image, overcoming the limitations of traditional methods that only focus on local information. This is especially beneficial when dealing with welds with complex shapes and blurred boundaries, effectively improving recognition accuracy. The self-attention mechanism automatically focuses on key regions in the image, making the model more sensitive to weld defect areas, especially in situations with high background noise, effectively reducing false positives and false negatives.
[0185] Compared to traditional deep learning methods, the model in this application not only ensures high-precision detection but also achieves faster inference speed and lower computational resource consumption through optimized computational architecture, adapting to the real-time detection needs of industrial environments. The application of this technology not only offers significant advantages in accuracy and efficiency but also provides stronger robustness and adaptability in complex pipeline weld inspection tasks, capable of handling different pipeline materials, welding processes, and imaging conditions. Through the application of this technology, the automation level of pipeline weld inspection will be effectively improved, reducing human error and inspection time, and enhancing the efficiency and safety of pipeline maintenance and monitoring.
[0186] Based on the same inventive concept as described above Figure 2 This is a schematic diagram of a deep learning-based weld inspection device provided in an embodiment of the present invention, as shown below. Figure 2 As shown, the device 200 is installed on a pipeline robot; the device 200 includes:
[0187] Acquisition unit 201 is used to acquire a first image in different regions of the pipeline; the different regions include straight pipe sections, bends, and joints.
[0188] Processing unit 202 is used to preprocess the first image to obtain first texture features and first morphological features; the preprocessing includes at least multi-scale feature extraction;
[0189] Determining unit 203 is used to determine the key region of the pipe based on the first texture feature and the first morphological feature;
[0190] The detection unit 204 is used to detect the weld area of the pipeline based on the critical area.
[0191] In some embodiments, the acquisition unit 201 is further configured to acquire the acquisition time, pipe position, and image number corresponding to the first image; the acquisition time, the pipe position, and the image number are used to locate the source of the first image.
[0192] In some embodiments, the preprocessing further includes denoising, feature fusion, image segmentation and embedding, and global feature processing; the processing unit 202 is further configured to denoise the first image to obtain a second image; extract multi-scale features from the second image to obtain first features at different scales; perform feature fusion on the first features to obtain second features; perform image segmentation and embedding based on the second features to obtain a third image; and perform global feature processing on the third image to obtain the first texture feature and the first morphological feature.
[0193] In some embodiments, the processing unit 202 is further configured to remove environmental noise or image interference from the first image to obtain a fourth image; and to perform brightness and contrast normalization processing on the fourth image to obtain the second image.
[0194] In some embodiments, the processing unit 202 is further configured to generate a global path for the target channel based on the dynamic safety distance model; and perform local dynamic obstacle avoidance processing based on the global path to obtain the path planning result for the target channel.
[0195] In some embodiments, the processing unit 202 is further configured to perform multi-scale feature extraction on the second image using a preset feature pyramid network (FPN) to obtain the first features at different scales.
[0196] In some embodiments, the processing unit 202 is further configured to perform image block processing on the second feature using a preset visual transformer ViT to obtain a fifth image; determine the position code of the fifth image; and obtain the third image based on the position code embedding processing.
[0197] In some embodiments, the processing unit 202 is further configured to perform image block processing on the second feature using a preset visual transformer ViT to obtain a fifth image; determine the position code of the fifth image; and obtain the third image based on the position code embedding processing.
[0198] In some embodiments, the processing unit 202 is further configured to perform global feature processing on the third image through a preset multi-layer self-attention network to obtain the first texture feature and the first morphological feature.
[0199] In some embodiments, the determining unit 203 is further configured to determine the similarity between various positions in the pipeline based on the first texture feature and the first morphological feature; perform normalization processing based on the similarity to obtain the attention weight of each region in the pipeline; and determine the key region of the pipeline based on the attention weight.
[0200] In some embodiments, the device 200 further includes a feedback unit; wherein,
[0201] The processing unit 202 is also used to perform defect classification processing on the weld area to obtain the defect type corresponding to the weld area; the defect type includes at least cracks, porosity, incomplete penetration, and slag inclusions.
[0202] The determining unit 203 is also used to determine the confidence score of the defect type;
[0203] The feedback unit is used to provide feedback on the defect type and the confidence score to the user corresponding to the pipeline robot; the defect type and the confidence score are used by the user to query and retrieve historical data.
[0204] It should be noted that the deep learning-based weld detection device provided in this embodiment of the invention and the deep learning-based weld detection method provided in the aforementioned embodiment of the invention belong to the same inventive concept. The meanings of the terms used here have been explained in detail above and will not be repeated here.
[0205] This invention also provides a storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0206] This invention also provides a deep learning-based weld inspection device, which includes a processor and a memory for storing a computer program that can run on the processor. When the processor runs the computer program, it executes the steps of the method embodiments described above stored in the memory.
[0207] Figure 3 This is a schematic diagram of a hardware structure of a deep learning-based weld inspection device according to an embodiment of the present invention. The deep learning-based weld inspection device 300 includes at least one processor 301 and a memory 302. Optionally, the deep learning-based weld inspection device 300 may further include at least one communication interface 303. The various components in the deep learning-based weld inspection device 300 are coupled together through a bus device 304. It is understood that the bus device 304 is used to realize the connection and communication between these components. In addition to a data bus, the bus device 304 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in… Figure 3 The general designated all buses as bus devices 304.
[0208] It is understood that memory 302 can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM); magnetic surface memory can be disk storage or magnetic tape storage. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), Sync Link Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM).The memory 302 described in this embodiment of the invention is intended to include, but is not limited to, these and any other suitable types of memory.
[0209] The memory 302 in this embodiment of the invention is used to store various types of data to support the operation of the deep learning-based weld inspection device 300. Examples of such data include any computer program for operating on the deep learning-based weld inspection device 300, and programs implementing the methods of this embodiment of the invention may be included in the memory 302.
[0210] The methods disclosed in the above embodiments of the present invention can be applied to processor 301, or implemented by processor 301. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. A general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of the present invention can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium, which is located in memory. The processor reads information from the memory and, in conjunction with its hardware, completes the steps of the aforementioned method.
[0211] In an exemplary embodiment, the deep learning-based weld inspection device 300 may be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers (MCUs), microprocessors, or other electronic components to perform the methods described above.
[0212] In the several embodiments provided by this invention, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another device, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed may be through some interfaces, and the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms. The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units; some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs. In addition, all functional units in the various embodiments of this invention can be integrated into one processing module, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated units can be implemented in hardware or in the form of hardware plus software functional units.
[0213] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention.
Claims
1. A weld inspection method based on deep learning, characterized in that, Applied to pipeline robots; the method includes: First images are acquired in different areas of the pipeline; the different areas include straight pipe sections, bends, and joints. The first image is preprocessed to obtain a first texture feature and a first morphological feature; the preprocessing includes at least multi-scale feature extraction. Determining the key regions of the pipeline based on the first texture feature and the first morphological feature; including: Construct a weld-background weight step transition branch network, process the first texture feature and the first morphological feature in parallel, and generate a feature map that causes the weld edge weight to jump step. Based on the feature map, graph tile nodes are established at multiple scales, and a graph similarity matrix is obtained through a similarity aggregation method based on graph tiles; The graph similarity matrix is subjected to iterative confidence-weight joint evolution: the feature representation is updated with the current weight, and the graph similarity is recalculated with the updated features. After multiple iterations, the final attention weight is obtained. The critical regions of the pipeline are determined based on the final attention weights. By fusing multi-scale features with global context within the key region, the weld area of the pipeline is detected.
2. The method according to claim 1, characterized in that, The method further includes: Obtain the acquisition time, pipe location, and image number corresponding to the first image; the acquisition time, pipe location, and image number are used to locate the source of the first image.
3. The method according to claim 1, characterized in that, The preprocessing also includes denoising, feature fusion, image segmentation and embedding, and global feature processing; The first image is preprocessed to obtain a first texture feature and a first morphological feature; The preprocessing includes at least multi-scale feature extraction, including: The first image is denoised to obtain the second image; Multi-scale feature extraction is performed on the second image to obtain first features at different scales; The first feature is subjected to feature fusion processing to obtain the second feature; Based on the second feature, image segmentation and embedding are performed to obtain the third image; Global feature processing is performed on the third image to obtain the first texture feature and the first morphological feature.
4. The method according to claim 3, characterized in that, The step of denoising the first image to obtain the second image includes: By removing environmental noise or image interference from the first image, a fourth image is obtained; The fourth image is then normalized in terms of brightness and contrast to obtain the second image.
5. The method according to claim 3, characterized in that, The step of performing multi-scale feature extraction on the second image to obtain first features at different scales includes: The second image is subjected to multi-scale feature extraction using a pre-defined feature pyramid network (FPN) to obtain the first features at different scales.
6. The method according to claim 3, characterized in that, The image segmentation and embedding process based on the second feature to obtain the third image includes: The second feature is processed into image blocks using a preset visual transformer (ViT) to obtain the fifth image; Determine the position encoding of the fifth image; The third image is obtained based on the location encoding embedding process.
7. The method according to claim 3, characterized in that, The step of performing global feature processing on the third image to obtain the first texture feature and the first morphological feature includes: The third image is subjected to global feature processing through a pre-defined multi-layered self-attention network to obtain the first texture feature and the first morphological feature.
8. The method according to any one of claims 1-7, characterized in that, The method further includes: The weld area is classified into defects to obtain the corresponding defect types; the defect types include at least cracks, porosity, incomplete penetration, and slag inclusions. Determine the confidence score for the defect type; The defect type and confidence score are fed back to the user corresponding to the pipeline robot; the defect type and confidence score are used by the user to query and retrieve historical data.
9. A weld inspection device based on deep learning, characterized in that, The device is installed on a pipeline robot; the device includes: An acquisition unit is used to acquire a first image in different regions of the pipeline; the different regions include straight pipe sections, bends, and joints. A processing unit is configured to preprocess the first image to obtain first texture features and first morphological features; the preprocessing includes at least multi-scale feature extraction. The determining unit is configured to determine the key regions of the pipe based on the first texture features and the first morphological features, including: Construct a weld-background weight step transition branch network, process the first texture feature and the first morphological feature in parallel, and generate a feature map that causes the weld edge weight to jump step. Based on the feature map, graph tile nodes are established at multiple scales, and a graph similarity matrix is obtained through a similarity aggregation method based on graph tiles; The graph similarity matrix is subjected to iterative confidence-weight joint evolution: the feature representation is updated with the current weight, and the graph similarity is recalculated with the updated features. After multiple iterations, the final attention weight is obtained. The critical regions of the pipeline are determined based on the final attention weights. A detection unit is used to fuse multi-scale features with global context within the critical area to detect the weld area of the pipeline.
Citation Information
Patent Citations
Submerged arc welding pipe weld defect judgment method based on deep learning
CN118967609A
Petroleum pipeline defect size measurement method based on HMCNet
CN120235831A