Weld defect detection method, electronic equipment and medium

By optimizing the weld defect detection features through spatial attention and dynamic attention mechanisms, and combining the Retinex-WeldNet and RE-YOLO modules, the difficulties of multi-scale and small target detection in low-light environments are solved, and efficient and accurate weld defect detection is achieved.

CN120125507BActive Publication Date: 2025-09-09HUNAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510139703.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-09-09
Estimated Expiration
2045-02-08

AI Technical Summary

Technical Problem

Existing weld defect detection technology has difficulty in effectively detecting multi-scale and small-target defects, especially microcracks and micropores, in low-light environments. Traditional methods are also susceptible to noise interference and uneven lighting, resulting in low detection accuracy.

Method used

The spatial attention mechanism is used to enhance key features, the cross-stage connection of the C3k2 module is used to optimize feature fusion, and the dynamic attention mechanism is combined to optimize the fusion of shallow and deep features. The Retinex-WeldNet module is used for illumination enhancement and detail restoration, the RE-YOLO module is introduced for multi-scale target detection, and a dynamic branching mechanism is constructed to adapt to different lighting conditions.

Benefits of technology

The accuracy and robustness of weld defect detection are improved, especially the detection capability of small targets in low-light environments, which significantly enhances the recognition and positioning capabilities of multi-scale targets and avoids unnecessary computational overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125507B_ABST
    Figure CN120125507B_ABST
Patent Text Reader

Abstract

The present invention provides a weld defect detection method, electronic device, and medium. The method of the present invention comprises: acquiring a weld image; performing enhancement processing on a weld image having an average brightness value lower than a threshold to obtain an enhanced image; using the weld image having an average brightness value higher than the threshold and the enhanced image as inputs to a deep learning network model, training the deep learning network model, and obtaining a weld defect detection model. The present invention enhances key features through a spatial attention mechanism, focusing on important spatial regions while suppressing irrelevant information. The processed features retain key information of each scale. Through the cross-stage partial connection of the C3k2 module, features are effectively optimized and refined, and the fusion and representation capabilities of features are improved, ensuring that information can be transmitted and fused between different levels and scales, thereby achieving accurate recognition and positioning of targets of various scales; the multi-level detection system can cover targets of multiple scales and improve the detection capability of small target defects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing and target detection, and in particular relates to a weld defect detection method, electronic equipment and medium. Background Art

[0002] Weld quality is a crucial quality control step in industrial production, and weld defect detection is of great significance in fields such as aviation, automotive, shipbuilding, and engineering machinery. However, weld defects are complex in nature, including welding slag, microcracks, and pores. Especially in low-light environments, weld surface details are difficult to clearly visualize, making defect detection extremely challenging. In recent years, with the advancement of industrial image detection technology, machine vision-based weld defect detection has gradually become a research hotspot. Efficient and automated weld defect detection can significantly improve industrial production efficiency and product quality.

[0003] At present, the technical solutions for weld defect detection mainly include the following categories:

[0004] (1) Traditional image enhancement methods

[0005] In low-light environments, image enhancement technology is a common means of improving image brightness and contrast. Common methods include: Histogram Equalization (HE): Improves overall brightness and contrast by adjusting the pixel value distribution of the image. CLAHE: Based on traditional histogram equalization, it adopts a local enhancement strategy to preserve local details to a certain extent. Gamma Correction: Enhances image brightness using nonlinear transformations. Retinex Theory: Improves the visual effects of low-light images based on a model that separates illumination and reflection.

[0006] Existing flaws: Traditional enhancement methods are prone to noise amplification, especially in low-light images. High ISO settings can cause significant photosensitive noise, and the noise and artifacts after enhancement can obscure the detailed features of the weld. Methods such as HE and CLAHE are prone to over-enhancement when dealing with complex weld boundaries, resulting in loss of boundary details. While the Retinex theoretical method has certain advantages in lighting optimization, existing implementations have errors in the lighting estimation process, resulting in unstable enhancement effects, especially in scenes with uneven low-light distribution.

[0007] (2) Traditional target detection methods

[0008] Traditional target detection methods detect weld defects through manual feature extraction and classifiers (such as support vector machines and random forests). These methods typically use edge detection, texture analysis, grayscale histograms, and other methods to extract features and combine them with machine learning models to classify defects.

[0009] Existing flaws: Limited feature expression capability: Manually designed features struggle to adapt to the complexity of weld defects, especially when the weld surface texture is complex and noise interference is high, which can easily lead to insufficient feature expression. Strong dependence on lighting: In low-light environments, traditional feature extraction methods significantly reduce their ability to describe boundaries and textures. Inadequate small target detection capability: Small targets such as microcracks and micropores in welds are difficult to detect, and are prone to missed detections and false detections.

[0010] (3) Deep learning target detection method

[0011] Deep learning-based object detection methods are increasingly being applied to weld defect detection, including classic detection algorithms such as Faster R-CNN, SSD, and YOLO. The YOLO family of algorithms has garnered significant attention due to its end-to-end design, high efficiency, and superior detection performance. YOLOv3 and subsequent versions have enhanced their multi-scale object detection capabilities by introducing modules such as the Feature Pyramid Network (FPN).

[0012] Existing defects: Poor adaptability to low-light images: The YOLO series of algorithms directly processes the input image and cannot solve the problems of insufficient brightness and uneven light distribution in low-light images. Boundary blur and noise interference: Weld defect detection requires precise boundary positioning, and existing algorithms have limited performance when dealing with blurred boundaries or complex backgrounds. Insufficient multi-scale and small target detection capabilities: Although FPN can fuse shallow and deep features, shallow features are easily affected by noise, while deep features have strong semantic expression but lack details, resulting in limited small target detection capabilities. Traditional multi-scale feature fusion methods cannot effectively coordinate the differences between shallow and deep features, resulting in low accuracy in multi-scale and small target detection, especially under complex background interference, which is prone to missed detection and false detection. Summary of the Invention

[0013] The purpose of the present invention is to address the deficiencies of the existing technology and provide a weld defect detection method, electronic equipment and medium to improve the detection capability of multi-scale and small target defects.

[0014] In order to achieve the above object, the technical solution adopted by the present invention is:

[0015] A method for detecting weld defects comprises the following steps:

[0016] S1, obtaining weld images;

[0017] S2. performing image enhancement processing on the weld images whose average brightness values ​​are lower than the threshold value to obtain an enhanced image;

[0018] S3. Using the weld image and the enhanced image whose average brightness value is higher than the threshold as inputs of a deep learning network model, training the deep learning network model, and obtaining a weld defect detection model;

[0019] The deep learning network model includes a backbone network, a neck network and a head network. The backbone network includes feature layers P2, P3, P4 and P5; the neck network includes a first spatial attention layer, a first C3k2 module, a second spatial attention layer, a second C3k2 module, a third spatial attention layer, a third C3k2 module, a fourth spatial attention layer and a fourth C3k2 module connected in sequence; the head network includes a first detection head, a second detection head, a third detection head, a fourth detection head and a Detect module;

[0020] The feature layer P2 is connected to the first spatial attention layer, the feature layer P3 is connected to the second spatial attention layer, the feature layer P4 is connected to the third spatial attention layer, and the feature layer P5 is connected to the fourth spatial attention layer;

[0021] The first C3k2 module is connected to the first detection head, the second C3k2 module is connected to the second detection head, the third C3k2 module is connected to the third detection head, and the fourth C3k2 module is connected to the fourth detection head;

[0022] The first detection head, the second detection head, the third detection head, and the fourth detection head are all connected to the Detect module.

[0023] The present invention enhances key features through a spatial attention mechanism, focusing on important spatial regions while suppressing irrelevant information. The processed features retain key information at each scale. Through the cross-stage partial connection of the C3k2 module, the features are effectively optimized and refined, and the fusion and representation capabilities of the features are improved, ensuring that information can be transmitted and integrated between different levels and scales, and achieving accurate recognition and positioning of targets of various scales. Based on the dynamic attention mechanism, the shallow and deep features are optimized and integrated, effectively improving the ability to express boundary details, and realizing the dynamic fusion of shallow detail features and deep semantic features, effectively solving the problems of shallow features being easily affected by noise and deep features losing details in traditional target detection methods. The multi-level detection system can cover extremely small, small, medium, and large targets, improving the detection capability of small target defects.

[0024] The present invention ensures the efficiency and robustness of the system under different lighting conditions by judging the illumination level of the image and selecting an illumination enhancement path or a direct target detection path; this mechanism improves the adaptability of the system and avoids unnecessary computational overhead.

[0025] Furthermore, the expression of the average brightness value is as follows:

[0026]

[0027] in, is the average brightness value of the weld image, M and N are the number of rows and columns of the weld image respectively, and I(i, j) is the brightness value of pixel (i, j) in the weld image.

[0028] Preferably, the threshold is 80.

[0029] Furthermore, the implementation process of S2 includes:

[0030] An image enhancement processing model is constructed, and weld images with average brightness values ​​lower than a threshold are input into the image enhancement processing model to obtain an enhanced image;

[0031] The image enhancement processing model includes an illumination estimator and a visual restoration unit connected in sequence;

[0032] The illumination estimator includes a first convolutional layer, a second convolutional layer, a third convolutional layer, and a first fusion layer connected in sequence;

[0033] The visual restoration unit includes a fourth convolutional layer, a multi-layer cascaded first neural network, a multi-layer cascaded second neural network, a seventh convolutional layer, and a second fusion layer connected in sequence; the first fusion layer is connected to the fourth convolutional layer and the second fusion layer respectively; the first neural network includes a DFM module and a fifth convolutional layer connected in sequence; the second neural network includes a deconvolution layer, a sixth convolutional layer, and a DFM module connected in sequence;

[0034] The DFM module includes a cross attention layer, a fifth spatial attention layer, a channel attention layer, a normalization layer, and a feedforward neural network connected in sequence.

[0035] This invention uses an illumination estimator to model the illumination distribution in low-light scenes, dynamically adjusting brightness to eliminate uneven illumination. A visual restoration unit restores detail and suppresses noise in the enhanced image, ensuring uniform brightness and clear details in the generated image, providing high-quality input for subsequent object detection. Furthermore, the visual restoration unit integrates multiple attention mechanisms to enhance the visibility and texture representation of weld defects in low-light environments.

[0036] Furthermore, in the cross-attention layer, the output features of the second convolutional layer are used as the query vector, and the input features of the cross-attention layer are used as the value vector and key vector.

[0037] Using illumination features as query vectors, we focus on local enhancement of low-frequency features in the image without redundant computation of the entire image. By fusing these two features through a cross-attention mechanism, we achieve precise information guidance and optimized attention distribution.

[0038] Furthermore, after the cross-attention layer, global and local illumination features are extracted through multi-directional scanning;

[0039] The expression for multi-directional scanning is as follows:

[0040]

[0041] Among them, h t is the hidden state of the current direction, x t is the output feature of the cross attention layer, y t is the output feature of the current direction, and is the learning parameter in the state update process, and C and D are the linear mapping parameters.

[0042] Multi-directional scanning can fully capture long-range dependencies and contextual information, effectively obtain the global information of the image, and integrate local features, thereby strengthening the expression of regional texture and structural information and enhancing the texture details of the weld area.

[0043] Furthermore, the first spatial attention layer, the second spatial attention layer, the third spatial attention layer, and the fourth spatial attention layer each include an eighth convolutional layer, a first batch normalization layer, a first ReLU activation function layer, a first sigmoid activation function layer, a ninth convolutional layer, a second batch normalization layer, a second ReLU activation function layer, a second sigmoid activation function layer, a third fusion layer, a tenth convolutional layer, a third batch normalization layer, and a third ReLU activation function layer;

[0044] The eighth convolutional layer, the first batch of normalization layers, the first ReLU activation function layer, and the first sigmoid activation function layer are connected in sequence;

[0045] The ninth convolutional layer, the second batch normalization layer, the second ReLU activation function, and the second sigmoid activation function layer are connected in sequence;

[0046] The second fusion layer, the tenth convolutional layer, the third batch normalization layer, and the third ReLU activation function layer are connected in sequence;

[0047] The first sigmoid activation function layer and the second sigmoid activation function layer are both connected to the third fusion layer.

[0048] Based on the same inventive concept, the present invention further provides an electronic device, comprising:

[0049] one or more processors;

[0050] A memory having one or more programs stored thereon, which, when executed by the one or more processors, enables the one or more processors to implement the steps of the weld defect detection method.

[0051] Based on the same inventive concept, the present invention also provides a computer-readable storage medium storing a computer program, which implements the steps of the weld defect detection method when executed by a processor.

[0052] Compared with the prior art, the present invention has the following beneficial effects:

[0053] The present invention enhances key features through a spatial attention mechanism, focusing on important spatial regions while suppressing irrelevant information. The processed features retain key information at each scale. Through the cross-stage partial connection of the C3k2 module, the features are effectively optimized and refined, and the fusion and representation capabilities of the features are improved, ensuring that information can be transmitted and integrated between different levels and scales, and achieving accurate recognition and positioning of targets of various scales. Based on the dynamic attention mechanism, the shallow and deep features are optimized and integrated, effectively improving the ability to express boundary details, and realizing the dynamic fusion of shallow detail features and deep semantic features, effectively solving the problems of shallow features being easily affected by noise and deep features losing details in traditional target detection methods. The multi-level detection system can cover extremely small, small, medium, and large targets, improving the detection capability of small target defects.

[0054] By determining the illumination level of the image and selecting either an illumination enhancement path or a direct target detection path, the present invention ensures the system's efficiency and robustness under varying illumination conditions. This mechanism significantly improves the system's adaptability and avoids unnecessary computational overhead. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 Schematic diagram of the weld defect detection method of the present invention;

[0056] Figure 2 Schematic diagram of the structure of the Retinex-WeldNet module of the present invention;

[0057] Figure 3 Schematic diagram of the structure of the RE-YOLO module of the present invention;

[0058] Figure 4 Schematic diagram of the structure of the recalibrated feature pyramid network of the present invention;

[0059] Figure 5 A visual comparison of different low-light enhancement methods in weld and steel defect scenes;

[0060] Figure 6 The comparison chart of the effects of low-light enhancement methods and HSV color distribution under different lighting conditions;

[0061] Figure 7 This is a comparison chart of the detection performance of different methods;

[0062] Figure 8 PR curve of RE-YOLO on NL-WELD dataset;

[0063] Figure 9 It is the PR curve of YOLOv11 on the NL-WELD dataset. DETAILED DESCRIPTION

[0064] The present invention will be described in detail below with reference to the following embodiments. It should be noted that the embodiments and features of the embodiments may be combined unless they conflict. For ease of description, the words "upper," "lower," "left," and "right" appearing below merely indicate the directions of upper, lower, left, and right relative to the accompanying drawings and do not limit the structure.

[0065] Example 1

[0066] like Figure 1 The weld defect detection method of this embodiment includes the following steps:

[0067] 1. Image acquisition: Use an industrial camera to acquire weld seam images and evaluate their overall brightness level using the following formula:

[0068]

[0069] in, represents the average brightness value of the weld image, M and N are the number of rows and columns of the weld image respectively, and I(i, j) represents the brightness value of pixel (i, j) in the weld image.

[0070] Determine whether the lighting is sufficient. If the average brightness value is greater than 80, it is normal lighting. If the average brightness value is less than 80, it is low lighting (insufficient lighting).

[0071] 2. Dynamic branching mechanism: When the image is not well-lit, it enters the Retinex-WeldNet module for illumination enhancement and detail restoration. When the image is well-lit, it enters the RE-YOLO module directly for defect detection.

[0072] 3. Illumination enhancement: Retinex-WeldNet separates the illumination component and the reflection component of the input image based on the Retinex theory, and then suppresses noise and restores details through the visual restoration unit.

[0073] 4. Multi-scale object detection: The enhanced image or the directly input image enters the RE-YOLO module, and features are optimized through the feature recalibration network (RC-FPN) and the selective boundary aggregation module (SBA) to perform multi-scale detection of weld defects.

[0074] 5. Defect marking and output: The final output of the inspection results includes the location, type and size of the weld defects.

[0075] (1) Retinex-WeldNet module

[0076] The Retinex-WeldNet module is a key module for image enhancement in low-light environments. It aims to solve problems such as insufficient brightness, uneven lighting, and blurred details in low-light scenes. Based on the Retinex theory, this module decomposes the input image into illumination components and reflection components. The illumination component describes the brightness distribution of the image, and the reflection component describes the texture and detail information of the image. By modeling and optimizing the illumination component and combining it with the detail restoration of the reflection component, an enhanced image with uniform brightness and clear details is generated. The Retinex-WeldNet module consists of two parts: the illumination estimator (IE) and the visual restoration unit (VRU). Figure 2 shown.

[0077] The illumination estimator is the first stage of Retinex-WeldNet, primarily responsible for modeling the illumination distribution in low-light environments and generating dynamically optimized illumination features. Through convolution operations, the illumination estimator combines the illumination information of the input image with the illumination prior features to generate an enhanced illumination component. Specifically, the illumination estimator first performs preliminary illumination feature extraction on the input image, then uses depthwise separable convolution operations to generate illumination distribution features. Finally, this is combined with the input image to produce a preliminary enhanced image.

[0078] According to the Retinex theory, the illumination estimator inputs the image in a low-light environment. Can be decomposed into reflection components and lighting components Its mathematical model is expressed as:

[0079] I=R⊙L

[0080] Here, ⊙ represents element-by-element multiplication (Hadamard product). In this model, R represents the reflective properties of the object surface, primarily describing the texture and detail of the image; L represents the illumination distribution in the scene, which is closely related to light intensity and uniformity.

[0081] However, the model assumes that the input image I is noise-free and distortion-free, which is inconsistent with underexposed scenes in low-light environments in actual engineering. Specifically, in low-light scenes, high ISO settings and long exposure times can significantly amplify the photosensitive noise in the image signal, thereby masking the defective texture on the weld surface. This degradation is particularly pronounced at the weld edge or other detail areas. At the same time, traditional image enhancement methods typically further amplify noise and artifacts when processing underexposed areas. This can not only lead to the loss of texture information in underexposed areas, but also cause color distortion in overexposed areas, further reducing the overall visibility of the weld defect area.

[0082] In order to solve the above problems, a perturbation modeling method is adopted. Based on the original Retinex model, combined with the noise and artifact problems in actual low-light scenes, the perturbation terms of the illumination component and the reflection component are introduced. and The improved mathematical model is expressed as:

[0083]

[0084] in, Describes the degradation of texture information caused by noise, which is particularly evident in the defect edge area. This reflects the uneven illumination and color deviation that may be caused by illumination estimation errors. After simplification, the illumination image can be represented as:

[0085]

[0086] Among them, I lu represents the image after illumination enhancement, and C is the combined effect of noise, illumination error, and texture residue. Therefore, the expression of Retinex-WeldNet is:

[0087] (I lu ,F lu )=IE(I,L p )

[0088] I en =VRE(I lu ,F lu )

[0089] RM=IE(I lu ,L p )+VRE(I lu ,F lu )

[0090] Among them, IE represents the illumination estimator and VRE represents the visual restoration module.

[0091] The main function of the illumination estimator is to extract the illumination prior L from the input image I. p , and combined with the lighting characteristics F in low-light scenes lu Generates the lighting estimation result of the image; the visual restoration module repairs the noise and artifact problems amplified during the low-light enhancement process by jointly modeling the lighting information and image details.

[0092] like Figure 2 As shown, first, IE uses a conv1×1 convolution layer (convolution kernel size is 1) to fuse I and L p Since the well-exposed area can provide semantic context information for the underexposed area, the input is upsampled using depthwise separable convolution conv5×5 to model the regional interaction under different lighting conditions, thereby generating the lighting feature F lu Finally, conv1×1 is used for downsampling to recover the three-channel illumination map L, which is then element-wise multiplied by the low-light image I to obtain the enhanced image I lu .

[0093] The visual restoration unit is the second stage of Retinex-WeldNet, responsible for enhancing image details and suppressing noise in low-light environments. This unit adopts a dynamic fusion model (DFM) and combines multiple attention mechanisms and feature fusion techniques, including an illumination-guided attention module (IGA), a two-dimensional selective scanning module (2D-SSM), and a convolutional block attention module (CBAM). Among them, IGA focuses on optimizing details in low-light areas by strengthening the correlation between illumination features and texture; 2D-SSM captures long-range dependencies through a multi-directional scanning mechanism to enhance texture details in the weld area; CBAM combines channel and spatial attention to further enhance image feature expression capabilities. After being processed by the visual restoration unit, the image has a uniform brightness distribution and effectively suppressed noise, providing high-quality input for the subsequent target detection module.

[0094] like Figure 2 As shown in Figure 2, the visual restoration unit achieves deep fusion of illumination features and spatial features through the dynamic fusion model (DFM) to improve detail restoration, noise suppression and texture restoration capabilities. lu First, it is downsampled by 3×3 convolution (stride=2) and combined with the illumination feature F luDimensions are aligned. Subsequently, the features continue to undergo two downsampling operations, each including a DFM module and a 4×4 convolution (stride=2), gradually reducing the width and height of the feature map and increasing the number of channels, thereby extracting deep semantic information and global illumination features. Next, the features are gradually restored to the original resolution of the image through a multi-layer step-by-step upsampling structure. Each stage includes a 2×2 deconvolution (stride=2), a 1×1 convolution and a DFM module to enhance detail representation and enhance the expression of texture information. At the same time, at each stage of feature processing, information of different scales is fused through a feature jump connection mechanism to compensate for the detail features that may be lost during the downsampling process. Finally, the output features are adjusted to a three-channel RGB format through a 3×3 convolution and compared with the brightness enhanced image I lu Perform residual fusion to generate an enhanced image with uniform brightness and rich details I en .

[0095] Dynamic Fusion Model (DFM) is the core module of the visual restoration unit, such as Figure 2 As shown, the input image first passes through the illumination guided attention module (IGA), using the illumination feature F lu The module models the illumination distribution in different regions, focusing on enhancing low-light areas and suppressing noise. Subsequently, it combines a two-dimensional selective spatial model (2D-SSM) to extract global and local illumination features. It also introduces spatial and channel attention via a convolutional block attention module (CBAM) to enhance texture details in weld defect areas. Finally, the module completes feature reconstruction through linear normalization (LN) and a feed-forward network (FFN), outputting an enhanced feature map that ensures uniform illumination distribution and clear texture details.

[0096] like Figure 2 As shown, IGA uses the lighting features F generated by the lighting enhancement module lu , together with the enhanced image extracted by feature extraction, is used as input. A design idea is proposed in the prior art, that is, lu The key vector K and the value vector V in the IGA input feature x are processed into tokens and then transformed to complete the fusion. However, in such methods, the key vector K and the value vector V are derived from different input features, which deviates from the input feature-centric attention assumption of the Transformer. To this end, this module adopts a more targeted structure to avoid data bias issues while significantly reducing computational overhead and parameter size.

[0097] To address this issue, this embodiment proposes to introduce Cross Attention layer into the IGA module to lu and x are used as the core to build a unified modeling framework, which effectively meets the requirements of the multi-head attention mechanism. The specific process can be formally expressed as follows:

[0098] X=[X1,X2,...,X k ],F lu =[F lu1 ,F lu2 ,...,F luk ]

[0099] Among them, X i represents the features of the i-th attention head, is the feature dimension of each attention head, and C is the total dimension of the input feature. Through this design, the input feature can be decomposed into k subspaces. The cross attention mechanism converts F lu As a query vector (Query), it focuses on the local enhancement of low-frequency features in the image without redundant calculation of the entire image. In the specific implementation, F lu It is defined as query Q, and the input feature x is defined as key K and value V respectively. The two features are fused through the cross-attention mechanism to achieve accurate information guidance and attention distribution optimization.

[0100] 2D-SSM uses multi-directional scanning and state modeling to effectively capture global image information and integrate local features, making it suitable for feature extraction and enhancement in complex areas under low-light conditions. Traditional unidirectional scanning modes struggle to fully capture the multi-dimensional information of welds and defective areas. However, 2D-SSM uses a four-directional scanning mechanism (top left to bottom right, bottom right to top left, top right to bottom left, and bottom left to top right) to effectively model long-range dependencies in the image, ensuring complete extraction of texture and detail information in complex areas.

[0101] 2D-SSM captures the temporal relationship between features by modeling the discrete state space in each scanning direction. Its calculation formula is as follows:

[0102]

[0103] Among them, h t is the hidden state of the current direction, x t is the input feature, y t is the output feature of the current direction, and are learned parameters during the state update process, while C and D are linear mapping parameters for features. This modeling approach allows the network to fully capture long-range dependencies and contextual information. Ultimately, the scan results from the four directions are fused to generate a unified two-dimensional feature map, further optimizing the illumination distribution and enhancing the representation of regional texture and structure.

[0104] (2) RE-YOLO module

[0105] like Figure 3As shown in Figure 1, the RE-YOLO module is a multi-scale object detection module optimized based on YOLOv11, used for accurate weld defect detection. Compared with traditional object detection methods, RE-YOLO designs a recalibrated feature pyramid network (RC-FPN) to target multi-scale objects (especially very small objects) in weld detection tasks. The main improvements include replacing the upsampling and splicing operations in the original YOLOv11 with a selective boundary aggregation module (SBA) and adding a new very small object detection head (P2 / 4). Compared with traditional detection models, the RE-YOLO module significantly improves the detection of small defects.

[0106] The Recalibrated Feature Pyramid Network (RC-FPN) is the core of the RE-YOLO module, used to dynamically fuse shallow and deep features. Shallow features contain rich detail but are susceptible to noise; deep features have strong semantic expression capabilities but suffer from severe loss of detail. RC-FPN uses a dynamic attention mechanism to generate fusion weights for shallow and deep features, enabling adaptive fusion of the two. During the fusion process, RC-FPN first uniformly maps the number of channels of shallow and deep features. It then dynamically optimizes the features through the Selective Boundary Aggregation (SBA) module to generate high-quality feature maps that combine detail and semantic consistency.

[0107] The RC-FPN part is mainly responsible for processing feature inputs of four different scales from the backbone network, namely P2, P3, P4 and P5, which are used to detect targets of different sizes. Among them, P2 corresponds to the smallest target, P3 corresponds to medium-sized targets, P4 corresponds to larger targets, and P5 processes the largest targets. For each feature map, the SBA module is first applied to enhance key features through the spatial attention mechanism, focusing on important spatial areas while suppressing irrelevant information. Subsequently, these attention-enhanced features are passed to the C3k2 module for further processing. The C3k2 module effectively optimizes and refines features through the cross-stage partial connection structure, thereby improving the feature fusion and representation capabilities.

[0108] Structurally, it consists of multiple alternating SBA and C3k2 modules, each pair corresponding to a feature map at a specific scale. These modules are connected in series to ensure smooth information transfer and fusion across different levels and scales. Furthermore, the head contains a Detect module, which receives features processed by the SBA and C3k2 modules at various scales (P2, P3, P4, and P5) and performs multi-scale object detection. By fusing features from different scales, the Detect module accurately identifies and locates objects of various sizes.

[0109] In terms of connectivity, each SBA module first performs spatial attention enhancement on the feature maps of the corresponding scale, then passes the enhanced features to the corresponding C3k2 module for further processing. These processed features not only retain key information at each scale but also achieve feature fusion and optimization through cross-stage connections in the C3k2 module. Finally, all optimized multi-scale features are integrated through the Detect module to form a unified detection output. This modular and hierarchical structural design enables efficient processing and fusion of features from different scales, improving the performance of the YOLO11 model in multi-scale object detection tasks.

[0110] Overall, the RC-FPN component leverages the multi-scale features of the backbone network through its structured modular composition and clear connectivity. The SBA module enhances the spatial representation of features, the C3k2 module optimizes feature fusion and representation, and the Detect module ultimately implements multi-scale object detection.

[0111] The specific structure of RC-FPN is as follows Figure 4 As shown in the figure, its core module is the Selective Boundary Aggregation module (SBA). The SBA module dynamically perceives the boundary details of shallow features and the semantic representation of deep features, establishing a complementary mechanism between features, thereby significantly improving the boundary detail representation ability and semantic consistency, and is used to improve the boundary representation ability under complex backgrounds. Specifically, before the shallow and deep features are input into the SBA module, they are first linearly mapped to unify the channel dimensions, thereby reducing the impact of feature mismatch and generating intermediate features for subsequent fusion operations. Next, the fused weights dynamically adjust the importance of shallow and deep features to achieve dynamic optimization of the boundary detail representation.

[0112] Based on a dynamic attention mechanism, SBA optimizes the fusion of shallow and deep features by screening and calibrating feature maps. Shallow features provide boundary details of weld defects, while deep features provide high-level semantic information. SBA dynamically generates fusion weights, effectively enhancing the ability to express boundary details of defects in complex scenarios.

[0113] The structure of SBA (Selective Boundary Aggregation) consists of multiple modules, which is mainly used for selective aggregation of feature boundaries. It receives two input feature maps F b and F s, representing boundary features at different levels respectively. These two feature maps are first input into two RAU (Re-calibration Attention Unit) modules for feature recalibration and enhancement. The RAU structure includes the following key parts: First, the input features are respectively passed through a 1×1 convolution layer to reduce the channel dimension to reduce the amount of computation while extracting important local features. Next, these features are subjected to batch normalization (BN) and ReLU activation function to ensure feature scale consistency and introduce nonlinearity. Subsequently, the convolutional features generate attention weights in each branch (implemented by the sigmoid activation function σ), obtaining T1′ and T2′ respectively. The input features T1 and T2 are then multiplied by their corresponding attention weights to complete the weighted calibration. Finally, the features of the two branches are fused through the addition operation to generate the calibrated output.

[0114] Features processed by the RAU module are fused through a concatenation operation. The fused features are further processed through a 3×3 convolutional layer to extract local spatial features. They are then processed using batch normalization (BN) and ReLU activation functions to ensure feature scale uniformity and enhance nonlinear expression capabilities. Ultimately, the output feature Z represents aggregated and calibrated boundary features, offering enhanced semantic expression and detail capture. SBA leverages the dynamic weighting and feature calibration mechanisms of the RAU, along with the integration capabilities of the aggregation module, to achieve selective aggregation of boundary information at different levels, enhancing the model's ability to capture and express boundary features.

[0115] To improve the detection capabilities of extremely small objects such as microcracks and micropores in welds, the RE-YOLO module has added an extremely small object detection head (P2 / 4). Traditional object detection networks typically only cover small, medium, and large objects, neglecting the detection needs of extremely small objects. The P2 / 4 detection head specifically uses shallow, high-resolution features as input to perform refined boundary description and feature representation of extremely small objects, further enhancing the ability to describe the boundaries and details of extremely small objects. Working in collaboration with the traditional P3 / 8, P4 / 16, and P5 / 32 detection heads, the P2 / 4 head effectively complements the detection capabilities of extremely small objects, achieving comprehensive coverage of extremely small, small, medium, and large objects. Shallow features, primarily generated by the P2 / 4 and P3 / 8 layers, focus on object boundaries and details, while deep features, generated by the P4 / 16 and P5 / 32 layers, focus on semantic classification accuracy. RE-YOLO builds a comprehensive multi-scale detection system covering extremely small, small, medium, and large objects, effectively improving the performance of weld defect detection.

[0116] Finally, in the RE-YOLO module, non-maximum suppression (NMS) is used to filter the detection results, remove bounding boxes with large overlaps, and retain the results with the highest confidence. The specific process is: First, based on the confidence score and category probability of each bounding box, the final score is calculated and sorted in descending order. Then, the bounding boxes with the highest scores are selected in turn, and the intersection over union (IoU) is calculated with the remaining bounding boxes. If the IoU is higher than the set threshold (such as 0.5), it is considered that the box overlaps too much and is removed; if the IoU is lower than the threshold, the box is retained and the remaining boxes are processed. The final output is a set of non-overlapping bounding boxes with the highest confidence, thereby improving the accuracy and robustness of the detection. The output results include the location, type and confidence value of the weld defect, which can meet the high-precision requirements of industrial weld detection.

[0117] (3) Dynamic branching mechanism

[0118] The Retinex-WeldNet and RE-YOLO modules achieve seamless collaboration through a dynamic branching mechanism. When the image is poorly illuminated, the Retinex-WeldNet module performs illumination enhancement and detail restoration on the image, generating a high-quality image with uniform brightness and clear details, providing optimized input for the RE-YOLO module. When the image is well-illuminated, the image directly enters the RE-YOLO module for target detection, avoiding unnecessary computational overhead. This collaborative design can adapt to varying lighting conditions, improving the robustness and efficiency of weld defect detection, particularly in low-light scenarios and small target detection, significantly outperforming existing technologies.

[0119] The end-to-end weld defect detection framework that combines illumination enhancement and target detection in this embodiment adapts to different lighting conditions through a dynamic branching mechanism. The illumination enhancement module (Retinex-WeldNet) designed based on Retinex theory solves problems such as insufficient brightness and noise amplification in low-light environments through illumination estimation and detail restoration. The optimized target detection module (RE-YOLO) significantly improves the detection capability of small targets such as weld microcracks and micropores through deep mining of shallow high-resolution features through feature recalibration pyramid network (RC-FPN) and extremely small target detection head (P2 / 4), and optimizes the fusion method of shallow and deep features; it effectively solves the problem that shallow features are easily interfered by noise and deep features lose details in traditional target detection methods, and effectively improves the ability to express defect boundaries under complex backgrounds.

[0120] This embodiment proposes an efficient and robust weld defect detection framework through dynamic illumination enhancement and optimized target detection module, which significantly outperforms the existing technology in low-light environments and small target detection.

[0121] Example 2

[0122] 1. Dataset Description

[0123] To verify the effectiveness and versatility of this method in industrial scenarios, the experiments used the self-developed low-light dataset LL-WELD for weld defects as the core dataset. Two public datasets, NEU-DET and PCB-DET, were also introduced for steel surface defect detection and PCB defect detection, respectively. Details of the datasets are shown in Table 1. The defects in these datasets share the following common characteristics: significant background interference, complex lighting conditions, small sample sizes, and extreme aspect ratios. These characteristics pose significant challenges for real-time and accurate defect detection.

[0124] (1) LL-WELD is a self-built low-light dataset for weld defect detection, which contains a total of 687 low-light images. Among them, 472 images are obtained by darkening images collected under normal lighting conditions using gamma correction technology. The formula is as follows:

[0125]

[0126] Among them, V out represents the output image, V in represents the input image. The remaining 215 images were directly collected in real-world low-light environments. Through data augmentation operations (such as flipping and cropping), the dataset was expanded to 2,655 images. This dataset covers typical weld defects such as weld overhangs, undercuts, and holes. Based on experimental requirements, it is divided into 2,124 training images, 266 validation images, and 265 test images.

[0127] (2) The NEU-DET dataset is a public dataset for steel surface defect detection released by Northeastern University, containing 1,800 images of hot-rolled steel strip defects. The dataset covers six typical steel surface defects, including rolling scale (Rs), pitting (Ps), inclusions (In), plaques (Pa), cracks (Cr), and scratches (Sc). In order to simulate low-light environments, the experiment uses gamma correction technology to adjust the brightness of all images to generate the low-light version of the dataset LL-NEU-DET. In the experiment, the dataset is divided into 1,260 training images, 270 verification images, and 270 test images.

[0128] (3) The PCB-DET dataset is a public dataset for PCB defect detection released by Peking University. It contains 693 defect images, covering six typical defects: mouse bite (Mb), short circuit (Sh), burr (Sp), pseudo copper (Spc), leak hole (Mh), and open circuit (Oc). To verify the performance of this method in low-light scenes, the experiment adjusted the brightness of all images using gamma correction technology to generate the low-light version of the dataset LL-PCB-DET. Finally, the dataset was divided into 555 training images, 69 verification images, and 69 test images.

[0129] Table 1 Detailed information about the NEU-DET, PCB-DET, and LL-WELD datasets

[0130]

[0131] 2. Experimental Setup and Performance Evaluation

[0132] The experiments were conducted on a high-performance server equipped with an RTX 4090 graphics card. The operating environment included Linux (Ubuntu 20.04), Python 3.8, PyTorch 2.0.0, and CUDA 11.8 to ensure computational efficiency and compatibility. The experimental design consisted of two parts: a low-light enhancement task and an object detection task. The low-light enhancement task was performed using the Retinex-WeldNet module, using the LL-WELD and LL-NEU-DET datasets, with a uniform training resolution of 128×128. The object detection task was performed using the RE-YOLO module, with an input resolution of 640×640. The specific experimental settings and parameters are shown in Table 2. To ensure fairness and reliability of the experimental results, all compared models were trained and tested under identical hardware and software conditions.

[0133] Table 2 Experimental settings and parameters

[0134]

[0135]

[0136] In terms of performance evaluation, the low-light enhancement task primarily uses three metrics to quantify the quality of the enhanced image: PSNR (Peak Signal-to-Noise Ratio), SSIM (Structural Similarity Index), and RMSE (Root Mean Square Error). PSNR measures the signal-to-noise ratio, reflecting the error level between the enhanced image and the reference image; higher values ​​indicate lower error. SSIM quantifies the structural similarity of images based on brightness, contrast, and structural information. Its value ranges from 0 to 1, with closer values ​​to 1 indicating higher structural similarity between the enhanced image and the reference image. RMSE calculates the pixel-level error between the predicted image and the reference image; lower values ​​indicate closer image quality to the ideal result.

[0137] The object detection task uses mAP (mean average precision) as the core performance metric, combining precision and recall to comprehensively evaluate the model's detection performance. Precision is defined as the proportion of samples predicted as positive by the model that are actually positive; recall represents the proportion of all positive samples that are successfully detected. mAP is calculated by averaging the detection precision of all categories, using the following formula:

[0138]

[0139] Among them, TP is the number of samples that are actually positive and predicted to be positive; FP is the number of false positives, which is the number of samples that are actually negative but predicted to be positive; FN is the number of false negatives, which is the number of samples that are actually positive but predicted to be negative; n is the number of samples, AP is the number of samples. i is the average precision.

[0140] 3. Low-light enhancement experiments and results

[0141] The performance of various low-light enhancement methods on the LL-WELD and LL-NEU-DET datasets was compared and analyzed, focusing on evaluating their performance in detail recovery, illumination balance, and noise suppression. The experiment used PSNR (peak signal-to-noise ratio), SSIM (structural similarity index), and RMSE (root mean square error) as evaluation indicators. Among them, the LL-WELD dataset is mainly composed of weld defect images, which is mainly used to test the ability of enhancement methods to restore details and suppress noise; while the LL-NEU-DET dataset contains steel defect images, focusing on evaluating the performance of enhancement methods in illumination balance and texture detail recovery. The experimental results show that Retinex-WeldNet outperforms other comparison methods in all indicators, showing significant advantages.

[0142] Performance Comparison. As shown in Table 3, on the LL-WELD dataset, Retinex-WeldNet achieves PSNR and SSIM of 34.86 and 0.971, respectively, with an RMSE of 4.183, achieving the best performance among all methods. This demonstrates that Retinex-WeldNet can effectively maintain illumination balance while enhancing image brightness, significantly improving image quality and reducing potential errors during the enhancement process. On the LL-NEU-DET dataset, Retinex-WeldNet's PSNR and SSIM further improve to 35.71 and 0.98, respectively, while its RMSE drops to 3.848, validating its robustness and generalization capabilities in complex low-light scenes. In comparison, URetinex-Net performs second best, but significantly lags behind Retinex-WeldNet in all three metrics: PSNR, SSIM, and RMSE. Furthermore, Retinex-Net and SCI methods perform poorly in illumination compensation and detail recovery, with low PSNR and SSIM levels, exposing their limitations in modeling image details during the enhancement process.

[0143] Table 3 Performance comparison of different methods on datasets

[0144]

[0145] Comparison of visual effects. To further verify the actual performance of each method in low-light scenes, Figure 5 This paper compares the visual effects of enhancing an input low-light image using four methods: Retinex-Net, URetinex-Net, SCI, and Retinex-WeldNet. The test images selected for this experiment contain weld defects and steel defects, representing typical industrial low-light inspection scenarios. In the original low-light image, defects (such as holes and cracks) are barely visible, and background texture information is severely lost.

[0146] Enhancement results show that Retinex-Net has some effectiveness in improving brightness, initially revealing defective areas, but its ability to recover detail is limited. Weld defects have blurred textures, and background information is easily obscured by noise. URetinex-Net demonstrates some advantages in illumination balance and detail recovery, improving the clarity of defective areas. However, background details remain blurred and exhibit minor artifacts, such as unnatural edge transitions. The SCI method slightly outperforms the previous two methods in detail recovery, improving the integrity of defective areas. However, it suffers from over-enhancement, resulting in distortion of background textures and ineffective suppression of random noise. In contrast, Retinex-WeldNet performs the best among all methods. Its illumination estimation module accurately models illumination distribution in low-light environments, avoiding the over-enhancement and color distortion common in SCI methods. Furthermore, its dynamic fusion module significantly improves detail recovery and noise suppression. Enhancement results show that welds and steel defects (such as holes and cracks) are clearly restored, with natural and artifact-free edge transitions, accurate background texture information, and significantly reduced random noise. The comprehensive and balanced performance of Retinex-WeldNet fully demonstrates its superiority in low-light enhancement tasks.

[0147] Performance under different lighting conditions. Further experiments explored the enhancement performance of Retinex-WeldNet under different lighting conditions (as shown in Table 4). Under extremely low lighting conditions (average brightness of 20), the PSNR and SSIM of Retinex-WeldNet were 34.86 and 0.971 respectively, and the RMSE was 4.183, which were significantly better than Retinex-Net's 25.09 and 0.689. Under medium and low lighting conditions (average brightness of 50), the PSNR and SSIM of Retinex-WeldNet were improved to 37.154 and 0.988 respectively, and the RMSE was reduced to 3.72, further demonstrating its illumination adaptability and detail recovery capabilities. In addition, by analyzing the HSV color distribution of the enhanced image and the Ground Truth ( Figure 6 ), the enhanced image using Retinex-WeldNet is closer to the ground truth in terms of brightness, saturation, and hue variations, while the enhanced result using Retinex-Net is concentrated in low-brightness and low-saturation areas, showing a significant difference from the reference image. This further validates Retinex-WeldNet's ability to restore detail and achieve illumination balance under various lighting conditions.

[0148] Table 4 Performance comparison of methods under different lighting conditions

[0149]

[0150] In summary, Retinex-WeldNet demonstrates exceptional performance in low-light enhancement tasks. Through precise modeling of the illumination estimation module and innovative design of the dynamic fusion module, this method generates clear, natural, high-fidelity enhanced images in both extreme low-light and medium-low-light scenarios, significantly improving brightness balance, detail recovery, and noise suppression. These features not only provide high-quality input for subsequent defect detection tasks but also lay an important technical foundation for industrial low-light applications.

[0151] 4. Test experiments and results

[0152] To validate the performance of the LIDet-Net framework (Retinex-WeldNet module + RE-YOLO module) in industrial defect detection, experiments were conducted on three datasets: weld (LL-WELD), steel (LL-NEU-DET), and PCB (LL-PCB-DET). The experiments focused on evaluating LIDet-Net's advantages in low-light environments, complex backgrounds, and multi-scale object detection, and explored the contributions of the illumination enhancement module and network modular design to performance improvements.

[0153] 4.1 Comparison of detection performance under low light conditions

[0154] Experiments compared the low-light performance of LIDet-Net with several classic object detection methods, including single-stage detection models (such as the YOLO series) and two-stage detection models (such as Faster R-CNN). To ensure fairness, all models were randomly trained from scratch on the same hardware without using pre-trained weights. The experimental results are shown in Table 5.

[0155] Experimental results show that the performance of traditional object detection methods degrades significantly in low-light environments, particularly in key metrics such as mAP and Recall. This is because traditional methods rely heavily on the quality of the input image, and in low-light conditions, common noise, low contrast, and blurred boundaries affect detection accuracy. On the LL-WELD dataset, LIDet-Net achieved a mAP of 91.1%, an 8.3% improvement over YOLOv11. Precision and Recall reached 88.1% and 87.1%, respectively, significantly outperforming the competing models. This performance improvement is primarily attributed to the dynamic illumination enhancement feature of the Retinex-WeldNet module, which mitigates uneven illumination and detail degradation in low-light scenes, providing high-quality feature input for the detection module. Furthermore, the multi-scale feature fusion mechanism and small object detection head design in RE-YOLO effectively improve the model's performance in complex backgrounds and for small object detection. Comparative results on other datasets also confirm the superior performance of LIDet-Net. On the LL-NEU-DET and LL-PCB-DET datasets, LIDet-Net's mAP increased by 3.8% and 15.1% respectively compared to YOLOv11, and its recall rate increased by 1.4% and 17.1% respectively, demonstrating its excellent robustness and versatility.

[0156] In order to further compare the performance of different models intuitively, Figure 7 The detection results of each model on three datasets are shown. As can be seen in the figure, traditional methods often suffer from significant missed detections and false detections in low-light environments, with low confidence levels and scattered distribution of detection boxes. They perform particularly poorly in detecting typical small objects (such as pores and missing holes). In contrast, LIDet-Net significantly reduces missed detections and false detections, demonstrating its superior detection capabilities in low-light environments.

[0157] Table 5 Defect detection results

[0158]

[0159] To evaluate the adaptability of LIDet-Net under different lighting conditions, the experiment compared its detection performance with YOLOv11 in extremely low light and medium-low light environments (see Table 6). The results show that LIDet-Net exhibits significant advantages in both lighting conditions, especially in extremely low light scenes. In extremely low light conditions, due to the lack of an illumination enhancement module, YOLOv11 is unable to effectively deal with uneven illumination and noise interference, resulting in a significant decrease in detection performance (mAP is only 43.5%). LIDet-Net dynamically optimizes illumination distribution, compensates for details, and suppresses noise through the Retinex-WeldNet module, increasing mAP to 91.1%, and Precision and Recall to 87.7% and 86.8% respectively. In medium-low light environments, although YOLOv11's mAP is improved to 81.9%, it is still lower than LIDet-Net's 91.2%. This performance is due to the Retinex-WeldNet module's ability to dynamically adjust optimization strategies according to lighting conditions and generate high-quality feature expressions, ensuring the efficiency and stability of LIDet-Net in different lighting environments.

[0160] Table 6 Performance comparison of LL-WELD dataset under different lighting conditions

[0161]

[0162] To further verify the enhancement effect of Retinex-WeldNet, the experiment compared the performance of combining different illumination enhancement methods with RE-YOLO (see Table 7). The mAP of Retinex-WeldNet reached 91.1%, significantly better than Retinex (53.7%) and SCI (82%). The traditional Retinex method does not model the noise in low-light scenes and is prone to introducing artifacts and random noise; the SCI method has improved in noise suppression, but the detail recovery effect is limited, especially in low-contrast areas. It is worth noting that the mAP of the Uretinex method is only 42.4%, even lower than the unenhanced YOLOv11 (43.5%). This phenomenon is attributed to the problems of over-enhancement and insufficient feature expression in the low-light enhancement process. The enhanced image has blurred edges and lost texture information, which seriously affects the feature extraction ability of the detection module. In contrast, Retinex-WeldNet uses an illumination estimation module (IE) to dynamically model the illumination distribution in low-light scenes, and combines it with a visual restoration unit (VRU) to restore detail and suppress noise. The resulting enhanced image has more balanced illumination, more complete details, and lower noise. This high-quality input significantly improves detection performance, enabling LIDet-Net to achieve significantly higher mAP, precision, and recall than other methods.

[0163] Table 7 Comparison of lighting enhancement methods combined with RE-YOLO on the LL-WELD dataset

[0164] Methods Precision(%) Recall (%) mAP (%) Retntinex+RE-YOLO 66 49.2 53.7 SCI+RE-YOLO 82.1 74.3 82 Uretntinex+RE-YOLO 62.2 38 42.4 LIDet-Net 87.7 86.8 91.1

[0165] 4.2 Analysis of RE-YOLO’s detection performance under normal lighting conditions

[0166] Table 8 compares the detection performance of YOLOv11 and RE-YOLO on the three datasets: NL-WELD, NEU-DET, and PCB-DET under normal lighting conditions. NL-WELD is a normal lighting dataset generated by applying Retinex-WeldNet to the low-light LL-WELD dataset. The results show that RE-YOLO significantly outperforms its baseline model, YOLOv11, on all datasets, with particularly significant improvements for small objects.

[0167] In the NL-WELD dataset, RE-YOLO achieved improvements in detection precision and recall across all defect categories. For example, for porosity and slag inclusion, two typical small-target defects, RE-YOLO achieved mAP improvements of 88.8% (up 19.7%) and 87.4% (up 4.2%), respectively, compared to YOLOv11. This demonstrates that RE-YOLO effectively improves the detection accuracy of small defects.

[0168] In order to further verify the small target detection capability, Figure 8 、 Figure 9 The PR curves for RE-YOLO and YOLOv11 on the NL-WELD dataset are shown. The area enclosed by the PR curve reflects the model's overall mAP performance. As can be seen from the figure, RE-YOLO's PR curve significantly outperforms YOLOv11, especially for small object categories such as porosity and slag inclusion.

[0169] In the NEU-DET dataset, RE-YOLO's mAP improved by 7.3% for the crack defect category, and its recall rate increased by 12.1% for the rolled-in scale defect category, indicating that RE-YOLO can more effectively capture contextual information and enhance the model's defect detection capabilities.

[0170] In the PCB-DET dataset, PCB defects typically feature numerous small objects and complex backgrounds. RE-YOLO achieved improvements in detection performance across all six defect categories, particularly for open circuit defects, with a 6.3% increase in mean average precision (mAP). This demonstrates RE-YOLO's superior detection performance on PCB datasets, which are typically characterized by small object defects. Furthermore, recall rates were improved for all defect categories, effectively reducing the risk of missed detections.

[0171] Compared with YOLOv11, RE-YOLO has shown significant advantages in performance on multiple datasets, especially in the detection of small target defects. First, on the NL-WELD dataset, RE-YOLO's detection accuracy for small target defects such as holes and slag inclusions has been particularly improved. For example, in the hole defect category, RE-YOLO's mAP increased from 69.1% of YOLOv11 to 88.8%, an improvement of 19.7%. This shows that RE-YOLO has significantly improved the detection accuracy of small-sized defects, especially for those with blurred boundaries and small targets. In contrast, large targets such as weld nodules have relatively limited accuracy improvements due to their clearer boundaries. RE-YOLO effectively enhances the model's ability to capture the boundaries of small targets through the SBA module, thereby significantly improving the detection accuracy of small targets.

[0172] In the NEU-DET dataset, RE-YOLO also demonstrated its advantages in complex defect detection, especially for small objects such as cracks and entrained oxide scale. RE-YOLO's mAP for cracks improved by 7.3%, while the recall rate for entrained oxide scale defects increased by 12.1%. This demonstrates RE-YOLO's powerful ability to capture contextual information and small objects in complex backgrounds, effectively reducing missed detections and improving the model's overall detection performance.

[0173] On the PCB-DET dataset, RE-YOLO continued to demonstrate its outstanding performance, particularly in defect detection tasks involving small objects and complex backgrounds. Taking open circuit defects as an example, RE-YOLO's mAP improved by 5.3%, from 90.8% to 97.1%. This demonstrates that RE-YOLO is able to maintain high detection accuracy while significantly reducing missed detections when dealing with small objects and complex backgrounds. For other defect categories, while the improvement for some targets was smaller, overall performance also improved in terms of recall, particularly for defect types such as burrs and pseudo-copper, demonstrating RE-YOLO's versatility and strong adaptability across a wide range of defect types.

[0174] Overall, RE-YOLO demonstrates significant advantages over YOLOv11 in small object detection, particularly in accurately detecting small objects in complex backgrounds. By combining the SBA module with the ultra-small object detection head, RE-YOLO effectively improves the boundary accuracy of small objects, reduces missed detections, and enhances overall detection performance. Whether detecting solder defects, material surface defects, or PCB defects, RE-YOLO demonstrates significant potential in practical applications, particularly for small object defect detection.

[0175] Table 8 Detection performance comparison of YOLOv11 and RE-YOLO under normal lighting conditions

[0176]

[0177]

[0178] Example 3

[0179] This embodiment provides an electronic device, including:

[0180] one or more processors;

[0181] A memory stores one or more programs, which, when executed by one or more processors, enable the one or more processors to implement the steps of the weld defect detection method.

[0182] In some implementations, the memory may be a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk storage.

[0183] In other implementations, the processor may be a central processing unit (CPU), a digital signal processor (DSP), or other general-purpose processors, which are not limited herein.

[0184] This embodiment provides a computer-readable storage medium storing a computer program, which implements the steps of the weld defect detection method when executed by a processor.

[0185] The contents illustrated in the above embodiments should be understood as these embodiments are only used to more clearly illustrate the present invention, and are not used to limit the scope of the present invention. After reading the present invention, various equivalent modifications of the present invention by those skilled in the art shall fall within the scope defined by the claims attached to this application.

Claims

1. A weld defect detection method, characterized in that: The following steps are involved: S1, obtaining weld images; S2. performing image enhancement processing on the weld images whose average brightness values ​​are lower than the threshold value to obtain an enhanced image; S3. Using the weld image and the enhanced image whose average brightness value is higher than the threshold as inputs of a deep learning network model, training the deep learning network model, and obtaining a weld defect detection model; The deep learning network model includes a backbone network, a neck network and a head network, and the backbone network includes feature layers P2, P3, P4 and P5; The neck network includes a first spatial attention layer, a first C3k2 module, a second spatial attention layer, a second C3k2 module, a third spatial attention layer, a third C3k2 module, a fourth spatial attention layer, and a fourth C3k2 module connected in sequence; the head network includes a first detection head, a second detection head, a third detection head, a fourth detection head, and a Detect module; The feature layer P2 is connected to the first spatial attention layer, the feature layer P3 is connected to the second spatial attention layer, the feature layer P4 is connected to the third spatial attention layer, and the feature layer P5 is connected to the fourth spatial attention layer; The first C3k2 module is connected to the first detection head, the second C3k2 module is connected to the second detection head, the third C3k2 module is connected to the third detection head, and the fourth C3k2 module is connected to the fourth detection head; The first detection head, the second detection head, the third detection head, and the fourth detection head are all connected to the Detect module; The implementation process of S2 includes: An image enhancement processing model is constructed, and a weld image having an average brightness value lower than a threshold is input into the image enhancement processing model to obtain an enhanced image; The image enhancement processing model includes an illumination estimator and a visual restoration unit connected in sequence; The illumination estimator includes a first convolutional layer, a second convolutional layer, a third convolutional layer, and a first fusion layer connected in sequence; The visual restoration unit includes a fourth convolutional layer, a multi-layer cascaded first neural network, a multi-layer cascaded second neural network, a seventh convolutional layer, and a second fusion layer connected in sequence; the first fusion layer is connected to the fourth convolutional layer and the second fusion layer respectively; the first neural network includes a DFM module and a fifth convolutional layer connected in sequence; the second neural network includes a deconvolution layer, a sixth convolutional layer, and a DFM module connected in sequence; The DFM module includes a cross attention layer, a fifth spatial attention layer, a channel attention layer, a normalization layer, and a feedforward neural network connected in sequence.

2. The weld defect detection method according to claim 1, characterized in that: The expression of the average brightness value is as follows: ; in, is the average brightness value of the weld image, M and N are the number of rows and columns of the weld image respectively, and I(i, j) is the brightness value of pixel (i, j) in the weld image.

3. The weld defect detection method according to claim 1, characterized in that: In the cross-attention layer, the output features of the second convolutional layer are used as the query vector, and the input features of the cross-attention layer are used as the value vector and key vector.

4. The weld defect detection method according to claim 1, characterized in that: After the cross-attention layer, global and local illumination features are extracted through multi-directional scanning; The expression for multi-directional scanning is as follows: ; in, h t is the hidden state of the current direction, x t is the output feature of the cross attention layer, y t is the output feature of the current direction, and is the learning parameter in the state update process, and C and D are the linear mapping parameters.

5. The weld defect detection method according to claim 1, characterized in that: The first spatial attention layer, the second spatial attention layer, the third spatial attention layer, and the fourth spatial attention layer each include an eighth convolutional layer, a first batch normalization layer, a first ReLU activation function layer, a first sigmoid activation function layer, a ninth convolutional layer, a second batch normalization layer, a second ReLU activation function layer, a second sigmoid activation function layer, a third fusion layer, a tenth convolutional layer, a third batch normalization layer, and a third ReLU activation function layer; The eighth convolutional layer, the first batch of normalization layers, the first ReLU activation function layer, and the first sigmoid activation function layer are connected in sequence; The ninth convolutional layer, the second batch normalization layer, the second ReLU activation function, and the second sigmoid activation function layer are connected in sequence; The third fusion layer, the tenth convolutional layer, the third batch normalization layer, and the third ReLU activation function layer are connected in sequence; The first sigmoid activation function layer and the second sigmoid activation function layer are both connected to the third fusion layer.

6. An electronic device, characterized in that: include: one or more processors; A memory having one or more programs stored thereon, which, when executed by the one or more processors, enables the one or more processors to implement the steps of the method according to any one of claims 1 to 5.

7. A computer-readable storage medium, characterized in that The computer program is stored therein, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Dam body defect identification method based on attention feature fusion enhancement network

    CN119027795A

  • Intelligent high-altitude camera cleaning method based on unmanned aerial vehicle

    CN119216290A