Multi-scale glue point detection method
By using the Multi-level Attention-Driven Fusion Network (MACAN) network model in laser glue point detection, the problems of dual-object detection complexity, multi-scale defect detection difficulty and complex background interference in glue point detection are solved, and real-time, robust and high accuracy of glue point detection is achieved.
Patent Information
- Application Number
- CN202510216817.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-26
AI Technical Summary
In the collimated lens glue point detection of high-power lasers, there is the complexity of dual-object detection, the difficulty of detection of multi-scale defects, and the interference of complex backgrounds, resulting in the limitation of the accuracy and robustness of glue point detection.
The Multi-level Attention-Driven Fusion Network (MACAN) network model is used to collect images through the intelligent laser lens dispensing platform, establish an LLAS data set, and use a multi-level attention-guided context aggregate network model to detect the glue point target. This model combines U-Net backbone network, Kolmogorov-Arnold representation module, multi-perceptual spatial attention module, KAN-enhanced channel attention module and multi-dimensional feature fusion module to enhance the multi-scale and robustness of detection.
Real-time and robust detection of glue points for laser collimated lenses is achieved, the accuracy of glue points detection and the automation level of the system are improved, and quality assurance is provided for the high coupling efficiency of high power lasers.
Smart Images

Figure CN120125547A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of glue dot detection, and relates to a multi-scale glue dot detection method. Background Art
[0002] High-power lasers have extensive application requirements in many fields such as optical communication, industrial processing, medical devices, and national defense. These fields have extremely high requirements for the efficiency, stability, and service life of lasers. As one of the core optical components, the collimating lens, its packaging process directly affects the beam quality, transmission efficiency, and overall system stability. However, the coupling and packaging technology of collimating lenses still faces many challenges. Especially during the dispensing process, factors such as the position, shape, and size of the glue dots significantly affect the coupling accuracy of the collimating lens, resulting in glue dot defects becoming one of the main factors affecting the performance of the optical system. Therefore, it is very necessary to detect the quality of the glue dots of the collimating lens during the coupling and packaging process of high-power lasers.
[0003] In recent years, thanks to the rapid development of convolutional neural networks, defect detection technologies using deep learning networks have shown remarkable results in many industrial defect recognition tasks. However, for the glue dot detection of laser collimating lenses, there are still three main challenges:
[0004] (1) Dual-object detection introduces more complexity to the glue dot detection task. The collimating lens is bonded and fixed through two glue dots. The coupling accuracy of the collimating lens is not only affected by the position and shape of a single glue dot, but more importantly, by the precise relative position relationship between the two glue dots. In dual-object detection, there is a spatial position association between the two objects. This association requires the model to effectively understand and consider the relationship between the two objects. Especially when there is a large difference in scale between the two objects, the design of the receptive field of the network and the detection framework becomes more complex.
[0005] (2) Multi-scale defects such as random shapes pose challenges to the precise detection of glue dots. Ideally, the shapes of the two glue dots are regular circles, and the objects have an appropriate scale in the entire field of view, and there is a high contrast between the objects and the background. However, in actual dispensing operations, the glue dots usually have potential problems such as different degrees of position deviation, size deviation, and random shapes. The glue dots have irregular shapes and relatively small overall sizes, which may pose challenges to the detection accuracy.
[0006] (3) Complex backgrounds may affect the accurate detection of glue dots. The features of the glue dots are not only difficult to distinguish from the background color of their target areas, but the large-area overflow of chip solder below further increases the difficulty of accurately identifying the glue dots from the background. This situation not only poses a challenge to traditional image processing methods based on grayscale information, but even for deep learning methods, it requires the network to have high robustness to ensure the accurate identification of the target. Summary of the Invention
[0007] Based on the deficiencies in the prior art, the present application provides a method for detecting multi-scale glue dots in the bonding of laser collimating lenses based on the Multi-level Attention-Driven Fusion Network (MADF-Net) network model, aiming to achieve more real-time and robust detection of the glue dots of laser collimating lenses and provide necessary quality assurance for achieving high coupling efficiency of high-power laser collimating lenses.
[0008] The present invention provides a multi-scale glue dot detection method, including the following steps:
[0009] Step 1: Use an intelligent laser lens dispensing platform to collect images of the laser chip before bonding; and establish an LLAS dataset based on the collected images;
[0010] Establish a multi-level attention-guided context aggregation network model;
[0011] Step 2: Apply the multi-level attention-guided context aggregation network model to process the images in the LLAS dataset and segment the glue dot targets;
[0012] Step 3: Based on the segmented glue dot targets, obtain the glue dot detection results.
[0013] The multi-level attention-guided context aggregation network model uses a U-Net model as the backbone network and fuses an interaction self-attention-guided Kolmogorov-Arnold representation module, a multi-perception spatial attention module, a KAN-enhanced channel attention module, and a multi-dimensional feature fusion module through an attention mechanism based on the Kolmogorov-Arnold representation.
[0014] The multi-level attention-guided context aggregation network model includes a downsampling layer, a fusion layer, and an upsampling layer connected in sequence. The fusion layer includes an interaction self-attention-guided Kolmogorov-Arnold representation module, a multi-perception spatial attention module, a KAN-enhanced channel attention module, and a multi-dimensional feature fusion module. The KAN-enhanced channel attention module is applied to the 4th layer in the downsampling process, the 5th layer in the downsampling process, and the 4th layer in the upsampling process. The multi-dimensional feature fusion module is applied to the skip connection layers of the 4th layer and the 5th layer.
[0015] The specific process of using the multi-level attention-guided context aggregation network model to process the images in the LLAS dataset is as follows:
[0016] The input feature image T ∈ (B, C, H, W) passes through three continuously set downsampling layers to obtain the reshaped feature map T 1 , the extreme value feature map T 2 and the mean feature map T 3 ;
[0017] Taking the extreme value feature map T 2 and the mean feature map T 3 as inputs and passing them to a 3×3 depth convolution block to obtain the image information T' 2 and the image information T' 3 ;
[0018] Halve the number of channels in the U-Net model;
[0019] Pass the image information T' 2 to the interaction self-attention-guided Kolmogorov-Arnold representation module, pass the image information T' 3 to the multi-perception spatial attention module. The image information T' 2 and the image information T' 3 are respectively processed by the interaction self-attention-guided Kolmogorov-Arnold representation module and the multi-perception spatial attention module to obtain high-resolution feature information;
[0020] The upsampling layer receives the high-resolution feature information processed by the interaction self-attention-guided Kolmogorov-Arnold representation module and the multi-perception spatial attention module, and based on the high-resolution feature information, obtains the segmented glue point targets.
[0021] The interaction self-attention-guided Kolmogorov-Arnold representation module includes an additive interaction self-attention module and a Kolmogorov-Arnold network, and the Kolmogorov-Arnold network has two layers;
[0022] The operation process of the additive interaction self-attention module is as follows:
[0023] 1) The input feature image \(T\in(B,C,H,W)\) is expanded in the channel dimension through a channel expansion convolution operation to obtain \(T'\in(B,3C,H,W)\);
[0024] 2) \(Q\), \(K\), and \(V\) are obtained through channel chunking operations, where \(Q\in(B,C,H,W)\), \(K\in(B,C,H,W)\), \(V\in(B,C,H,W)\), \(Q\) is the Query, \(K\) is the Key, and \(V\) is the Value;
[0025] 3) The similarity function is defined as the sum of the context scores of \(Q\) and \(K\):
[0026] \(\text{Sim}(Q,K)=F(Q)+F(K)\), s.t. \(F(Q)=S(C(Q))\);
[0027] where: \(F(\cdot)\) is the context mapping function and it contains necessary information interaction; \(F(\cdot)\) is specifically the spatial attention \(S(\cdot)\in(B,C,H,W)\) and the channel attention \(C(\cdot)\in(B,C,H,W)\);
[0028] 4) Based on the similarity function defined as the sum of the context scores of \(Q\) and \(K\), the final output feature \(\text{Attn}(Q,K,V)\) is obtained;
[0029] After the output feature \(\text{Attn}(Q,K,V)\) is further processed by depthwise separable convolution, the original input image \(T\in(B,C,H,W)\) is added as a residual to obtain the final output feature \(T''\in(B,C,H,W)\) of the additive interaction self-attention module;
[0030] The operation process of the Kolmogorov - Arnold network is as follows:
[0031] The final output feature \(T''\in(B,C,H,W)\) of the additive interaction self-attention module is reconstructed into two-dimensional patch features \(T''\) r \(\in(B*H*W,C)\);
[0032] The two-dimensional patch features \(T''\) r \(\in(B*H*W,C)\) are projected to a low dimension and then embedded back into the original dimension \(C\) and reshaped to the original tensor dimension;
[0033] The final output feature \(T''\in(B,C,H,W)\) of the additive interaction self-attention module is added as a residual to obtain the final IKA output image \(T\) IKA \(\in(B,C,H,W)\).
[0034] The operation process of the multi-sensory spatial attention module is as follows:
[0035] (ⅰ). The input feature image T ∈ (B, C, H, W) is successively subjected to feature extraction through 4 convolutional kernels with sizes of 1×1, 3×3, 5×5, and 7×7 respectively to obtain image features.
[0036] (ⅱ). Use average pooling to compress the information of the image features extracted in step (ⅰ), and then restore the image to the original size through nearest neighbor interpolation to obtain the reshaped feature map T 1 ∈ (B, 1, H*W);
[0037] (ⅲ). On the second and third branches of the multi-sensory spatial attention module, perform global max pooling and global average pooling operations respectively to obtain the extreme value feature information and mean feature information of the input feature image T ∈ (B, C, H, W), and reconstruct the extreme value feature information and mean feature information into the extreme value feature map T 2 ∈ (B, H*W, 1) and the mean feature map T 3 ∈ (B, 1, H*W);
[0038] (ⅳ). Perform matrix multiplication on the extreme value feature map T 2 and the mean feature map T 3 to obtain the feature map T 4 ∈ (B, H*W, H*W);
[0039] Perform matrix multiplication on the reshaped feature map T 1 and the feature map T 4 . After the result is reconstructed and subjected to Sigmoid activation operation, score the input feature image T ∈ (B, C, H, W) to obtain the spatial attention feature map T 5 ∈ (B, C, H, W).
[0040] The operation process of the KAN-enhanced channel attention module is as follows:
[0041] (Ⅰ). The input feature image T ∈ (B, C, H, W) is respectively subjected to feature compression through global max pooling and global average pooling.
[0042] (Ⅱ). Reconstruct the compressed image features to obtain the reconstructed feature maps T' 1 ∈ (B, 1, C) and the reconstructed feature map T' 2 ∈ (B, 1, C);
[0043] The reconstructed feature map T' 1 ∈ (B, 1, C) and the reconstructed feature map T' 2∈(B, 1, C) is fused through a one-dimensional convolutional kernel of size 3, and after being activated by the Sigmoid function and reconstructed, it scores the original input feature image T∈(B, C, H, W) to obtain the final channel attention output feature T'. 3 ∈(B, C, H, W).
[0044] The multi-dimensional feature fusion module includes a low layer, a current layer, and a high layer, the low layer, the current layer, and the high layer;
[0045] Let the low-layer input feature be The current layer input feature is T∈(B, C, H, W), and the high layer input feature is
[0046] The operation process of the multi-dimensional feature fusion module is as follows:
[0047] Align the low-layer input feature between the current layer input feature T∈(B, C, H, W) and the high layer input feature and the current layer input feature T∈(B, C, H, W) to ensure the same feature size;
[0048] The low-layer input feature First, perform downsampling through detail-preserving downsampling to obtain the low-layer feature T' low ∈(B, C, H, W);
[0049] After the current layer input feature T∈(B, C, H, W) extracts features through a 1×1 convolutional layer, the current layer feature T'∈(B, C, H, W) is obtained;
[0050] The high layer input feature Performs upsampling using the method of linear interpolation to obtain the high layer feature T' high ∈(B, C, H, W);
[0051] The low-layer feature T' low The current layer feature T' and the high layer feature T' high After passing through the 1×1 convolutional operation and being activated by the Sigmoid function respectively, generate the spatial weight matrices Wlow, Wcur, and Whigh from different layers. Each element in the spatial weight matrices Wlow, Wcur, and Whigh represents the possibility that its corresponding point belongs to the target;
[0052] Concatenate the spatial weight matrices Wlow, Wcur, and Whigh, and after passing through the 1×1 convolutional operation and being activated by the Sigmoid function, generate the final spatial weight matrix W;
[0053] The low-layer feature T' low, the current layer feature T′ and the high - level feature T′ high All pass through the intelligent channel integration module for channel fusion to obtain the ICI fusion feature T ICI ∈(B, C, H, W);
[0054] Use the spatial weight matrix W and the T after channel fusion 1 Perform element - by - element multiplication to achieve joint weighted fusion in both spatial and channel dimensions, obtaining the final output result T of the multi - dimensional feature fusion module MFF ∈(B, C, H, W).
[0055] The multi - dimensional feature fusion module includes a detail - preserving downsampling module and an intelligent channel integration module;
[0056] The operation process of the detail - preserving downsampling module is as follows:
[0057] Cut the input low - level input feature In a symmetric manner along the height and width directions, cutting the input picture into 4 parts;
[0058] Concatenate the four features after cutting and compression, as well as the features after max - pooling and average - pooling, and perform fine feature extraction through a 1×1 convolutional kernel, and finally output the downsampled low - level feature T′ low ∈(B, C, H, W);
[0059] The operation process of the intelligent channel integration module is as follows:
[0060] For the low - level feature T′ low that is already aligned in dimensions, the current layer feature T′, and the high - level feature T′ high , use the channel splitting technology to split the feature map into 4 equal parts in the channel dimension, respectively obtaining the low - level segmentation feature the current layer segmentation feature and the high - level segmentation feature where, T′ low_i is the i - th divided part of the low - level feature, T′ i is the i - th divided part of the current layer feature, T′ high_i is the i - th divided part of the high - level feature, i=(1, 2, 3, 4);
[0061] For the low - level segmentation feature the current layer segmentation feature and the high - level segmentation feature Calculate according to the following formula, so that the intelligent channel integration module can intelligently select which layer's channels to increase the weight according to the target features:
[0062] p i= Sigmoid(T′ i );
[0063]
[0064] T ICI = Concat(T′ g_1 , T′ g_2 , T′ g_3 , T′ g_4 );
[0065]
[0066] Where p i is the feature weight obtained after the feature T′ i passes through the activation function Sigmoid, and T′ g_i represents the intelligent aggregation result of each segmentation;
[0067] Concatenate the 4 equal parts after segmentation in the channel dimension to obtain T ICI ∈R B×C×H×W ;
[0068] When p i > 0.5, the multi-level attention-guided context aggregation network model enhances the weight of the fine-grained features transmitted by the low-level network;
[0069] When p i < 0.5, the multi-level attention-guided context aggregation network model enhances the weight of the globally semantic-rich features transmitted by the high-level network.
[0070] Compared with the prior art, the present invention has the following beneficial effects:
[0071] This application proposes a Multi-level Attention-guided Context Aggregation Network (MACAN) based on the self-built LLAS dataset. Its purpose is to achieve real-time and robust detection of the gluing points on the laser chip and provide the necessary quality guarantee for achieving high coupling efficiency of high-power laser lenses. The multi-level attention-driven mechanism enhances various attention mechanisms by widely applying the KAN network in different convolutional layers, effectively guiding complex non-linear information to focus on the target area. The fusion network improves the encoding efficiency and robustness of the network by adaptively selecting and fusing the target information from multi-level features in the convolutional layer. IKA reduces the multiplication operations in the traditional self-attention module by adding an Interactive Self-Attention (AISA) sub-module, thus reducing the computational cost. This sub-module introduces channel and spatial attention mechanisms, enhancing the interaction between spatial and channel information. In addition, the application of the KAN network effectively captures global context information, can accurately extract target features, and at the same time suppresses the interference of complex background noise. MPSA enhances the model's ability to capture targets at different scales by introducing a multi-sensing mechanism in the spatial attention module to handle the dual-target task of this application. KECA enhances the model's ability to model complex non-linear relationships by integrating the KAN network into the channel attention module, and this module is applied to the deep network, thereby improving the model's performance when dealing with complex tasks and high-dimensional channel data challenges. The MFF module suppresses the noise in the fusion process by adaptively selecting and fusing the target information from multi-level features, enhances the feature representation ability, and significantly improves the encoding efficiency and robustness of the network.
[0072] In addition to the purposes, features, and advantages described above, the present invention has other purposes, features, and advantages. The following will refer to the drawings for a further detailed description of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] The drawings constituting a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0074] Figure 1 is a schematic structural diagram of the MACAN model in an embodiment of the present invention;
[0075] Figure 2 is a schematic structural diagram of the IKA module in an embodiment of the present invention;
[0076] Figure 3 is a schematic structural diagram of the MPSA module in an embodiment of the present invention;
[0077] Figure 4 is a schematic structural diagram of the KECA module in an embodiment of the present invention;
[0078] Figure 5 It is a schematic structural diagram of the MFF module in an embodiment of the present invention;
[0079] Figure 6 It is a schematic structural diagram of the DPD module in an embodiment of the present invention;
[0080] Figure 7 It is a schematic structural diagram of the ICI module in an embodiment of the present invention;
[0081] Figure 8 It is a schematic diagram of the results of the ablation experiment in an experimental example of the present invention;
[0082] Figure 9 It is a schematic diagram of the results of the comparative experiment in an experimental example of the present invention;
[0083] Figure 10 It is an ROC curve graph in an experimental example of the present invention;
[0084] Figure 11 It is a schematic diagram of the measurement results of the glue dot parameters in an embodiment of the present invention. Detailed implementation manners
[0085] To make the above objects, features, and advantages of the present invention more clearly understandable, the following will describe in detail the specific implementation manners of the present invention with reference to the accompanying drawings. It should be noted that the drawings of the present invention are all in simplified forms and use non-precise scales, only for conveniently and clearly assisting in explaining the implementation of the present invention; the several mentioned in the present invention are not limited to the specific quantities in the drawing examples; the orientation or positional relationships indicated by 'front','middle', 'back', 'left', 'right', 'up', 'down', 'top', 'bottom','middle', etc. in the present invention are all based on the orientation or positional relationships shown in the drawings of the present invention, and do not indicate or imply that the devices or components referred to must have a specific orientation, nor can it be understood as a limitation to the present invention.
[0086] Embodiment:
[0087] A multi-scale glue dot detection method provided by the present invention includes the following steps:
[0088] Step 1: Use an intelligent laser lens dispensing platform to collect images of the laser chip before bonding; and establish an LLAS dataset based on the collected images;
[0089] Establish a multi-level attention-guided context aggregation network model;
[0090] Step 2: Apply the multi-level attention-guided context aggregation network model to process the images in the LLAS dataset, segment the glue dot targets, so as to realize more real-time and robust detection of the glue dots of the laser collimating lens, and provide necessary quality guarantee for realizing high coupling efficiency of the high-power laser collimating lens.
[0091] Step 3: Calculate the parameters of the segmented glue dot targets to obtain the glue dot detection results.
[0092] Specifically, by calculating the parameters of the segmented glue dot targets, it is further determined whether the glue dots are qualified to prompt the operator what operation to perform next.
[0093] Furthermore, the method for determining whether the glue dots are qualified is as follows:
[0094] Glue dot area;
[0095] Calculate the total area within the contour according to the segmented glue dot contour. An overly large glue dot area will cause glue overflow and significant glue offset after bonding the collimating lens. An overly small glue dot area will pose a risk of the collimating lens becoming loose.
[0096] Glue dot roundness;
[0097] Fit the glue dot into an ellipse, and then judge the dimensions of the major axis and minor axis of the ellipse. In an ideal state, the glue dot is circular. An overly long major axis or an overly short minor axis will both pose risks such as glue overflow or insecure bonding.
[0098] Glue dot spacing;
[0099] Measure the glue dot spacing based on the elliptical centers of two glue dots, including the horizontal spacing and the vertical spacing. An overly large horizontal distance will exceed the lens length, and an overly small distance will cause the glue dots to concentrate in the middle section of the lens while leaving the two ends suspended, affecting the bonding accuracy and reliability. The ideal vertical distance is 0. An overly large vertical distance will pose risks such as the glue dots detaching from the lens or causing the lens to shift.
[0100] Further preferably, according to the design dimensions of the collimating lens and long-term practical engineering experience, the index thresholds shown in Table 1 are determined.
[0101] Table 1: Comparison and evaluation of the measurement results of glue dot quality
[0102]
[0103] In actual production, most of the glue dots meet the design requirements. To highlight the ability of MACAN in segmenting abnormal glue dots, Figure 11Among them, five typical pictures containing normal and abnormal glue dots were selected, and the parameters of the glue dots segmented by the MACAN network and the real (ground truth) glue dots were measured respectively. Table 1 quantifies the measurement results and evaluates the quality of the glue dots. It can be seen from Table 1 that in the measurement results of parameters such as the area, roundness, and spacing of the target, the deviation between the measurement results of the method proposed in this application and the real results is small. At the same time, the glue dot detection can cover a wide range of scales from 0.002 to 1.421 mm 2 and has a small deviation from the real value. Compared with the traditional method, the MACAN network reduces the need for manual inspection and improves the automation level of the system.
[0104] Furthermore, as shown in Figure 1 , MACAN uses the commonly used U-Net model in the field of image segmentation as the backbone network, and fuses the IKA module, MPSA module, KECA module, and MFF module through a new attention mechanism based on the Kolmogorov-Arnold representation, enhancing the multi-scale and multi-level information fusion. MACAN is based on the commonly used U-Net backbone network in the field of image segmentation, aiming to achieve the segmentation of glue dot targets.
[0105] Furthermore, the MACAN includes a downsampling layer, a fusion layer, and an upsampling layer.
[0106] Furthermore, the process of the MACAN for processing images is as follows:
[0107] The input feature image T ∈ (B, C, H, W) passes through three continuously set downsampling layers to obtain the reshaped feature map T 1 , the extreme value feature map T 2 and the mean feature map T 3 ;
[0108] Taking the extreme value feature map T 2 and the mean feature map T 3 as inputs and passing them to a 3×3 depth convolution block to obtain the image information T′ 2 and the image information T′ 3 respectively; specifically, this 3×3 depth convolution block can enhance the expression ability of local features at low computational cost;
[0109] Halve the number of channels in the U-Net model;
[0110] Pass the image information T′ 2 to the IKA module, pass the image information T′ 3 to the MPSA module, the image information T′ 2 and the image information T′ 3After being processed by the IKA module and the MPSA module respectively, the target features are directed to a specific spatial position range to obtain high-resolution feature information. Among them, the IKA module can directly capture the relationships between pixels within the global scope, without being restricted by the size of the convolution kernel. When segmenting boundaries or complex shapes, the global information can help to more accurately identify different regions. The MPSA module, based on the spatial attention mechanism, can highlight important regions and target positions in the image, thereby reducing interference from the background and irrelevant regions and improving the accuracy of segmentation. After the feature extraction, for the image information processed by the IKA module and the MPSA module, the target features are directed to a specific spatial position range.
[0111] The upsampling layer receives the high-resolution feature information processed by the IKA module and the MPSA module, which helps MACAN to retain more details when restoring the resolution.
[0112] The KECA module is applied to the 4th layer in the downsampling process, the 5th layer in the downsampling process, and the 4th layer in the upsampling process, which is used to enhance the attention to key features while suppressing irrelevant or redundant features during the process of processing rich abstract features in the deep network, and improve the expressive ability of features.
[0113] The MFF module is applied to the skip connection layers of the 4th layer and the 5th layer. By adaptively selecting and fusing the target information in multi-level features, it can effectively suppress noise, retain target details, and enhance the information encoding efficiency and robustness of the model.
[0114] Further, as shown in Figure 2 The IKA module (Interactive Self-Attention Guided Kolmogorov-Arnold Representation Module) includes the AISA module and the KAN network. The AISA module (Additive Interactive Self-Attention Module) constructs a novel additive similarity function and uses channel and spatial attention as new forms of information interaction, thus avoiding complex matrix multiplication and Softmax operations.
[0115] Further preferably, the operation process of the AISA module is as follows:
[0116] 1) The input feature image T ∈ (B, C, H, W) undergoes a channel expansion convolution operation to be expanded in the channel dimension to obtain T′ ∈ (B, 3C, H, W);
[0117] 2) Through channel block operations, Query (Q), Key (K), and Value (V) are obtained, where Q ∈ (B, C, H, W), K ∈ (B, C, H, W), and V ∈ (B, C, H, W);
[0118] 3) The similarity function is defined as the sum of the context scores of Q and K:
[0119] Sim(Q,K) = F(Q) + F(K), s.t. F(Q) = S(C(Q));
[0120] Where: F(·) is a context mapping function that contains necessary information interaction. Specifically, F(·) is specified as spatial attention S(·) ∈ (B, C, H, W) and channel attention C(·) ∈ (B, C, H, W), and residual connections are used for each spatial attention and channel attention, adding the information before the attention module processing as residuals; then, through the combination of spatial and channel attention, the AISA module can promote multi-information interaction and effectively capture global context information. At the same time, the similarity function adds the results of spatial and channel attention to replace the traditional dot product attention mechanism, and depthwise separable convolution is used to further extract the added information. These measures significantly reduce the computational complexity, making the model more suitable for deployment on resource-constrained devices.
[0121] 4) The sum of the context scores obtained by the additive similarity function is further processed by depthwise separable convolution and used to weight V, thereby obtaining the final output feature Attn(Q, K, V);
[0122] The expression of the output feature Attn(Q, K, V) is as follows:
[0123]
[0124] Where: Ψ(·) is a linear transformer used to integrate context information; is a matrix multiplication operation.
[0125] After the output feature Attn(Q, K, V) is further processed by depthwise separable convolution, the original input image T ∈ (B, C, H, W) is added as a residual to obtain the final output feature T″ ∈ (B, C, H, W) of the AISA module.
[0126] Specifically, the calculation process of the entire AISA module is summarized as:
[0127] AISA(T) = DWConv(Attn(Chunk(Conv(T)))) + T.
[0128] Further preferably, in this application, 2 layers of KAN networks (Kolmogorov - Arnold networks) are set; the operation process of the KAN network is as follows:
[0129] The final output feature T″ of the AISA module ∈ (B, C, H, W) is reconstructed into a series of flattened two-dimensional patches T″ r ∈ (B * H * W, C);
[0130] The two-dimensional patch T″ r ∈ (B * H * W, C) is projected into a lower dimension to reduce the parameter overhead, then it is embedded and projected back to the original dimension C, and finally reshaped into the original tensor dimension;
[0131] Then, the final output feature T″ of the AISA module ∈ (B, C, H, W) is added as a residual to obtain the final IKA output image T IKA ∈ (B, C, H, W).
[0132] Specifically, the entire process of the KAN network can be described as:
[0133] KAN(T″) = (Reshape(K 2 (K 1 (Reshape(T″)))))+T″;
[0134] Furthermore, the entire calculation process of the IKA module can be summarized as:
[0135] IKA(T) = KAN((AISA(T))).
[0136] Furthermore, in the dual-object detection task of this application, the size of the glue dots has a certain degree of randomness. Even on the same chip, there are both large-size and small-size targets. Therefore, in the channel attention module, a multi-receptive field perception mechanism is introduced to enhance the model's ability to capture targets of different scales, which is called the multi-perception spatial attention module (MPSA).
[0137] Specifically, as shown in Figure 3 Let the input feature image T ∈ (B, C, H, W), and the detailed steps of the MPSA module (multi-perception spatial attention module) are as follows:
[0138] (ⅰ) The input feature image T ∈ (B, C, H, W) is successively passed through 4 convolutional kernels with sizes of 1×1, 3×3, 5×5, and 7×7 for feature extraction to obtain image features; by using convolutional kernels of different sizes on the input feature image T ∈ (B, C, H, W), the perception ability of multiple receptive fields is incorporated into the MACAN model.
[0139] (ii) Use average pooling to compress the information of the image features extracted in step (i), and then restore the image to its original size through nearest neighbor interpolation (NNI) to obtain the reshaped feature map T 1 ∈(B, 1, H*W); Specifically, average pooling effectively compresses the redundant information in the image by taking the average of the pixel values in the local area, and at the same time helps to remove some noise, which enables the network to focus on higher-level features. When restoring the image size through nearest neighbor interpolation, no new pixel values or grayscale information are introduced, but the nearest pixels are simply copied. This method avoids introducing unnecessary assumptions, preserves the original features of the image, and reduces information distortion during the image operation process.
[0140] (iii) On the second and third branches of the MPSA module, perform global max pooling (GMP) and global average pooling (GAP) operations respectively to obtain the extreme value feature information and mean feature information of the input feature image T∈(B, C, H, W), and reconstruct the extreme value feature information and mean feature information into the extreme value feature map T 2 ∈(B, H*W, 1) and the mean feature map T 3 ∈(B, 1, H*W).
[0141] (iv) Perform matrix multiplication on the extreme value feature map T 2 and the mean feature map T 3 to obtain the feature map T 4 ∈(B, H*W, H*W).
[0142] Perform matrix multiplication on the reshaped feature map T 1 ∈(B, 1, H*W) and the feature map T 4 ∈(B, H*W, H*W). After the result is reconstructed and passed through the Sigmoid activation operation, score the input feature image T∈(B, C, H, W) to obtain the spatial attention feature map T 5 ∈(B, C, H, W).
[0143] Further preferably, performing multiple matrix multiplications in the MPSA module can effectively calculate the spatial correlation between different regions of the input feature map, enabling the model to dynamically adjust the importance of each region in the feature map, thereby highlighting the target region and suppressing background interference. This operation not only efficiently fuses global information but also enhances the model's perception ability of the target region and improves the accuracy of segmentation or detection.
[0144] Preferably, the specific process of the MPSA module can be summarized by the following equations:
[0145] T 1 = Reshape(NNI(AvgP(Conv1(T) + Conv3(T) + Conv5(T) + Conv7(T))))
[0146] T 2 = Reshape(GMP(T))
[0147] T 3 = Reshape(GAP(T))
[0148]
[0149] Where: Conv1(·), Conv3(·), Conv5(·), and Conv7(·) respectively represent convolutional kernels with sizes of 1×1, 3×3, 5×5, and 7×7; AvgP(·) is the average pooling operation; NNT(·) is the operation of the NNT module; Reshape is to reconstruct the shape of the feature tensor; GMP(·) is the global maximum pooling; GAP(·) is the global average pooling; represents the matrix multiplication operation; Sigmoid is the activation function.
[0150] Furthermore, as shown in Figure 4 the detailed steps of the KECA module (KAN-enhanced channel attention module) are as follows:
[0151] (Ⅰ) The input feature image T ∈ (B, C, H, W) is respectively subjected to feature compression through global maximum pooling (GMP) and global average pooling (GAP), which can avoid overemphasis on a certain type of feature that may be caused by using a single pooling method.
[0152] (Ⅱ) Reconstruct the compressed image features to obtain reconstructed feature maps T' 1 ∈ (B, 1, C) and reconstructed feature map T' 2 ∈ (B, 1, C); By reconstructing the compressed image features, two KAN networks (the two KAN networks are respectively named KAN1 and KAN2) can be used to more effectively capture the complex non-linear dependencies between channel features. The specific working process of the KAN network is the same as that in the IKA module.
[0153] Reconstructed feature map T' 1 ∈ (B, 1, C) and reconstructed feature map T' 2∈(B, 1, C) is fused through a one-dimensional convolutional kernel of size 3, then activated and reconstructed through the Sigmoid function, and the original input feature image T∈(B, C, H, W) is scored to obtain the final channel attention output feature T'. 3 ∈(B, C, H, W).
[0154] Preferably, the specific process of the KECA module can be summarized by the following equation:
[0155] T' 1 = Reshape(KAN1(Reshape(GMP(T))));
[0156] T' 2 = Reshape(KAN2(Reshape(GAP(T))));
[0157]
[0158] Among them, the specific calculation processes of KAN1(·) and KAN2(·) are the same as the specific calculation process of KAN(·) in the IKA module; Conv1D()Conv1D(·) is a one-dimensional convolutional kernel with a size of 3.
[0159] Furthermore, in the laser collimating lens glue point detection of this application, the glue points have random shapes. This requires that the deep learning network can not only accurately identify the positions of the glue points, but also clearly identify the irregular contours of the glue points and effectively suppress noise interference. For this reason, a multi-dimensional feature fusion module (MFF) is set on the skip connection layers of the fourth and fifth layers of the MACAN network, aiming to fuse high-level image features and low-level image features, so that the extracted features contain both clear texture and edge information and are not affected by noise.
[0160] See Figure 5 As shown, the MFF module (multi-dimensional feature fusion module) is used to fuse image features from 3 different layers; that is, the MFF module includes a low layer, a current layer, and a high layer, a low layer, a current layer, and a high layer;
[0161] Let the low-layer input feature be The current layer input feature is T∈(B, C, H, W), and the high-layer input feature is
[0162] The specific operation process is as follows:
[0163] 1), between the low-layer feature and the image feature T∈(B, C, H, W) of the current layer, and the high-layer feature are aligned with the image features T ∈ (B, C, H, W) of the current layer to ensure the same feature size;
[0164] 2) Generate a spatial weight matrix and perform fusion;
[0165] The low-level feature T' low , the current layer feature T', and the high-level feature T' high After passing through a 1×1 convolution operation and activation by the Sigmoid function respectively, spatial weight matrices Wlow, Wcur, and Whigh from different layers are generated. Each element in the spatial weight matrices Wlow, Wcur, and Whigh represents the possibility that its corresponding point belongs to the target;
[0166] The spatial weight matrices Wlow, Wcur, and Whigh are concatenated, and after passing through a 1×1 convolution operation and activation by the Sigmoid function, the final spatial weight matrix W is generated. The spatial weight matrix W combines enhanced target details and precise target localization while effectively suppressing background noise.
[0167] 3) Intelligently fuse the channels of the image features of each layer;
[0168] The low-level feature T' low , the current layer feature T', and the high-level feature T' high All pass through the intelligent channel integration module (ICI) for channel fusion to obtain the ICI fusion feature T ICI ∈ (B, C, H, W); The ICI module can intelligently select channels from different layers for fusion according to the features of the target.
[0169] Perform element-wise multiplication of the spatial weight matrix W and the T after channel fusion 1 to achieve joint weighted fusion in the spatial and channel dimensions, and obtain the output result T of the final multi-dimensional feature fusion module (MFF) MFF ∈ (B, C, H, W). Performing element-wise multiplication of the spatial weight matrix and the channel fusion feature can effectively integrate spatial information and channel information, improve the accuracy and robustness of the model, and perform better especially in the task of accurately identifying target details from complex backgrounds.
[0170] Further preferably, in order to retain target details, the low-level image features first undergo downsampling through detail-preserving downsampling (DPD) to obtain the low-level feature T' low ∈ (B, C, H, W);
[0171] The image features T ∈ (B, C, H, W) of the current layer are refined and feature-extracted through a 1×1 convolution layer to obtain the current layer feature T' ∈ (B, C, H, W);
[0172] High-level image features Upsample using linear interpolation to obtain the high-level feature T' high ∈(B, C, H, W).
[0173] Preferably, see Figure 6 As shown, the operation process of the DPD module (detail-preserving downsampling module) is as follows:
[0174] Cut the input low-level image features Symmetrically along the height and width directions, and cut the input picture into 4 parts;
[0175] The input features also undergo max-pooling operation and average-pooling operation for feature extraction, taking into account the characteristics that max-pooling operation can extract the most significant features in the local area and the average-pooling operation can smooth details, enhancing the model's expressive ability and improving robustness;
[0176] Concatenate the four features after cutting and compression, and the features after max-pooling and average-pooling, and perform fine feature extraction through a 1×1 convolutional kernel, and finally output the downsampled feature T' low ∈(B, C, H, W).
[0177] Further preferably, the calculation process of the DPD module is summarized by the formula:
[0178] DPD(T low ) = Conv(Concat(Cut(T low ), Maxpool(T low ), Avgpool(T low ))).
[0179] Preferably, see Figure 7 As shown, the operation process of the ICI module (intelligent channel integration module) is as follows:
[0180] For the low-level feature T' low that has been aligned in dimensions, the current layer feature T', and the high-level feature T' high , use the channel splitting technology to split the feature map into 4 equal parts in the channel, and respectively obtain the low-level segmentation feature The current layer segmentation feature And the high-level segmentation feature Among them, T' low_i Is the i-th divided part of the low-level feature, T' i Is the i-th divided part of the current layer feature, T' high_i Is the i-th divided part of the high-level feature, i = (1, 2, 3, 4);
[0181] For low-level segmentation features Current layer segmentation features and high-level segmentation features Calculate according to the following formula so that the ICI module can intelligently select which layer of channels to increase the weight according to the target features:
[0182] p i = Sigmoid(T′ i );
[0183]
[0184] T ICI = Concat(T′ g_1 , T′ g_2 , T′ g_3 , T′ g_4 );
[0185]
[0186] Among them, p i is the feature weight obtained after the feature T′ i passes through the activation function Sigmoid, and T′ g_i represents the intelligent aggregation result of each segmentation;
[0187] Concatenate the 4 equal parts after segmentation in the channel dimension to obtain T ICI ∈R B×C×H×W . In such a setting, when p i > 0.5, MACAN will enhance the weight of the fine-grained features transmitted by the low-level network; if p i < 0.5, MACAN will enhance the weight of the semantically rich global features transmitted by the high-level network, so as to achieve the intelligent aggregation of features at different levels.
[0188] Preferably, the calculation process of the entire MFF can be summarized by the formula:
[0189] W low (T low ) = Sigmoid(Conv(DPD(T low )));
[0190] W cur (T) = Sigmoid(Conv(Conv(T)));
[0191] W high (T high ) = Sigmoid(Conv(Up(T high )));
[0192] W(W low ,W cur ,W high ) = Sigmoid(Conv(Concat(W low ,W cur ,W high )));
[0193]
[0194] Experimental Example:
[0195] (I) Experimental Results and Analysis
[0196] Based on the images taken before bonding of the laser chip, the LLAS dataset established contains a total of 3,070 glue dot pictures, and the size of the pictures is 256×192 (width×height).
[0197] Specifically, 2,520 (82%) of them are used for training, 280 (9%) for testing, and the remaining 270 (9%) for validation.
[0198] All the experiments of the present invention are calculated through a GPU of 12G NVIDIA RTX 3060. The calculation process is completed based on the PyTorch architecture. The Adam optimizer is used for parameter optimization, and the binary cross-entropy loss function (BinaryCross-EntropyLoss, BCE Loss) is used to calculate the loss, which is used to measure the gap between the model prediction value and the true label. And indicators such as intersection over union (IoU), mean intersection over union (mIoU), F1 score, Dice, receiver operating characteristic curve (ROC), etc. are used to evaluate the performance of MACAN.
[0199] Specifically, the IoU indicator is used to measure the overlapping degree of the prediction result and the true label, and evaluate the spatial consistency; the mIoU averages the IoU of all samples and is used for overall performance evaluation; F1 comprehensively considers precision and recall and reflects the balanced performance of the classification model; Dice divides twice the intersection by the total area to evaluate the overlapping degree of the prediction and the true value and is more sensitive to small targets. The larger the value of the above indicators, the more accurate the model prediction. The ROC describes the classification ability of the model at different thresholds, shows the change relationship between the true positive rate (TPR) and the false positive rate (FPR), and the larger the area enclosed by the curve and the X-axis, the better the performance of the model. The specific calculation formulas are as follows:
[0200]
[0201]
[0202] Among them, A ∩ and A ∪ represent the areas of the intersection and union respectively. N and j represent the total number of test samples and the j-th sample respectively. True Positive represents the number of samples correctly predicted as positive classes, False Positive represents the number of samples that are actually negative classes but are wrongly predicted as positive classes, False Negative represents the number of samples that are actually positive classes but are wrongly predicted as negative classes, and True Negative represents the number of samples correctly predicted as negative classes. Recall is also called TPR in some scenarios.
[0203] (2) Ablation experiment research
[0204] To explore the effectiveness of each module designed in MACAN, ablation experiment research was carried out, and the evaluation results are shown in Table 1 and Figure 8 .
[0205] Specifically, the classic image segmentation network U-Net was used as the baseline. The results of the ablation experiment show that the numerical evaluation performance of only using the baseline network is poor. By adding the KECA module to the deep layer of the U-Net backbone network, the performance of the model in processing complex high-dimensional channel data is significantly improved, and the mIoU, F1, and Dice metrics are improved by 1.6%, 0.7%, and 0.9% respectively. When the IKA module is further introduced into the skip connection, the mIoU, F1, and Dice metrics are further increased by 1.3%, 0.5%, and 0.7%. When the MPSA module is introduced into the skip connection, the mIoU, F1, and Dice metrics are increased by 1.0%, 0.2%, and 0.5%. However, if the processing results of the IKA and MPSA modules are concatenated on the skip connection at the same time, the mIoU, F1, and Dice metrics are increased by 2.0%, 0.7%, and 1.1% compared with only adding the KECA module. This shows that since the IKA module can improve the global perception ability and the MPSA module has the fine positioning ability for multi-scale targets, when the two are combined, a powerful synergistic effect can be formed, which can more effectively cope with complex backgrounds and target scale differences. Finally, the setting of the MFF module adaptively fuses the target information in multi-level features, and further improves the mIoU, F1, and Dice metrics by 0.8%, 0.3%, and 0.4%. Finally, compared with the baseline network, the MACAN network improves the mIoU, F1, and Dice metrics by 4.4%, 1.7%, and 2.4% respectively.
[0206] In Figure 8Some qualitative results are also given to further demonstrate the effects of each module. In the pictures of sequences (1) and (2), there are no glue dots themselves, but the baseline network incorrectly identifies the existence of glue dots. In particular, due to the more serious dirt condition on the chip in sequence (1), the baseline network also identifies more false targets. When the KECA module is added, with the improvement of the module's performance in processing complex high-dimensional channel data, the false alarm problem is significantly improved. After adding the IKA and MPSA modules respectively, the false alarm problem is further improved. Especially for series (2), the occurrence of the false alarm problem has been completely eliminated. When both the IKA and MPSA modules are added to the model, the false alarm in sequence (1) with more complex background information is also completely suppressed. In the pictures of sequence (3), there are glue dots in the dispensing target areas. However, the pixel brightness of the glue dots on the left is similar to that of the background area, with insufficient contrast. It is difficult for the baseline network to segment an accurate boundary, and there are also some false alarms to a certain extent. With the addition of the KECA module, the false alarm problem is eliminated. At the same time, the difference between the recognized glue dot area and the true label decreases, but the contour boundary is still not accurate enough. When the IKA and MPSA modules are added respectively, the recognition accuracy of the contour boundary is improved, and the recognition result of the IKA module for the contour boundary is relatively more accurate. When both of these modules are added, especially when the MFF module is further added, the recognition accuracy of the contour boundary is significantly improved, and the final recognition result of the MACAN network for the contour boundary of the target in sequence (3) is highly consistent with the true label. In the pictures of sequence (4), there are two very small glue dots. The glue dot target on the left is particularly small. The baseline network can identify the existence of the glue dot on the right, but the contour is quite different from the true label, and the small target on the left cannot be recognized. The addition of the KECA and IKA modules still fails to recognize the small target on the left. When the MPSA module with a multi-receptive field perception mechanism is added, the small target on the left is detected. In the subsequent experiments of adding both the IKA and MPSA modules and further adding the MFF module, the accuracy of glue dot recognition gradually improves. In the pictures of sequence (5), the glue dot target is relatively ideal. Even the baseline network can segment the glue dot target relatively accurately. However, by carefully comparing the contour shapes, the advantage of the module proposed in this application in improving the segmentation accuracy can still be seen.
[0207] Table 2: Results of ablation experiments
[0208]
[0209] (III) Comparison with other advanced methods
[0210] In existing research reports, there is no segmentation algorithm for the glue points of the laser micro-lens. To evaluate the performance of MACAN on the LLAS dataset, networks that have achieved excellent results in fields such as infrared small target recognition and medical image segmentation are used to train and validate the LLAS dataset for comparison. The selected methods include TransUNet, MTU-Net, DNA-Net, HCF-Net, MRF3Net, and U-KAN. TransUNet introduces the ViT architecture into the U-Net network for medical image segmentation tasks. MTU-Net further constructs a multi-pole feature extraction network based on ViT. DNA-Net fuses a densely nested attention network structure for infrared small target detection tasks. HCF-Net improves the detection performance of infrared small targets by integrating a hierarchical context fusion network. MRF3Net fuses multi-receptive field perception and effective feature fusion strategies for infrared small target detection. U-KAN achieves medical image segmentation by integrating the KAN network into U-Net. The quantitative results of the comparison experiment are shown in Table II. Benefiting from the outstanding advantages of the multi-level attention-driven fusion network in dealing with complex non-linear problems, among the image segmentation results of all networks, the three indicators of mIoU, F1, and Dice are the highest. The multi-pole feature extraction network MTU-Net based on ViT achieved the second-best result. Due to the integration of the KAN network with strong non-linear expression ability, U-KAN also achieved results close to MTU-Net. TransUNet and HCF-Net performed at a medium level among all networks, while DNA-Net and MRF3Net achieved relatively poor results. The possible reason is that DNA-Net has a large width and MRF3Net has a relatively shallow depth. Although they have achieved good detection results in their respective tasks, the depth and width of the network need to be properly balanced to adapt to different task characteristics. After a series of optimized calculation strategies, MACAN achieved the best performance in terms of calculation time. However, the number of its parameters exceeds those of the MRF3Net, DNA-Net, and U-KAN networks. However, compared with the TransUNet and MTU-Net networks that also integrate the ViT module, MACAN has fewer parameters due to the efficiency improvement of the traditional ViT module in the IKA module.
[0211] Table 3: Comparison Results with Six Other Advanced Methods on the LLAS Dataset
[0212]
[0213] In Figure 9Six pictures with typical characteristics were selected to qualitatively compare the verification results of different networks. In the pictures of sequence (1), there are typical irregular shapes, especially the glue dots on the right side with relatively random shapes; in the pictures of sequence (2), there is a large amount of overflowing flux under the chip, making it difficult to distinguish the contour of the glue dots from the background; the pictures in sequence (3) were also shown in sequence (4) of the ablation experiment, which has the characteristic of small-sized targets in a complex background and can be used to well verify the effects of different networks in identifying tiny targets. The chips shown in sequences (4) and (5) are another relatively common type of chip in the LLAS dataset. This type of chip is attached to a black substrate and is extremely close in color to the glue dots, greatly increasing the difficulty of image recognition. And sequence (6) is a relatively ideal image recognition picture with the target being relatively clear in the figure. Thanks to the advantages of the KAN network in MACAN in dealing with complex non-linear problems, and the powerful ability of the ViT module inherited by IKA in capturing context information and long-range dependencies, the model can clearly identify the boundary contour of the target. By integrating the multi-receptive field perception mechanism, MACAN can accurately extract the target at different scales. In addition, the added multi-layer fusion module further enhances MACAN's ability to accurately extract the target from a complex background. Combining these advantages, MACAN achieved excellent recognition results in all images. The DNA-Net has relatively poor performance in the accuracy of target contour recognition. The possible reason is that the sample diversity of the LLAS dataset is limited, making it difficult to fully train the DNA-Net network with a dense nested design, and its advantages cannot be fully exerted. The MRF3Net network also performed unsatisfactorily in the recognition tasks of these typical images. Its relatively shallow network layers make it particularly difficult to accurately segment the target in complex backgrounds such as sequences (4) and (5), but the design of its multi-receptive field perception mechanism enables it to successfully identify the target in the tiny target perception task of sequence (3). The image segmentation results of TransUNet, HCF-Net, and U-KAN networks are relatively close, but U-KAN has higher segmentation accuracy when dealing with image tasks where the gray values of the target and the background are extremely close, such as sequences (4) and (5). And the MTU-Net network constructed based on modules such as ViT multi-level feature extraction achieved a segmentation effect second only to MACAN.
[0214] In Figure 10 the ROC curve was used to provide another intuitive way to evaluate the classification performance of each network. The larger the area enclosed by the curve and the X-axis, that is, the larger the value of the area under the curve (AUC), the better the performance of the network. As can be seen from the figure, the curve of the proposed network is closest to the upper left corner and has the largest AUC value, indicating that MACAN can meet the highest true positive rate at the lowest false positive rate.
[0215] To address the challenge of detecting the glue point defects of laser collimating lenses in complex backgrounds, especially in the presence of random shapes and complex backgrounds, this application proposes MACAN based on deep learning. MACAN introduces a multi-pole attention mechanism and an adaptively selected multi-layer fusion module to extract defects with random shapes and multi-scales, and optimizes the traditional attention mechanism through modules such as KAN and multi-receptive field perception mechanism to further suppress the interference brought by complex backgrounds. Ablation experiments and extensive comparison tests based on the self-built dataset LLAS have proved the superiority and robustness of MACAN. However, the current research still has certain limitations. For example, since the dataset is the glue point images collected during the formal production after the dispensing equipment has been fully debugged, most of the glue points in the dataset are relatively ideal and easy to identify. Therefore, there are relatively few pictures containing obvious defect features, and the sample diversity is insufficient. When the debugging effect of the dispensing equipment is not ideal and more complex features appear, the probability of manual re-inspection will increase. At the same time, although the MACAN network has achieved excellent segmentation accuracy results, compared with other more lightweight networks, it is not dominant in terms of operation time and the number of parameters. Therefore, in future research, the type and quantity of the dataset will be further expanded to further improve the training effect of the model. At the same time, lightweight research, while not reducing the accuracy, improves the advantages of the model in terms of calculation time and the number of parameters, which is more conducive to the deployment of the model.
[0216] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A multi-scale glue spot detection method, characterized in that: The following steps are involved: Step 1: Use the intelligent laser lens dispensing platform to collect images of the laser chip before bonding; and establish a LLAS data set based on the collected images; Establish a multi-level attention-guided context aggregation network model; Step 2: Use the multi-level attention-guided context aggregation network model to process the images in the LLAS dataset and segment the glue point targets; Step 3: Based on the segmented glue point targets, obtain the glue point detection results.
2. The multi-scale glue spot detection method according to claim 1, characterized in that: The multi-level attention-guided context aggregation network model adopts the U-Net model as the backbone network, and integrates the interactive self-attention-guided Kolmogorov-Arnold representation module, the multi-sensory spatial attention module, the KAN-enhanced channel attention module and the multi-dimensional feature fusion module through the attention mechanism based on the Kolmogorov-Arnold representation.
3. The multi-scale glue spot detection method according to claim 2, characterized in that: The multi-level attention-guided context aggregation network model includes a downsampling layer, a fusion layer and an upsampling layer connected in sequence, and the fusion layer includes an interactive self-attention-guided Kolmogorov-Arnold representation module, a multi-sensory spatial attention module, a KAN-enhanced channel attention module and a multi-dimensional feature fusion module; Apply the KAN-enhanced channel attention module to the 4th layer during downsampling, the 5th layer during downsampling, and the 4th layer during upsampling; Apply the multi-dimensional feature fusion module to the skip connection layers of the 4th and 5th layers; The specific process of applying the multi-level attention-guided context aggregation network model to process images in the LLAS dataset is as follows: The input feature image T∈(B,C,H,W) is subjected to three consecutive downsampling layers to obtain the reshaped feature map T1, the extreme feature map T2 and the mean feature map T3 respectively; The extreme feature map T2 and the mean feature map T3 are taken as input and passed to a 3×3 depth convolution block to obtain image information T2′ and image information T3′ respectively; Halve the number of channels in the U-Net model; The image information T2′ is passed to the interactive self-attention guided Kolmogorov-Arnold representation module, and the image information T3′ is passed to the multi-sensory spatial attention module. The image information T2′ and the image information T3′ are processed by the interactive self-attention guided Kolmogorov-Arnold representation module and the multi-sensory spatial attention module respectively to obtain high-resolution feature information; The upsampling layer receives Kolmogorov-Arnold representation modules and multi-sensory spatial attention from interactive self-attention guidance. The force module processes the high-resolution feature information, and based on the high-resolution feature information, obtains the segmented glue point target.
4. The multi-scale glue spot detection method according to claim 3, characterized in that: The interactive self-attention guided Kolmogorov-Arnold representation module includes an additive interactive self-attention module and a Kolmogorov-Arnold network, wherein the Kolmogorov-Arnold network has two layers; The operation process of the additive interactive self-attention module is as follows: 1) The input feature image T∈(B,C,H,W) is expanded in the channel dimension through the channel expansion convolution operation to obtain T′∈(B,3C,H,W); 2) Obtain Q, K and V through channel block operation, where Q∈(B,C,H,W), K∈(B,C,H,W), V∈(B,C,H,W), Q is Query, K is Key, and V is Value; 3) The similarity function is defined as the sum of the context scores of Q and K: Sim(Q,K)=F(Q)+F(K), stF(Q)=S(C(Q)); Where: F(·) is the context mapping function, and it contains the necessary information interaction; F(·) is concretized as spatial attention S(·)∈(B,C,H,W) and channel attention C(·)∈(B,C,H,W); 4) Based on the similarity function defined as the sum of the context scores of Q and K, the final output feature Attn(Q, K, V) is obtained; After the output feature Attn(Q,K,V) is further processed by depthwise separable convolution, the original input image T∈(B,C,H,W) is added as the residual to obtain the final output feature T″∈(B,C,H,W) of the additive interactive self-attention module; The operation process of the Kolmogorov-Arnold network is as follows: Reconstruct the final output feature T″∈(B,C,H,W) of the additive interactive self-attention module into a two-dimensional patch feature T″ r ∈(B*H*W,C); The two-dimensional patch feature T″ r ∈(B*H*W,C) projected to low dimension Then embed it back to the original dimension C and reshape it to the original tensor dimension; The final output feature T″∈(B,C,H,W) of the additive interactive self-attention module is added as the residual to obtain the final IKA output image T IKA ∈(B,C,H,W).
5. The multi-scale glue spot detection method according to claim 4, characterized in that: The operation process of the multi-sensory spatial attention module is as follows: (i) The input feature image T∈(B,C,H,W) is sequentially subjected to four convolution kernels of sizes 1×1, 3×3, 5×5, and 7×7 for feature extraction to obtain image features; (ii) Use average pooling to compress the image features extracted in step (i), and then restore the image to its original size through nearest neighbor interpolation to obtain a reshaped feature map T1∈(B,1,H*W); (iii) In the second branch and the third branch of the multi-sensory spatial attention module, the extreme feature information and the mean feature information of the input feature image T∈(B,C,H,W) are obtained by global maximum pooling and global average pooling respectively, and the extreme feature information and the mean feature information are reconstructed into the extreme feature map T2∈(B,H*W,1) and the mean feature map T3∈(B,1,H*W) respectively; (iv) Perform matrix multiplication on the extreme value feature map T2 and the mean feature map T3 to obtain the feature map T4∈(B,H*W,H*W); The reshaped feature map T1 is multiplied by the feature map T4 by matrix multiplication. The result is reconstructed and activated by Sigmoid. Then the input feature image T∈(B,C,H,W) is scored to obtain the spatial attention feature map T5∈(B,C,H,W).
6. The multi-scale glue spot detection method according to claim 5, characterized in that: The operation process of the KAN enhanced channel attention module is as follows: (I), the input feature image T∈(B,C,H,W) is subjected to feature compression through global maximum pooling and global average pooling respectively; (II) Reconstruct the compressed image features to obtain the reconstructed feature map T′1∈(B,1,C) and the reconstructed feature map T′2∈(B,1,C); The reconstructed feature map T′1∈(B,1,C) and the reconstructed feature map T′2∈(B,1,C) are fused through a one-dimensional convolution kernel of size 3. After being activated and reconstructed by the Sigmoid function, the original input feature image T∈(B,C,H,W) is scored to obtain the final channel attention output feature T′3∈(B,C,H,W).
7. The multi-scale glue spot detection method according to claim 6, characterized in that: The multi-dimensional feature fusion module includes a low layer, a current layer and a high layer, a low layer, a current layer and a high layer; Assume that the low-level input feature is The current layer input feature is T∈(B,C,H,W), and the high-level input feature is The operation process of the multi-dimensional feature fusion module is as follows: Input low-level features Between the current layer input feature T∈(B,C,H,W) and the high-level input feature Align with the current layer input feature T∈(B,C,H,W) to ensure the same feature size; Low-level input features First, downsample the sample by detail-preserving downsampling to obtain the low-level feature T′ low ∈(B,C,H,W); The current layer input feature T∈(B,C,H,W) is refined and extracted through a 1×1 convolution layer to obtain the current layer feature T′∈(B,C,H,W); High-level input features Use linear interpolation method to upsample and get high-level features T′ high ∈(B,C,H,W); Low-level features T′ low , current layer features T′ and high-level features T′ high After 1×1 convolution operation and Sigmoid function activation, the spatial weight matrices Wlow, Wcur and Whigh from different layers are generated. Each element in the spatial weight matrices Wlow, Wcur and Whigh represents the possibility that the corresponding point belongs to the target. The spatial weight matrices Wlow, Wcur and Whigh are concatenated and activated by a 1×1 convolution operation and a Sigmoid function to generate the final spatial weight matrix W. Low-level features T′ low , current layer features T′ and high-level features T′ high All channels are fused through the intelligent channel integration module to obtain the ICI fusion feature T ICI ∈(B,C,H,W); The spatial weight matrix W is used to perform element-by-element multiplication with the channel-fused T1 to achieve joint weighted fusion in the spatial and channel dimensions, and the final multi-dimensional feature fusion module output result T is obtained. MFF ∈(B,C,H,W).
8. The multi-scale glue spot detection method according to claim 6, characterized in that: The multi-dimensional feature fusion module includes a detail-preserving downsampling module and an intelligent channel integration module; The operation process of the detail-preserving downsampling module is as follows: Input low-level features Cut the input image into 4 parts in a symmetrical way along the height and width directions; The four features after cutting and compression, as well as the features after maximum pooling and average pooling are concatenated, and after fine feature extraction through a 1×1 convolution kernel, the downsampled low-level feature T′ is finally output. low ∈(B,C,H,W); The operation process of the intelligent channel integration module is as follows: For the low-level features T′ that have been aligned in dimension low , current layer features T′ and high-level features T′ high , using channel splitting technology, the feature map is split into 4 equal parts on the channel to obtain low-level segmentation features Current layer segmentation features and high-level segmentation features Among them, T′ low_i is the i-th partition of the low-level feature, T′ i is the i-th partition of the current layer feature, T′ high_i is the i-th partition of the high-level feature, i = (1, 2, 3, 4); For low-level segmentation features Current layer segmentation features and high-level segmentation features The calculation is performed according to the following formula, so that the intelligent channel integration module can intelligently select which layer of channels to increase the weight according to the characteristics of the target: p i =Sigmoid(T′ i ); T ICI =Concat(T′ g_1 ,T′ g_2 ,T′ g_3 ,T′ g_4 ); Among them, p i is the characteristic T′ i The feature weight obtained after the activation function Sigmoid, T′ g_i Represents the smart aggregation results for each segmentation; The four equal parts after segmentation are concatenated in the channel dimension to obtain T ICI ∈R B×C×H×W ; When p i When >0.5, the multi-level attention guides the context aggregation network model to enhance the weights of fine-grained features transmitted by the low-level network; When p i When <0.5, the multi-level attention-guided context aggregation network model enhances the weight of the semantically rich global features transmitted by the high-level network.
Citation Information
Patent Citations
SAR image road segmentation method based on attention mechanism
CN112883934A
Solar cell module defect EL detection method based on deep learning
CN113780434A
Double-feature fusion semantic segmentation system and method based on internet of things perception
WO2022227913A1