PTC thermistor defect detection method based on YOLOv11 improvement

By constructing an LMSF-YOLO detection network through an improved YOLOv11 algorithm, and combining it with RS-HGNetV2, ASF-V2, and Detect-SEAM modules, the problems of low efficiency and insufficient accuracy of traditional detection methods are solved, achieving efficient and accurate detection of defects in PTC thermistors, which is suitable for industrial applications.

CN121685497APending Publication Date: 2026-03-17FUZHOU UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511901370.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing PTC thermistor defect detection methods rely on manual visual inspection or traditional automated inspection, which are inefficient and easily affected by subjective factors. They are difficult to meet the high-precision and high-stability requirements of modern production, especially in complex background scenarios where the detection accuracy and robustness are insufficient.

Method used

A defect detection method for PTC thermistors based on improved YOLOv11 is adopted. By constructing an LMSF-YOLO detection network, combined with an RS-HGNetV2 backbone network, an ASF-V2 neck network and a Detect-SEAM detection head, the SimAM attention mechanism is introduced, and training parameters and data augmentation techniques are optimized to improve the detection accuracy and robustness of the model in complex backgrounds.

Benefits of technology

It enables efficient and accurate detection of defects in PTC thermistors under complex backgrounds, improves detection accuracy and robustness, is suitable for industrial testing, and reduces hardware requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121685497A_ABST
    Figure CN121685497A_ABST
Patent Text Reader

Abstract

The invention provides a PTC thermistor defect detection method based on YOLOv11 improvement. The PTC thermistor defect detection method comprises the following steps: acquiring and constructing a PTC surface image data set containing defects such as defects, color spots, bottom exposure, missing printing and the like; a lightweight detection network LMSF-YOLO is designed based on a YOLO series, and the network comprises an improved backbone RS-HGNetV2, an ASF-V2 neck structure in which an additional P2 feature layer and a SimAM attention mechanism are introduced, and a lightweight SEAM detection head; and training the LMSF-YOLO by using the data set and obtaining the optimal weight, inputting an image to be detected into the trained model, and outputting the category, the position and the confidence coefficient of each defect. The method has high detection precision and robustness while keeping the light weight of the model, and is suitable for online industrial detection and embedded deployment of the PTC thermistor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of defect image processing technology, and in particular to a defect detection method for PTC thermistors based on YOLOv11. Background Technology

[0002] With the rapid iteration of electronic information technology and the booming development of intelligent manufacturing, the quality and performance of PTC thermistors directly affect the safety and reliability of end products, and are of great significance to the safe and stable operation of the entire electronics industry chain. Accurate defect detection of PTC thermistors is a core link in ensuring their quality, which can not only effectively reduce product failure rates and improve production efficiency, but also enhance the core competitiveness of enterprises in the market. However, ceramic sheets are easily affected by various external factors during production and processing, resulting in defects such as defects, color spots, cracks, and exposed substrate. These defects lead to a decline in product quality, thereby damaging the performance of the material itself and increasing enterprise costs. Therefore, research on surface defect detection of PTC thermistors has significant engineering application value.

[0003] Current detection methods largely rely on manual visual inspection or traditional automated inspection based on fixed rules. This is not only labor-intensive and inefficient, but also susceptible to subjective factors leading to missed or false detections. Traditional rule-based methods are ill-suited to the complex and diverse defect morphologies, failing to meet the demands of high precision and stability in modern production. In recent years, deep learning-based target detection algorithms have gradually become mainstream. Among them, the single-stage YOLO series, balancing accuracy and real-time performance, has shown promising application prospects in electronic component defect detection. However, the general-purpose YOLOv11 model is not specifically optimized for small targets, low contrast, and complex backgrounds in PTC thermistors. Under conditions of scarce datasets, small target sizes, surface reflection interference, and industrial vibration, its detection accuracy, robustness, and computational efficiency still have room for improvement.

[0004] To address the problems existing in the prior art, this invention proposes a defect detection method for PTC thermistors based on YOLOv11. Summary of the Invention

[0005] This invention proposes an improved PTC thermistor defect detection method based on YOLOv11. This method can efficiently and accurately detect and extract defect features of PTC thermistors under complex background conditions, and the model exhibits good transferability and generalization. It also lowers the hardware requirements for model training and detection, making it widely applicable to industrial inspection.

[0006] The present invention adopts the following technical solution.

[0007] A defect detection method for PTC thermistors based on YOLOv11, comprising the following steps;

[0008] Step S1: Take photos of PTC thermistor images, perform data enhancement on images containing thermistor defects, and then perform various operations on the images to construct a PTC thermistor dataset.

[0009] Step S2: Construct the LMSF-YOLO defect detection network based on the YOLOv11 algorithm for PTC thermistor defect detection.

[0010] Step S3: Train the PTC thermistor defect detection algorithm using the PTC thermistor dataset, set the training rounds, learning rate and optimizer, and obtain the optimal weight file;

[0011] Step S4: Load the optimal weight file for the PTC thermistor defect detection algorithm, detect the PTC thermistor, and output the defect detection results of the PTC thermistor. The defect detection results include the location of the defect area in the image to be detected and the category corresponding to each defect area.

[0012] In step S1, firstly, PTC thermistor images are acquired using an industrial camera. The images include PTC thermistor samples with and without defects, and the acquisition process covers various imaging conditions such as different light intensities and shadow variations. Subsequently, the acquired image data is augmented using methods including image translation, horizontal flipping, vertical flipping, scaling, adding Gaussian noise, and color dithering. Then, the enhanced PTC thermistor images are annotated using the image annotation tool LAbelImg to obtain the category and bounding box information of each defect target, thereby constructing a PTC thermistor image dataset.

[0013] The LMSF-YOLO defect detection network described in step S2 is an improved YOLOv11s model. It uses its backbone network RS-HGNetV2, neck fusion network ASF-V2, and head network Detect-SEAM, and introduces the SimAM attention mechanism module to enhance attention to image detail enhancement areas.

[0014] The backbone network RS-HGNetV2 includes the HGStem preprocessing layer, the RS-HGBlock data processing module, the DWConv depthwise separable convolution, the SPPF spatial pyramid pooling layer, and the C2PSA advanced feature extraction module.

[0015] The neck fusion network ASF-V2 includes the Zoomcat multi-scale fusion module, the ScalSeq unified scale sequence module, the ASF-LearnAtt adaptive attention fusion module, and the C3k2-SimAM (CS) feature extraction module, where C3k2 is the feature extraction component of YOLOv11.

[0016] The SEAM attention unit, which is based on a feature selection mechanism constructed by depthwise separable convolution, is introduced through the head network Detect-SEAM.

[0017] To optimize the LMSF-YOLO network model, the training parameters were set as follows: the input image size was automatically adjusted to 640×640; the algorithm model was trained using a stochastic gradient descent (SGD) optimizer; the initial learning rate was set to 0.01; the weight decay was set to 0.0005; the momentum was set to 0.937; the final learning rate was set to 0.001; the batch size was 4; and the total number of iterations was set to 200. The Mosaic data augmentation method was used to obtain the optimal parameter weights of the network model and reconstruct the LMSF-YOLO network model to obtain the optimized LMSF-YOLO network model.

[0018] The HGStem initial module is a stem structure for efficient feature extraction, located at the network input, used to extract sufficient initial target features from the original input with low parameter and computational costs. The RS-HGBlock is the core module of the backbone network RS-HGNetV2, integrating a reparameterized structure (Rep), an HGNet hierarchical feature modeling structure, and a channel recalibration mechanism (SE). Reparameterization employs a multi-branch structure during training to enhance feature extraction capabilities, and during inference, it effectively merges multiple branches into a single branch to improve inference speed. SE uses global pooling to evaluate and scale the responses of each channel to enhance useful channels and suppress noisy channels, improving the discriminability of small targets and weak textures. The DWConv is a depthwise separable convolutional structure used to significantly reduce parameter and computational costs during downsampling.

[0019] SimAM is a three-dimensional, parameter-free attention module based on differences in neuron activation energy. It characterizes the importance of each feature unit by measuring the energy difference, requiring no additional convolution operations and exhibiting good structural adaptability and interpretability. To enhance the C3k2 module's ability to focus on key features, SimAM is embedded into C3k2 to form a CS feature extraction structure, which is then integrated into the ASF-V2 feature fusion network, thereby improving the model's ability to focus on and extract key information.

[0020] The ASF-LearnAtt adaptive attention fusion module is the core aggregation unit of the neck structure of the ASF-V2 feature fusion network, including channel attention (CA), local spatial attention (LA), and learnable modulation branch.

[0021] The CA module enhances the feature channels related to defects and suppresses redundant and noise channels in the channel dimension to improve the ability to express key information such as texture and edges.

[0022] The LA module weights each position of the feature map in the spatial dimension, focusing on small defects such as minor flaws and local exposure of the substrate, while reducing interference from large-area uniform ceramic substrates and complex backgrounds, which is beneficial for the detection of small targets and low-contrast defects.

[0023] Learnable modulation branches are used to assign learnable weights to features at different scales and paths, enabling the network to adaptively amplify useful scales and weaken redundant scales, avoiding information conflicts or ineffective redundancy caused by simple splicing.

[0024] The head network Detect-SEAM integrates the SEAM attention module into the detection head of the defect detection network to improve the detection capability of the defect detection network model in complex scenes such as multi-scale, low contrast and partial occlusion. By selectively enhancing and suppressing features, it strengthens effective features related to defect targets and weakens redundant information, thereby improving the capture accuracy of multi-scale targets.

[0025] The feature extraction backbone network is used to process the input PTC thermistor image, outputting four feature maps: large, medium, small, and micro, which are then input into the feature fusion neck network. The target detection head performs detection based on the fused features of the four scales (large, medium, small, and micro) output by the feature fusion neck network, and outputs the final target detection result.

[0026] Figure 7 As shown, a new Detect-SEAM detection head is constructed by introducing a SEAM attention module into the detection head. The SEAM module is based on a feature selection mechanism built on depthwise separable convolution. First, depthwise convolution processes each channel independently, preserving channel-specific information and significantly reducing computational overhead while initially learning channel importance. Subsequently, pointwise convolution reassembles features, effectively compensating for the information fragmentation between channels caused by depthwise convolution and promoting cross-channel interaction. To enhance channel correlation, the module further introduces two fully connected layers to construct dense information transmission paths, achieving deep fusion of channel dimensions, and finally forming a feature enhancement unit that combines lightweight design with strong representational capabilities.

[0027] Embedding the SEAM module into the detection head significantly enhances the model's ability to focus on target regions. In small target detection scenarios, this module effectively amplifies the feature signals of weak targets, reducing the false negative rate caused by feature blurring or background interference. Simultaneously, its suppression of background regions reduces interference from irrelevant information, thereby suppressing false detections. This design, while inheriting the efficiency of depthwise separable convolution, improves feature quality through a targeted information fusion mechanism, enabling Detect-SEAM to exhibit superior performance in multi-scale target detection tasks in complex scenes.

[0028] like Figure 8 As shown in the image, a defective PTC thermistor is displayed. The detection method described targets surface defects in the PTC thermistor, including defects such as breaks, color spots, exposed substrate, and missing print. Breaks are caused by vibrations during transport or collisions between ceramic sheets during chamfering, and are mainly distributed on the outer surface of the PTC thermistor, especially at the chamfered edges. Color spots are circular, grayish-brown spots, often caused by poor manufacturing processes. Exposed substrate occurs when the resistive printing layer of the PTC thermistor is damaged, exposing the ceramic substrate. Missing prints are caused by poor printing.

[0029] like Figure 9 As shown, after data augmentation, due to the scarcity of defective samples, various operations are performed on the image, such as brightness adjustment, flipping, Gaussian noise, contrast enhancement, and rotation, thereby effectively expanding the dataset.

[0030] In the detection method described above, firstly, HGStem performs two convolutions on the input image to obtain a 1 / 4 downsampled P2 feature map. Then, it uses RS-HGBlock (containing two 3×3 convolution layers, residual connections, and reparameterized branches) to deeply extract low-level semantic features at the P2 level. The P2 features are downsampled to 1 / 8 by DWConv and input into RS-HGBlock to obtain mid-to-low-level semantics at the P3 level. At P3, DWConv is used again for downsampling, and three sets of RS-HGBlocks are alternately stacked on the lightweight and normal branches to enhance mid-to-high-level semantic expression. Subsequently, the P5 feature map is downsampled by DWConv, and after channel expansion by one set of RS-HGBlocks, it is connected to the SPPF three-scale pooling concatenation module to extract multi-scale contextual information. Finally, the C2PSA channel and spatial attention module highlight key response regions, thereby outputting multi-scale features.

[0031] The ASF-V2 feature fusion network first performs channel dimensionality reduction on the P4 and P5, and P3 and P4 feature maps respectively using 1×1 convolution. After concatenation, the fused P3 and P5 feature maps are obtained through two CS modules. Then, the fused P3 features are downsampled to the P4 scale and concatenated with the original P4 feature map, and feature extraction is performed through two CS layers. At the same time, the fused P5 features are downsampled to the P4 scale to obtain the fused P4 feature map. The three fused features of P3, P4, and P5 are input into the ScalSeq module for adaptive weighted combination, and ASF attention is applied to suppress redundant information and highlight key response regions. Then, the multi-directional fusion result is upsampled to the P2 scale, concatenated with the original P2 feature map, and after extracting detailed features through two CS layers, it is input into the detection head along with the fused features of the three scales of P3, P4, and P5.

[0032] like Figure 5 As shown, in the SimAM module, H, W, and C represent the height, width, and number of channels of the feature map, respectively; as Figure 4 As shown, the SimAM module dynamically adjusts the weights of the energy function of each neuron to adapt to the feature representations of different channels and spatial locations;

[0033] In visual neuroscience, neurons carrying the most information signals often exhibit firing patterns different from their neighbors; furthermore, highly active neurons can suppress responses to their surroundings, a phenomenon known as spatial inhibition. Therefore, assigning higher weights to neurons exhibiting spatial inhibition allows the model to emphasize subtle-scale cues, reduce background noise in small object detection, and enhance attention to salient information.

[0034] The minimum energy function is obtained by measuring the linear separability between the target neuron and other neurons. The formula for its calculation is as follows:

[0035] Formula 1;

[0036] In the formula, t represents the target neuron. is the regularization coefficient, and M is the number of neurons in each channel. In Equation 3, for neurons t with lower energy, the difference between them and the surrounding active neurons is greater, and they should be given higher weights to highlight their importance. Finally, the SimAM module performs feature enhancement processing through sigmoid, as shown in Equation 2 below.

[0037] Formula 2;

[0038] In the formula, E is the set of minimum energy functions of neurons; X is the input feature; Multiplying matrix elements.

[0039] like Figure 6As shown, the ASF-V2 feature fusion network first performs channel adjustment and preliminary processing on the multi-scale features P2, P3, P4, and P5 from the backbone. The high-scale features P4 and P3 output from the backbone are adjusted for channel count using 1x1 convolutions to match the other features to be fused. Upsampling and downsampling are used to align features of different resolutions, providing inputs with the same spatial size for subsequent fusion and ensuring that features of different scales have the basic conditions for fusion. Next, the multi-scale fusion module Zoomcat performs preliminary integration of the three input features at different scales, aligning and concatenating features of different scales, with the height / width of the mid-scale features as the target size. Large-scale features are compressed to the target size using max pooling and average pooling to enhance key region information. Small-scale features are enlarged to the target size using nearest neighbor interpolation to retain detailed information. Channel dimensions are concatenated, and the three features are fused to integrate multi-scale information.

[0040] After initial fusion by Zoomcat, the features are further unified into a scale sequence through the ScalSeq module. This module converts features from three scales into a feature sequence with a unified channel and scale, and models cross-scale correlation through 3D convolution.

[0041] Let the Zoomcat multi-scale fusion module take three features at different scales as input. , , These represent large-scale, medium-scale, and small-scale features, respectively; large-scale features Adjustments were made to compress the data to a mesoscale, and the results were summed to retain more information; the relevant formulas are as follows:

[0042] Formula 3;

[0043] in, This indicates that feature F is adaptively pooled to... size, Indicates adaptive average pooling. express The result after compression;

[0044] For small-scale features (High resolution) Upsampled to a mesoscale size via nearest neighbor interpolation:

[0045] Formula 4;

[0046] in, This indicates that feature F is upsampled to size;

[0047] The adjusted large-scale, medium-scale, and small-scale features are concatenated along the channel dimension, using the following formula:

[0048] Formula 5.

[0049] The ScalSeq module takes the input features P3, P4, and P5 at three consecutive scales; uses a 1x1 convolution to unify the number of channels of the three features to the target number of channels C; and upsamples the P4 and P5 features to the spatial size of P3; the formula is as follows:

[0050] Formula 6;

[0051] in: This indicates that the number of channels for a 1x1 convolution is adjusted to C;

[0052] Upsample the features of P4 and P5 to the spatial size of P3;

[0053] Formula 7;

[0054] in, For a dynamic sampling function, the input features F, target size (H, W), and scaling factor s, the output is... , and The dimensions are consistent.

[0055] Will , and The features are expanded into 3D features, and three features are concatenated along the scale dimension to form a 3D feature volume. Finally, 3D convolution is used to capture inter-scale correlations, and the scale dimension is compressed through batch normalization, activation, and 3D pooling to obtain the final cross-scale fused feature output. ;

[0056] Formula 8;

[0057] in: This indicates a compression scale dimension operation;

[0058] Finally, the ASF-LearnAtt feature aggregation module dynamically adjusts feature weights through channel attention (CA) and local spatial attention (LA) to highlight defect areas, including defects and color spots in PTC thermistors, while suppressing background interference. First, for the two input features to be fused—one a basic feature and the other a cross-scale fusion feature—channel attention is used to enhance defect-related channels, generating channel weights and suppressing irrelevant channels. Then, feature fusion is performed, adding the two features together. Next, local spatial attention is used to enhance the spatial response of the defect area, generating spatial weights to highlight the defect location. The process is as follows:

[0059] Basic features of input features Applying channel attention:

[0060] Formula 9;

[0061] Features with cross-scale fusion After addition, multiplying by a small per-channel scale allows for learnable modulation, enabling small-scale learnable amplification or suppression to enhance expressive power. Finally, local attention is applied.

[0062] Formula 10;

[0063] in, This is the final output feature map of the local spatial attention module. The intermediate feature map obtained by adding the base features enhanced by channel attention and the cross-scale features. This represents a 1x1 convolution operation used to learn the height-oriented attention weights. This represents a 1x1 convolution operation used to learn the attention weights in the width direction. This represents the Sigmoid activation function. This indicates element-wise multiplication.

[0064] The specific method in step S3 is as follows: Set the training parameters, adaptively adjust the input image size to 640×640, train the LMSF-YOLO model using the stochastic gradient descent (SGD) optimizer, with an initial learning rate of 0.01, a final learning rate of 0.001, a weight decay of 0.0005, a momentum of 0.937, a batch size of 4, and a total of 200 iterations; during the training process, use Mosaic data augmentation to improve generalization ability, obtain the optimal parameter weights and reconstruct the LMSF-YOLO network model, and obtain the optimized LMSF-YOLO network by comprehensively evaluating the indices of mAP, mAP@50:95, and loss function.

[0065] The specific method in step S4 is as follows: load the trained weight file into the network model of the defect detection network LMSF-YOLO; input the PTC thermistor images in the test set into the LMSF-YOLO model with loaded weights to perform defect detection, and obtain the center coordinates, width and height of the bounding box of each defect target and the corresponding confidence score; output the detection results in a predetermined format and save them as a txt tag file.

[0066] This invention proposes a defect detection method for PTC thermistors based on an improved YOLOv11. It collects and constructs a dataset of PTC surface images containing defects such as defects, color spots, exposed substrate, and missing prints. A lightweight detection network, LMSF-YOLO, is designed based on the YOLO series. This network includes an improved backbone RS-HGNetV2, an ASF-V2 neck structure with an added P2 feature layer and a SimAM attention mechanism, and a lightweight SEAM detection head. LMSF-YOLO is trained using the dataset to obtain optimal weights. The image to be inspected is input into the trained model, which outputs the category, location, and confidence level of various defects. This invention maintains a lightweight model while achieving high detection accuracy and robustness, making it suitable for online industrial inspection and embedded deployment of PTC thermistors.

[0067] Compared with the prior art, the present invention has the following beneficial effects:

[0068] ① In this invention, an efficient backbone network RS-HGNetV2 is designed, which integrates hierarchical gradient feature extraction, reparameterized convolution technology and channel recalibration, thereby enhancing feature extraction capabilities while reducing model complexity.

[0069] ② This invention proposes a novel neck network ASF-V2 attention-scale sequence fusion network, which creates an additional cross-scale feature fusion path and introduces the SimAM module, adding a P2 micro-target detection layer to achieve efficient fusion of feature information at different scales.

[0070] ③ This invention introduces a SEAM spatial enhanced attention module into the detection head, constructing a new Detect-SEAM detection head, reducing model parameters and improving network detection performance.

[0071] ④ This invention can collect data according to the needs of enterprises, construct a dedicated dataset of PTC thermistor surface defects, and use data augmentation methods to expand the sample. Attached Figure Description

[0072] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments:

[0073] Appendix Figure 1 This is a flowchart illustrating an embodiment of the present invention;

[0074] Appendix Figure 2 This is a schematic diagram of the LMSF-YOLO model according to an embodiment of the present invention;

[0075] Appendix Figure 3 This is a schematic diagram of the RS-HGNetV2 backbone network structure according to an embodiment of the present invention (where (a) is a schematic diagram of the HGStem module structure, and (b) is a schematic diagram of the RS-HGBlock module structure).

[0076] Appendix Figure 4 This is a schematic diagram of the RepConv module structure according to an embodiment of the present invention;

[0077] Appendix Figure 5 This is a schematic diagram of the SimAM module structure according to an embodiment of the present invention;

[0078] Appendix Figure 6 This is a schematic diagram of the Zoomcat, ScalSeq, and ASF-LearnAtt module structures in an embodiment of the present invention;

[0079] Appendix Figure 7 This is a schematic diagram of the SEAM module structure according to an embodiment of the present invention;

[0080] Appendix Figure 8 The following is a schematic diagram illustrating four categories in the dataset of this invention (where (a) represents defects, (b) represents colored dots, (c) represents exposed background, and (d) represents missing prints).

[0081] Appendix Figure 9 This is a schematic diagram illustrating the data enhancement effect of an embodiment of the present invention (where (a) is the original image, (b) is rotation, (c) is flipping, (d) is scaling, (e) is color jitter, (c) is Gaussian noise, and (c) is rotation + color jitter).

[0082] Appendix Figure 10 This is a schematic diagram of a comparative experiment of different algorithms in an embodiment of the present invention;

[0083] Appendix Figure 11 This is a schematic diagram of a comparative experiment between LMSF-YOLO and baseline YOLOv11s in an embodiment of the present invention (where (a) is a training process diagram, and (b) is a Grad-CAM heatmap analysis).

[0084] Appendix Figure 12 This is a schematic diagram comparing the detection images of different algorithms in an embodiment of the present invention. Detailed Implementation

[0085] As shown in the figure, a defect detection method for PTC thermistors based on YOLOv11 includes the following steps;

[0086] Step S1: Take photos of PTC thermistor images, perform data enhancement on images containing thermistor defects, and then perform various operations on the images to construct a PTC thermistor dataset.

[0087] Step S2: Construct the LMSF-YOLO defect detection network based on the YOLOv11 algorithm for PTC thermistor defect detection.

[0088] Step S3: Train the PTC thermistor defect detection algorithm using the PTC thermistor dataset, set the training rounds, learning rate and optimizer, and obtain the optimal weight file;

[0089] Step S4: Load the optimal weight file for the PTC thermistor defect detection algorithm, detect the PTC thermistor, and output the defect detection results of the PTC thermistor. The defect detection results include the location of the defect area in the image to be detected and the category corresponding to each defect area.

[0090] Please refer to Figure 2 In this embodiment, the LMSF-YOLO defect detection network includes a backbone network, a feature fusion network, and a detection head; the backbone network is RS-HGNetV2, the feature fusion network is ASF-V2 feature fusion network, the feature fusion network includes a SimAM module, and the detection head is a Detect-SEAM detection head.

[0091] This embodiment employs the efficient backbone network RS-HGNetV2 in the backbone network, significantly reducing the number of model parameters and computational complexity while improving detection accuracy, which is beneficial for real-time detection in industrial deployments. In the feature fusion network, an ASF-V2 neck network is proposed, creating an additional cross-scale feature fusion path and introducing the ASF-LearnAtt module and the parameterless attention mechanism SimAM. This dynamically learns the weights of different feature layers, enhancing the contribution of key features, suppressing redundant information, and improving fusion efficiency. A dedicated P2 small target detection layer is added to improve the recognition accuracy of small defects and optimize the feature pyramid structure. To reduce model parameters and improve network detection performance, a SEAM attention unit is introduced into the detection head, constructing a new Detect-SEAM detection head to improve localization accuracy in small targets, partially occluded environments, and noisy environments. The LMSF-YOLO algorithm uses the efficient backbone network RS-HGNetV2 and introduces the ASF-V2 neck network, designing a new Detect-SEAM detection head to improve feature extraction and feature fusion effects while significantly reducing the number of parameters.

[0092] Please refer to Figure 3The high-efficiency backbone network RS-HGNetV2 first performs two 3×3 convolutions on the input image using HGStem to generate a 1 / 4 downsampled feature map. Then, it uses six RS-HGBlock modules, each containing two 3×3 convolutions, residual connections, and a reparameterized branch, to extract low-level semantics at the P2 level. The P2 feature map is downsampled to 1 / 8 using DWConv, and then six RS-HGBlocks are used to extract mid-to-low-level semantics at the P3 level. At P3, DWConv is performed again, with three sets of RS-HGBlocks alternately stacked on the lightweight and normal branches, resulting in 18 3×3 convolutions to enhance mid-to-high-level semantics. After further downsampling to P5 using DWConv, one set of RS-HGBlocks is used to expand the channels, followed by SPPF three-scale pooling to extract multi-scale context. Key response regions are highlighted through C2PSA channels and spatial attention; finally, multi-scale features are output.

[0093] Please refer to Figure 4 RepConv, a reparameterized convolution, learns and mines diverse representations of data in parallel through 3×3 convolutions, 1×1 convolutions and BN branches during training. During inference, it is integrated into a single 3×3 convolution kernel with the same number of parameters as the standard convolution. It inherits the advantages of multi-branch fusion. After incorporating HGBlock, training can enhance the extraction of fine-grained features of complex targets. During inference, the single branch does not increase the amount of computation and retains the multi-scale fusion capability and network efficiency.

[0094] In step S1, firstly, PTC thermistor images are acquired using an industrial camera. The images include PTC thermistor samples with and without defects, and the acquisition process covers various imaging conditions such as different light intensities and shadow variations. Subsequently, the acquired image data is augmented using methods including image translation, horizontal flipping, vertical flipping, scaling, adding Gaussian noise, and color dithering. Then, the enhanced PTC thermistor images are annotated using the image annotation tool LAbelImg to obtain the category and bounding box information of each defect target, thereby constructing a PTC thermistor image dataset.

[0095] In this embodiment, an industrial camera is used to acquire images of PTC thermistors. The acquired images include both defective and defect-free PTC thermistor images. The acquisition process includes various conditions, such as changes in light intensity and shadow. Data enhancement is performed on the image data, including image translation, horizontal flipping, vertical flipping, scaling, adding Gaussian noise, and color dithering.

[0096] The LMSF-YOLO defect detection network described in step S2 is an improved YOLOv11s model. It uses its backbone network RS-HGNetV2, neck fusion network ASF-V2, and head network Detect-SEAM, and introduces the SimAM attention mechanism module to enhance attention to image detail enhancement areas.

[0097] The backbone network RS-HGNetV2 includes the HGStem preprocessing layer, the RS-HGBlock data processing module, the DWConv depthwise separable convolution, the SPPF spatial pyramid pooling layer, and the C2PSA advanced feature extraction module.

[0098] The neck fusion network ASF-V2 includes the Zoomcat multi-scale fusion module, the ScalSeq unified scale sequence module, the ASF-LearnAtt adaptive attention fusion module, and the C3k2-SimAM (CS) feature extraction module, where C3k2 is the feature extraction component of YOLOv11.

[0099] The SEAM attention unit, which is based on a feature selection mechanism constructed by depthwise separable convolution, is introduced through the head network Detect-SEAM.

[0100] To optimize the LMSF-YOLO network model, the training parameters were set as follows: the input image size was automatically adjusted to 640×640; the algorithm model was trained using a stochastic gradient descent (SGD) optimizer; the initial learning rate was set to 0.01; the weight decay was set to 0.0005; the momentum was set to 0.937; the final learning rate was set to 0.001; the batch size was 4; and the total number of iterations was set to 200. The Mosaic data augmentation method was used to obtain the optimal parameter weights of the network model and reconstruct the LMSF-YOLO network model to obtain the optimized LMSF-YOLO network model.

[0101] The HGStem initial module is a stem structure for efficient feature extraction, located at the network input, used to extract sufficient initial target features from the original input with low parameter and computational costs. The RS-HGBlock is the core module of the backbone network RS-HGNetV2, integrating a reparameterized structure (Rep), an HGNet hierarchical feature modeling structure, and a channel recalibration mechanism (SE). Reparameterization employs a multi-branch structure during training to enhance feature extraction capabilities, and during inference, it effectively merges multiple branches into a single branch to improve inference speed. SE uses global pooling to evaluate and scale the responses of each channel to enhance useful channels and suppress noisy channels, improving the discriminability of small targets and weak textures. The DWConv is a depthwise separable convolutional structure used to significantly reduce parameter and computational costs during downsampling.

[0102] SimAM is a three-dimensional, parameter-free attention module based on differences in neuron activation energy. It characterizes the importance of each feature unit by measuring the energy difference, requiring no additional convolution operations and exhibiting good structural adaptability and interpretability. To enhance the C3k2 module's ability to focus on key features, SimAM is embedded into C3k2 to form a CS feature extraction structure, which is then integrated into the ASF-V2 feature fusion network, thereby improving the model's ability to focus on and extract key information.

[0103] The ASF-LearnAtt adaptive attention fusion module is the core aggregation unit of the neck structure of the ASF-V2 feature fusion network, including channel attention (CA), local spatial attention (LA), and learnable modulation branch.

[0104] The CA module enhances the feature channels related to defects and suppresses redundant and noise channels in the channel dimension to improve the ability to express key information such as texture and edges.

[0105] The LA module weights each position of the feature map in the spatial dimension, focusing on small defects such as minor flaws and local exposure of the substrate, while reducing interference from large-area uniform ceramic substrates and complex backgrounds, which is beneficial for the detection of small targets and low-contrast defects.

[0106] Learnable modulation branches are used to assign learnable weights to features at different scales and paths, enabling the network to adaptively amplify useful scales and weaken redundant scales, avoiding information conflicts or ineffective redundancy caused by simple splicing.

[0107] The head network Detect-SEAM integrates the SEAM attention module into the detection head of the defect detection network to improve the detection capability of the defect detection network model in complex scenes such as multi-scale, low contrast and partial occlusion. By selectively enhancing and suppressing features, it strengthens effective features related to defect targets and weakens redundant information, thereby improving the capture accuracy of multi-scale targets.

[0108] The feature extraction backbone network is used to process the input PTC thermistor image, outputting four feature maps: large, medium, small, and micro, which are then input into the feature fusion neck network. The target detection head performs detection based on the fused features of the four scales (large, medium, small, and micro) output by the feature fusion neck network, and outputs the final target detection result.

[0109] Figure 7As shown, a new Detect-SEAM detection head is constructed by introducing a SEAM attention module into the detection head. The SEAM module is based on a feature selection mechanism built on depthwise separable convolution. First, depthwise convolution processes each channel independently, preserving channel-specific information and significantly reducing computational overhead while initially learning channel importance. Subsequently, pointwise convolution reassembles features, effectively compensating for the information fragmentation between channels caused by depthwise convolution and promoting cross-channel interaction. To enhance channel correlation, the module further introduces two fully connected layers to construct dense information transmission paths, achieving deep fusion of channel dimensions, and finally forming a feature enhancement unit that combines lightweight design with strong representational capabilities.

[0110] Embedding the SEAM module into the detection head significantly enhances the model's ability to focus on target regions. In small target detection scenarios, this module effectively amplifies the feature signals of weak targets, reducing the false negative rate caused by feature blurring or background interference. Simultaneously, its suppression of background regions reduces interference from irrelevant information, thereby suppressing false detections. This design, while inheriting the efficiency of depthwise separable convolution, improves feature quality through a targeted information fusion mechanism, enabling Detect-SEAM to exhibit superior performance in multi-scale target detection tasks in complex scenes.

[0111] like Figure 8 As shown in the image, a defective PTC thermistor is displayed. The detection method described targets surface defects in the PTC thermistor, including defects such as breaks, color spots, exposed substrate, and missing print. Breaks are caused by vibrations during transport or collisions between ceramic sheets during chamfering, and are mainly distributed on the outer surface of the PTC thermistor, especially at the chamfered edges. Color spots are circular, grayish-brown spots, often caused by poor manufacturing processes. Exposed substrate occurs when the resistive printing layer of the PTC thermistor is damaged, exposing the ceramic substrate. Missing prints are caused by poor printing.

[0112] like Figure 9 As shown, after data augmentation, due to the scarcity of defective samples, various operations are performed on the image, such as brightness adjustment, flipping, Gaussian noise, contrast enhancement, and rotation, thereby effectively expanding the dataset.

[0113] In the detection method described above, firstly, HGStem performs two convolutions on the input image to obtain a 1 / 4 downsampled P2 feature map. Then, it uses RS-HGBlock (containing two 3×3 convolution layers, residual connections, and reparameterized branches) to deeply extract low-level semantic features at the P2 level. The P2 features are downsampled to 1 / 8 by DWConv and input into RS-HGBlock to obtain mid-to-low-level semantics at the P3 level. At P3, DWConv is used again for downsampling, and three sets of RS-HGBlocks are alternately stacked on the lightweight and normal branches to enhance mid-to-high-level semantic expression. Subsequently, the P5 feature map is downsampled by DWConv, and after channel expansion by one set of RS-HGBlocks, it is connected to the SPPF three-scale pooling concatenation module to extract multi-scale contextual information. Finally, the C2PSA channel and spatial attention module highlight key response regions, thereby outputting multi-scale features.

[0114] The ASF-V2 feature fusion network first performs channel dimensionality reduction on the P4 and P5, and P3 and P4 feature maps respectively using 1×1 convolution. After concatenation, the fused P3 and P5 feature maps are obtained through two CS modules. Then, the fused P3 features are downsampled to the P4 scale and concatenated with the original P4 feature map, and feature extraction is performed through two CS layers. At the same time, the fused P5 features are downsampled to the P4 scale to obtain the fused P4 feature map. The three fused features of P3, P4, and P5 are input into the ScalSeq module for adaptive weighted combination, and ASF attention is applied to suppress redundant information and highlight key response regions. Then, the multi-directional fusion result is upsampled to the P2 scale, concatenated with the original P2 feature map, and after extracting detailed features through two CS layers, it is input into the detection head along with the fused features of the three scales of P3, P4, and P5.

[0115] like Figure 5 As shown, in the SimAM module, H, W, and C represent the height, width, and number of channels of the feature map, respectively; as Figure 4 As shown, the SimAM module dynamically adjusts the weights of the energy function of each neuron to adapt to the feature representations of different channels and spatial locations;

[0116] In visual neuroscience, neurons carrying the most information signals often exhibit firing patterns different from their neighbors; furthermore, highly active neurons can suppress responses to their surroundings, a phenomenon known as spatial inhibition. Therefore, assigning higher weights to neurons exhibiting spatial inhibition allows the model to emphasize subtle-scale cues, reduce background noise in small object detection, and enhance attention to salient information.

[0117] The minimum energy function is obtained by measuring the linear separability between the target neuron and other neurons. The formula for its calculation is as follows:

[0118] Formula 1;

[0119] In the formula, t represents the target neuron. is the regularization coefficient, and M is the number of neurons in each channel. In Equation 3, for neurons t with lower energy, the difference between them and the surrounding active neurons is greater, and they should be given higher weights to highlight their importance. Finally, the SimAM module performs feature enhancement processing through sigmoid, as shown in Equation 2 below.

[0120] Formula 2;

[0121] In the formula, E is the set of minimum energy functions of neurons; X is the input feature; Multiplying matrix elements.

[0122] like Figure 6 As shown, the ASF-V2 feature fusion network first performs channel adjustment and preliminary processing on the multi-scale features P2, P3, P4, and P5 from the backbone. The high-scale features P4 and P3 output from the backbone are adjusted for channel count using 1x1 convolutions to match the other features to be fused. Upsampling and downsampling are used to align features of different resolutions, providing inputs with the same spatial size for subsequent fusion and ensuring that features of different scales have the basic conditions for fusion. Next, the multi-scale fusion module Zoomcat performs preliminary integration of the three input features at different scales, aligning and concatenating features of different scales, with the height / width of the mid-scale features as the target size. Large-scale features are compressed to the target size using max pooling and average pooling to enhance key region information. Small-scale features are enlarged to the target size using nearest neighbor interpolation to retain detailed information. Channel dimensions are concatenated, and the three features are fused to integrate multi-scale information.

[0123] After initial fusion by Zoomcat, the features are further unified into a scale sequence through the ScalSeq module. This module converts features from three scales into a feature sequence with a unified channel and scale, and models cross-scale correlation through 3D convolution.

[0124] Let the Zoomcat multi-scale fusion module take three features at different scales as input. , , These represent large-scale, medium-scale, and small-scale features, respectively; large-scale features Adjustments were made to compress the data to a mesoscale, and the results were summed to retain more information; the relevant formulas are as follows:

[0125] Formula 3;

[0126] in, This indicates that feature F is adaptively pooled to... size, Indicates adaptive average pooling. express The result after compression;

[0127] For small-scale features (High resolution) Upsampled to a mesoscale size via nearest neighbor interpolation:

[0128] Formula 4;

[0129] in, This indicates that feature F is upsampled to size;

[0130] The adjusted large-scale, medium-scale, and small-scale features are concatenated along the channel dimension, using the following formula:

[0131] Formula 5.

[0132] The ScalSeq module takes the input features P3, P4, and P5 at three consecutive scales; uses a 1x1 convolution to unify the number of channels of the three features to the target number of channels C; and upsamples the P4 and P5 features to the spatial size of P3; the formula is as follows:

[0133] Formula 6;

[0134] in: This indicates that the number of channels for a 1x1 convolution is adjusted to C;

[0135] Upsample the features of P4 and P5 to the spatial size of P3;

[0136] Formula 7;

[0137] in, For a dynamic sampling function, the input features F, target size (H, W), and scaling factor s, the output is... , and The dimensions are consistent.

[0138] Will , and The features are expanded into 3D features, and three features are concatenated along the scale dimension to form a 3D feature volume. Finally, 3D convolution is used to capture inter-scale correlations, and the scale dimension is compressed through batch normalization, activation, and 3D pooling to obtain the final cross-scale fused feature output. ;

[0139] Formula 8;

[0140] in: This indicates a compression scale dimension operation;

[0141] Finally, the ASF-LearnAtt feature aggregation module dynamically adjusts feature weights through channel attention (CA) and local spatial attention (LA) to highlight defect areas, including defects and color spots in PTC thermistors, while suppressing background interference. First, for the two input features to be fused—one a basic feature and the other a cross-scale fusion feature—channel attention is used to enhance defect-related channels, generating channel weights and suppressing irrelevant channels. Then, feature fusion is performed, adding the two features together. Next, local spatial attention is used to enhance the spatial response of the defect area, generating spatial weights to highlight the defect location. The process is as follows:

[0142] Basic features of input features Applying channel attention:

[0143] Formula 9;

[0144] Features with cross-scale fusion After addition, multiplying by a small per-channel scale allows for learnable modulation, enabling small-scale learnable amplification or suppression to enhance expressive power. Finally, local attention is applied.

[0145] Formula 10;

[0146] in, This is the final output feature map of the local spatial attention module. The intermediate feature map obtained by adding the base features enhanced by channel attention and the cross-scale features. This represents a 1x1 convolution operation used to learn the height-oriented attention weights. This represents a 1x1 convolution operation used to learn the attention weights in the width direction. This represents the Sigmoid activation function. This indicates element-wise multiplication.

[0147] The specific method in step S3 is as follows: Set the training parameters, adaptively adjust the input image size to 640×640, train the LMSF-YOLO model using the stochastic gradient descent (SGD) optimizer, with an initial learning rate of 0.01, a final learning rate of 0.001, a weight decay of 0.0005, a momentum of 0.937, a batch size of 4, and a total of 200 iterations; during the training process, use Mosaic data augmentation to improve generalization ability, obtain the optimal parameter weights and reconstruct the LMSF-YOLO network model, and obtain the optimized LMSF-YOLO network by comprehensively evaluating the indices of mAP, mAP@50:95, and loss function.

[0148] In this embodiment, to optimize the LMSF-YOLO network model, the training parameters are set as follows: the input image size is automatically adjusted to 640×640, the algorithm model is trained using a stochastic gradient descent (SGD) optimizer, the initial learning rate is set to 0.01, the weight decay is set to 0.0005, the momentum is set to 0.937, the final learning rate is set to 0.001, the batch size is 4, the total number of iterations is set to 200, and the Mosaic data augmentation method is used to obtain the optimal parameter weights of the network model and reconstruct the LMSF-YOLO network model to obtain the optimized LMSF-YOLO network model.

[0149] The specific method in step S4 is as follows: load the trained weight file into the network model of the defect detection network LMSF-YOLO; input the PTC thermistor images in the test set into the LMSF-YOLO model with loaded weights to perform defect detection, and obtain the center coordinates, width and height of the bounding box of each defect target and the corresponding confidence score; output the detection results in a predetermined format and save them as a txt tag file.

[0150] Includes the following steps:

[0151] S41: First, save the trained weight file to the LMSF-YOLO network model.

[0152] S42: Input the test set into the trained LMSF-YOLO network, load the trained LMSF-YOLO weight file, perform defect detection on the PTC thermistor images in the test set, and obtain the center point position and width and height of the defect bounding box, as well as the confidence level of the defect.

[0153] S43: Output the defect detection results and save them as a txt tag file.

[0154] The following comparative experiments with other YOLO series algorithms verify the detection performance of the LMSF-YOLO algorithm in the field of PTC thermistor defect detection. The experimental environment was built on Ubuntu 22.04 operating system, programmed using Python 3.10, and using PyTorch 2.7.1 and CUDA 12.6 to build the deep learning framework. Regarding hardware, the graphics card was an NVIDIA RTX4060 (7.8GB VRAM). In this experiment, the input image size was automatically adjusted to 640×640. The algorithm model was trained using a stochastic gradient descent (SGD) optimizer, with an initial learning rate of 0.01, weight decay of 0.001, momentum of 0.937, and a final learning rate of 0.001. The batch size was 4. During training, the total number of iterations was set to 200. Mosaic data augmentation was used during training, and the optimal training weights were selected to calculate the final detection accuracy.

[0155] To comprehensively evaluate the model's performance, the LMSF-YOLO algorithm proposed in this invention selects precision, recall, average precision (AP), average AP (mAP), billion floating-point operations per second (GFLOPS), and number of parameters as key evaluation metrics. Params reflect model size, while GFLOPS reflect computational complexity. Lower numbers of params and GFLOPS indicate lower hardware resource requirements and easier deployment of the detection algorithm.

[0156] To verify the effectiveness of the proposed model, this paper selected several mainstream object detection models for comparative experiments, including the YOLO series (v5s, v8s, v9s, v10s, v11s, v12s). In the experiments, the performance of each model on the dataset was evaluated, including metrics such as P, R, mAP50, mAP50:95, computational cost, and parameter count. These results clearly demonstrate that LMSF-YOLO outperforms existing mainstream models in several key performance indicators.

[0157] like Figure 10 As shown, in order to more intuitively display the metrics of each model and highlight the superiority of LMSF-YOLO, a bar chart is used to show the results of the comparative experiments of each model, further verifying the advantages of the proposed model in terms of accuracy and efficiency.

[0158]

[0159] This embodiment conducted comparative experiments on a self-built dataset. The experimental results are shown in Table 1. LMSF-YOLO performed excellently across all metrics, with P and R rates reaching 92.6% and 96.0% respectively, and mAP50 and mAP50:95 reaching 95.7% and 65.0% respectively, significantly outperforming other models. Furthermore, LMSF-YOLO has only 7.9M parameters, effectively reducing computational complexity while improving performance. Compared to the latest models such as YOLOv11s and YOLOv12s, it further demonstrates its efficiency and robustness. Compared to mainstream target detection models, LMSF-YOLO achieves the optimal balance between performance and resource consumption in PTC thermistor defect detection tasks, proving the effectiveness of its improvements and demonstrating high engineering practicality.

[0160] like Figure 11 As shown in the figure, the training process diagrams and heatmaps of LMSF-YOLO and the baseline model YOLOv11s are displayed respectively. By comparing and analyzing the two, the effectiveness of the improved algorithm LMSF-YOLO of this invention can be more comprehensively evaluated. Figure 11 As clearly shown in (a), with the increase of training iterations, precision, recall, and MAP change accordingly, and the curve of LMSF-YOLO is consistently higher than that of YOLOv11, highlighting the performance gains of the proposed model in the detection task. To more comprehensively evaluate the effectiveness of the LMSF-YOLO algorithm, we conducted a heatmap-based visualization comparison analysis of the LMSF-YOLO algorithm and the baseline YOLOv11s model. Gradient-weighted class activation mapping was used to visualize the output layers of different models, and the gradients of specific layers of each model were calculated to reveal the impact of different models on the detection decisions of image regions. In the heatmap, darker colors indicate a greater contribution to the final prediction result; red areas represent the parts the model focuses on most, while blue areas have a smaller role in prediction. The more concentrated the red areas are within the target box, the more the model can focus on defects. Experimental results are as follows: Figure 11 As shown in (b), the bounding boxes indicate the key regions used for defect detection. Visually, LMSF-YOLO demonstrates superior localization capabilities: the activation in defect regions is more concentrated and stronger (red areas), indicating improved localization accuracy and enhanced robustness. Higher activation intensity corresponds to stronger model attention.

[0161] like Figure 12As shown in the figure, the performance of LMSF-YOLO and mainstream algorithm models in PTC thermistor defect detection is illustrated, with colored bounding boxes representing predicted categories and confidence scores. A comparison of LMSF-YOLO with mainstream algorithms reveals that YOLOv11 fails to detect instances of missing defect categories, while LMSF-YOLO not only identifies these instances but also locates them more precisely on the PTC thermistor surface. Among the four defect categories, LMSF-YOLO performs best: for missing defects, it detects all targets with the highest confidence. For color spots and exposed substrate defects, only YOLOv5s identifies instances in the upper left corner, while LMSF-YOLO achieves a higher confidence score. YOLOv11s produces false positives by misclassifying color spots as exposed substrate, while LMSF-YOLO correctly identifies all relevant areas. Overall, LMSF-YOLO provides reliable and high-precision detection for a wide range of PTC thermistor surface defects.

[0162] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention to other forms. Those skilled in the art can utilize the disclosed technical content to make changes or modifications to obtain equivalent embodiments. Any simple modifications, equivalent changes, and alterations made to the above embodiments without departing from the technical solution of the present invention and based on the technical essence of the present invention shall still fall within the protection scope of the present invention.

Claims

1. A PTC thermistor defect detection method based on YOLOv11 improvement, characterized by: The method comprises the following steps: Step S1, image data of PTC thermistor is acquired by shooting, data enhancement is performed on the image with defects, and various operations are performed on the image to construct a PTC thermistor data set; Step S2, a defect detection network LMSF-YOLO for a PTC thermistor defect detection algorithm is constructed based on a YOLOv11 algorithm; Step S3, the PTC thermistor defect detection algorithm is trained by using the PTC thermistor data set, the training round, the learning rate and the optimizer are set, and the best weight file is obtained; Step S4, the best weight file is loaded for the PTC thermistor defect detection algorithm, the PTC thermistor is detected, and the defect detection result of the PTC thermistor is output, wherein the defect detection result includes the position of the defect area in the image to be detected and the corresponding category of each defect area.

2. The PTC thermistor defect detection method based on the improved YOLOv11 according to claim 1, characterized in that: In step S1, first, an industrial camera is used to collect PTC thermistor images, the images include PTC thermistor samples with defects and without defects, and the collection process covers various imaging conditions such as different light intensities and shadow changes; then the collected image data is subjected to data enhancement, the data enhancement methods include image translation, horizontal flipping, vertical flipping, scaling, adding Gaussian noise and color jittering; then the enhanced PTC thermistor images are labeled by using an image labeling tool LAbelImg, the category and boundary box information of each defect target are obtained, and thus a PTC thermistor image data set is constructed.

3. The PTC thermistor defect detection method based on the improved YOLOv11 according to claim 1, characterized in that: The LMSF-YOLO defect detection network in step S2 is an improved YOLOv11s model, which comprises a backbone network RS-HGNetV2, a neck fusion network ASF-V2, a head network Detect-SEAM, and a SimAM attention mechanism module is introduced to strengthen the attention to the enhanced image detail area; The backbone network RS-HGNetV2 comprises an HGStem preprocessing layer, an RS-HGBlock data processing module, a DWConv depth separable convolution, an SPPF spatial pyramid pooling layer and a C2PSA advanced feature extraction module; The neck fusion network ASF-V2 comprises a Zoomcat multi-scale fusion module, a ScalSeq unified scale sequence module, an ASF-LearnAtt adaptive attention fusion module and a C3k2-SimAM feature extraction module, wherein C3k2 is a feature extraction component of YOLOv11; The SEAM attention unit based on the depth separable convolution is introduced through the head network Detect-SEAM.

4. The PTC thermistor defect detection method based on the improved YOLOv11 according to claim 3, characterized in that: The HGStem initial module is a stem structure for efficient feature extraction, which is arranged at the input end of the network and is used to extract sufficient initial target features from the original input under the premise of low parameter quantity and calculation amount. The RS-HGBlock is the core module of the backbone network RS-HGNetV2, which integrates the reparameterization structure Rep, the HGNet hierarchical feature modeling structure and the channel recalibration mechanism SE. The reparameterization is used to enhance the feature extraction capability in the training stage by using a multi-branch structure, and in the inference stage, the multi-branch is equivalent to a single branch to improve the inference speed. The SE enhances the useful channels and suppresses the noise channels by evaluating and scaling the channel responses through global pooling, thereby improving the distinguishability of small targets and weak textures. The DWConv is a deep separable convolution structure, which is used to significantly reduce the parameter quantity and computational quantity in the downsampling process. The SimAM is a three-dimensional parameter-free attention module based on the energy difference of neuron activation. The importance of each feature unit is represented by measuring the energy difference, and the SimAM is embedded in C3k2 to form a CS feature extraction structure, which is integrated into the ASF-V2 feature fusion network, thereby improving the focusing and extraction capability of the model on key information. The ASF-LearnAtt is a self-adaptive attention fusion module, which is the core aggregation unit of the neck structure of the ASF-V2 feature fusion network, including a channel attention CA, a local spatial attention LA and a learnable modulation branch. The CA module enhances the feature channels related to defects and suppresses the redundant and noise channels in the channel dimension, thereby improving the expression capability of key information. The LA module weights each position of the feature map in the spatial dimension, focuses on small defects, local defects and small area defects, and weakens the interference of large area uniform ceramic substrates and complex backgrounds. The learnable modulation branch is used to assign learnable weights to features of different scales and different paths, so that the network can adaptively amplify useful scales and weaken redundant scales, and avoid information conflict or invalid redundancy caused by simple concatenation. The head network Detect-SEAM integrates the SEAM attention module in the detection head of the defect detection network, which is used to improve the detection capability of the model of the defect detection network in complex scenes such as multi-scale, low contrast and partial occlusion. By selectively enhancing and suppressing features, the effective features related to defect targets are strengthened, and the redundant information is weakened, thereby improving the capture accuracy of multi-scale targets. The feature extraction backbone network is used to process the input PTC thermistor image, and outputs four feature maps of large, medium, small and micro, which are input into the feature fusion neck network. The target detection head detects based on the fusion features of large, medium, small and micro scales output by the feature fusion neck network, and outputs the final target detection result. The detection method detects the defects on the surface of the PTC thermistor, including defects, color spots, exposed bottoms and missing prints. The defects are distributed on the outer surface of the PTC thermistor, especially at the four edge chamfers. The color spot is a round grayish brown spot, and the exposed bottom is the exposed ceramic bottom surface due to the damage of the resistance printing layer of the PTC thermistor.

5. The PTC thermistor defect detection method based on the improved YOLOv11 according to claim 4, characterized in that: ​ In the detection method, first, the HGStem performs twice convolution on an input image to obtain a 1 / 4 down-sampled P2 feature map; then, a P2-level low-layer semantic feature is extracted through the RS-HGBlock depth; the P2 feature is down-sampled to 1 / 8 through the DWConv, and is input into the RS-HGBlock to obtain a P3-level middle-low layer semantic; the P3 is again down-sampled through the DWConv, and three groups of RS-HGBlock are alternately stacked on the lightweight branch and the ordinary branch to enhance the middle-high layer semantic expression. Then, the P5 feature map is obtained through the DWConv down-sampling, and the multi-scale context information is extracted through the SPPF three-scale pooling splicing module after the channel expansion of the RS-HGBlock group, and then the key response area is highlighted through the C2PSA channel and spatial attention module, so as to output the multi-scale feature; The ASF-V2 feature fusion network first performs channel dimension reduction on the P4 and P5, P3 and P4 feature maps through 1x1 convolution respectively, and obtains the fused P3 and P5 feature maps through two layers of CS modules after splicing; then, the fused P3 feature is down-sampled to the P4 scale and spliced with the original P4 feature map, and the feature is extracted through two layers of CS; at the same time, the fused P5 feature is down-sampled to the P4 scale to obtain the fused P4 feature map; the P3, P4 and P5 three-way fusion features are input into the ScalSeq module for adaptive weighted combination, and the ASF attention is applied to suppress redundant information and highlight the key response area; then, the multi-directional fusion result is up-sampled to the P2 scale, spliced with the original P2 feature map, and the detail features are extracted through two layers of CS, and then input into the detection head together with the fused features of P3, P4 and P5 of three scales.

6. The PTC thermistor defect detection method based on the improved YOLOv11 according to claim 4, characterized in that: In the SimAM module, H, W and C represent the height, width and channel number of the feature map respectively; the SimAM module adjusts the weight of the energy function of each neuron dynamically, so as to adapt to the feature representation of different channels and spatial positions; The neurons showing spatial inhibition are given higher weights, so that the model can emphasize the fine-scale clues, reduce the background noise in small target detection, and strengthen the attention to significant information; The minimum energy function is obtained by measuring the linear separability between the target neuron and other neurons, and the calculation formula is as follows: Formula 1 ; where t is the target neuron, is the regularization coefficient, M is the number of neurons in each channel; in formula 3, for neurons t with lower energy, the more different from the surrounding active neurons, the higher weight should be given to highlight its importance. Finally, the SimAM module performs feature enhancement processing through sigmoid, as shown in the following formula 2. Formula 2; In the formula, E is the minimum energy function set of neurons; X is input features; is the matrix element multiplication.

7. The PTC thermistor defect detection method based on the improved YOLOv11 according to claim 6, characterized in that: The ASF-V2 feature fusion network first adjusts the channels and preliminarily processes the multi-scale features P2, P3, P4 and P5 from the backbone, adjusts the channel number of the high-scale features P4 and P3 output by the backbone through 1x1 convolution, so as to match with other features to be fused; through up-sampling and down-sampling, the features of different resolutions are scaled to the same size to provide the basis for subsequent fusion; then, the multi-scale feature preliminary integration is realized through the multi-scale fusion module Zoomcat, the features of different scales are aligned and spliced, and the high / width of the medium-scale feature is taken as the target size; the large-scale feature is compressed to the target size through maximum pooling and average pooling, and the key area information is enhanced; Small-scale features are enlarged to the target size by nearest neighbor interpolation, retaining detailed information, and channel dimension splicing, integrating multi-scale information from the three features; After preliminary fusion by Zoomcat, the features are further unified in scale sequence by the ScalSeq module, which converts the features of three scales into a unified channel and scale feature sequence, modeling cross-scale correlation through 3D convolution; The multiscale fusion module Zoomcat inputs three features of different scales , , respectively represent large, medium and small scale features large-scale features Adjustment is made, compress to mesoscale size, the results are added to retain more information; the relevant formula is as follows: Formula 3; wherein, denotes adaptive pooling of feature F to size, denotes adaptive average pooling, denotes the compressed result; For small-scale features Up-sampled to meso-scale size by nearest-neighbor interpolation: Formula 4; wherein, represents up-sampling the feature F to size; The adjusted large-scale, medium-scale, and small-scale features are spliced in the channel dimension, with the formula being: Formula 5.

8. The PTC thermistor defect detection method based on the improved YOLOv11 according to claim 7, characterized in that: The ScalSeq module converts the input features P3, P4, and P5 of three consecutive scales into a unified channel number C through 1x1 convolution; P4 and P5 are upsampled to the spatial size of P3; the formula is as follows: Official 6; wherein: represents a 1x1 convolution to adjust the number of channels to C; P4 and P5 are upsampled to the spatial size of P3; Formula 7; wherein, is a dynamic sampling function, input feature F, target size (H, W), scaling factor s, output , is consistent with size; the , is consistent with feature expansion to 3D feature, splicing 3 features in scale dimension, forming 3D feature body, finally, through 3D convolution, capturing inter-scale association, through batch normalization, activation and 3D pooling compression scale dimension, obtaining the final output of cross-scale fusion feature ; Formula 8; wherein: denotes a compression scale dimension operation; Finally, the ASF-LearnAtt feature aggregation module dynamically adjusts the feature weight through channel attention CA and local spatial attention LA, highlights the defect area, including the defect of the PTC thermistor and color spots, and suppresses background interference. First, input two features to be fused, one is the basic feature, and the other is the cross-scale fusion feature. The channel attention enhances the defect-related channel, generates the channel weight, and suppresses the irrelevant channel. Then, the features are fused and added to integrate the two features. Then, the local spatial attention enhances the spatial response of the defect area, generates the spatial weight, and highlights the defect position. The process is as follows: to input features base features application channel attention: Formula 9; with cross-scale fusion features Fcross After addition, multiply by small-range per-channel scale learnable modulation, small-range learnable amplification or suppression, enhance expression ability, and finally apply local attention: Formula 10; wherein, is the final output feature map of the local spatial attention module, is the intermediate feature map after adding the base feature enhanced by the channel attention and the cross-scale feature, represents a 1x1 convolution operation for learning the height direction attention weight, represents a 1x1 convolution operation for learning the width direction attention weight, represents a Sigmoid activation function, represents an element-wise multiplication.

9. The PTC thermistor defect detection method based on the improved YOLOv11 according to claim 4, characterized in that: The specific method in step S3 is to set the training parameters, adaptively adjust the input image size, use the stochastic gradient descent optimizer to train the LMSF-YOLO model, use Mosaic data augmentation to improve the generalization ability during training, obtain the optimal parameter weight, and reconstruct the LMSF-YOLO network model. The optimized LMSF-YOLO network is obtained by comprehensive evaluation of the indicators of mAP, mAP@50:95, and loss function.

10. The PTC thermistor defect detection method based on the improved YOLOv11 according to claim 4, characterized in that: The specific method in step S4 is to load the trained weight file into the network model of the defect detection network LMSF-YOLO; input the PTC thermistor image in the test set into the LMSF-YOLO model with loaded weights for defect detection, and obtain the center coordinates, width and height of each defect target bounding box, and the corresponding confidence; the detection results are output in the predetermined format and saved as a txt label file.

Citation Information

Cited By

  • Weld defect real-time detection method for industrial welding scene

    CN121982032A