Defect or fault detection method and device based on improved YOLOv5
By combining the improved YOLOv5 network model with the Swin Transformer's shifted window multi-head self-attention mechanism, the problems of low efficiency and poor accuracy in detecting circuit and component defects or faults are solved, and efficient positioning of tiny and hidden defects or faults is achieved. It is suitable for rapid detection of high-reliability and miniaturized electronic devices.
Patent Information
- Application Number
- CN202510859984.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-09-19
AI Technical Summary
In the existing technology, the detection efficiency and accuracy of defects or faults in circuits and components during processing and use are low, and it is difficult to locate small and hidden defects or faults. Traditional detection methods rely on manual labor and are highly subjective, which makes it difficult to meet the needs of high-reliability, miniaturized and multifunctional electronic equipment.
An improved YOLOv5 network model is adopted, combined with the shift window multi-head self-attention mechanism of Swin Transformer, and infrared image training samples are used for defect or fault detection. The image is enhanced by histogram equalization, and an improved YOLOv5 network model is constructed to improve the precision and accuracy of small target detection.
It improves the ability to extract small target features and the detection accuracy in infrared images, realizes the accurate identification and location of defects or faults in the processing and use of power electronic products, ensures product quality and life, and is suitable for rapid detection in complex environments.
Smart Images

Figure CN120673171A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of circuit defects or faults and fault detection, and in particular to a defect or fault detection method and device based on an improved YOLOv5. Background Art
[0002] With the development of information and electronic devices and the widespread use of high-density integrated circuits, the long-term reliability of PCBs, as critical structures, has become a key research topic. PCBs are crucial components of electronic systems, but they are also prone to failure. Their complex multilayer structure is prone to microcracks and hidden solder bumps during processing and use. In high-frequency and high-voltage circuits, the complex multilayer structure of PCBs can lead to defects such as bumps, layer separation, delamination, voids, microcracks, breakdown damage, and micropores during processing (including soldering) and use (for example, when soldering the PCB to the sampling element in smart water, electricity, and gas meters). These defects affect the performance and lifespan of flip-chip electronic devices and PCBs, increasing circuit risk. Therefore, defect detection is crucial. Furthermore, soldering techniques used in manufacturing, such as those for smart electricity (water, gas) meter sampling elements, utilize a combination of projection soldering and brazing. A bump is formed at the soldering site, filled with a suitable solder filler, and then pressed with a high-temperature electrode to form a solid solder joint, enhancing the strength of the solder joint. During the welding process, compression can cause the bump to flatten, resulting in weld cracks that are imperceptible to the naked eye. Furthermore, there can be defects such as solder overflow and missing bumps. These solder joint and weld defects can lead to open solder joints, short circuits, displacement, and other circuit failures, including circuit breaks and wire corrosion. Furthermore, other factors (such as changes in environmental or parameter conditions, stress, and corrosion) can cause functional and performance failures in products, circuits, or power systems during use. Thermal infrared imaging detects temperature changes and the color of images at different temperatures to detect and identify faults. Currently, small targets and minor faults are difficult to detect, and their recognition rate is limited.
[0003] Existing thermal infrared images are sensitive to product changes and temperature distribution and thermal effects of welding parts, and can effectively display product changes and temperature changes of welding parts. In the current defect (fault or abnormal point) detection, visual, automatic optical imaging, X-ray, CT imaging, ultrasonic, laser ultrasonic and terahertz imaging detection can only detect surface defects or faults. The traditional defect or fault detection effect of internal defects or faults and faults including power system defects or faults is limited. With the development of information electronic equipment towards high reliability, miniaturization, lightweight and multifunctionality, high-density integrated circuits with many functional components have been widely used. The development of ultra-large-scale integration technology has increased the requirements for silicon wafer manufacturing processes, and the occurrence and detection difficulty of existing defects or faults are increasing with the trend. The resulting tiny and hidden defects or faults are even more difficult to detect.
[0004] Furthermore, due to the complexity of the production environment and potential random factors in the image acquisition process, smart electricity (water, gas) meter and PCB images are often affected by noise and other interference factors. These interferences include color smear, shadows, and changes in lighting conditions, resulting in the diversity and variability of PCB defects or faults. Therefore, efficient and rapid PCB defect or fault detection is imperative. In actual production environments, meeting the urgency of the production line requires timely completion of PCB and smart electricity (water, gas) meter defect or fault detection and eliminating welding defects or faults, which directly affects product safety and reliability. Therefore, the detection and control of welding quality is of great significance. However, traditional welding quality detection methods rely on manual inspection and empirical judgment, and suffer from problems such as low detection efficiency, poor accuracy, and strong subjectivity. Summary of the Invention
[0005] In view of this, the present invention provides a defect or fault detection method and device based on an improved YOLOv5 to solve the problems in the prior art of low efficiency and accuracy in defect or fault (fault) detection of circuits and components during processing and use, and difficulty in locating, detecting and identifying tiny and hidden defects or faults (faults).
[0006] In a first aspect, the present invention provides a defect or fault detection method based on an improved YOLOv5, the method comprising: obtaining infrared images containing defect or fault parts and parts without defect or fault, and constructing training samples; constructing an improved YOLOv5 network model based on a shifted window multi-head self-attention mechanism of a Swin Transformer; using the training samples to train the improved YOLOv5 network model to obtain a defect or fault detection model; inputting the infrared image of the part to be detected into the defect or fault detection model to obtain a defect or fault detection result.
[0007] In the present invention, an improved YOLOv5 network model is constructed by adopting infrared images as training samples and adding a shifted window multi-head self-attention mechanism based on Swin Transformer into the YOLOv5 network model; thereby, the improved network model is capable of better capturing complex patterns and detailed features, and thus adopting the trained defect or fault (fault) detection module for detection, which not only solves the problems of low efficiency, poor accuracy and strong subjectivity of traditional manual detection, but also improves the model's feature extraction capability and detection precision and accuracy for small targets in infrared images, especially in detecting abnormal points, defects or faults in infrared images of potential processing (including welding parts) of power electronic products and changes in faults, performance and thermal characteristics during use, so as to achieve accurate identification, detection and tracking of defects or faults and faults, which helps to ensure product quality and life guarantee and fault prediction and prevention.
[0008] In an optional embodiment, infrared images containing defective or faulty parts and parts without defective or faulty parts are obtained to construct training samples, including: obtaining infrared images containing defective or faulty parts and parts without defective or faulty parts; enhancing the infrared images using histogram equalization; labeling the enhanced infrared images using preset labels, and constructing training samples based on the labeled images.
[0009] In the present invention, the infrared image is enhanced by adopting histogram equalization, thereby improving the effect of detecting small targets such as welding abnormal points.
[0010] In an optional embodiment, the preset labels include three types of welding abnormality point labels and non-abnormality point labels, and the three types of welding abnormality point labels include cracks, solder overflow and missing convex hulls; the defect or fault detection results include cracks, solder overflow, missing convex hulls and no abnormalities.
[0011] In the present invention, infrared images are annotated with labels such as cracks, solder overflow, and convex hull loss, so that the trained model can detect different defects or fault types.
[0012] In an optional embodiment, the shifted window multi-head self-attention mechanism based on Swin Transformer is used to divide the image into multiple windows, calculate the self-attention weight in each window, and implement a shift mechanism for each window to calculate the attention weights between different windows.
[0013] In the present invention, the shifted window multi-head self-attention mechanism based on Swin Transformer improves the correlation of features in local areas by calculating the self-attention weight in each window. At the same time, by calculating the attention weights between different windows, the model fully learns the features across windows, enhances the model's learning ability for global features, and improves the model's receptive field and context perception. Thus, by adopting the shifted window multi-head self-attention mechanism based on Swin Transformer, the size of the feature map can be maintained unchanged, the feature expression ability is significantly enhanced, and the complex information and small target feature points in the image can be better detected.
[0014] In an optional embodiment, an improved YOLOv5 network model is constructed based on a shifted window multi-head self-attention mechanism of a Swin Transformer, including: obtaining a YOLOv5 network model, wherein the YOLOv5 network model includes a backbone network, a neck network, and a head network; adding a shifted window multi-head self-attention mechanism based on a Swin Transformer to the backbone network of the YOLOv5 network model to obtain an improved backbone network; adding a global information fusion module to the head network of the YOLOv5 network model to obtain an improved head network, wherein the global information fusion module is used to learn and fuse global features of a feature map output by the improved neck network using a global self-attention mechanism; constructing an improved YOLOv5 network model based on the improved backbone network, the neck network, and the improved head network, wherein the improved YOLOv5 network model learns different attention distributions in different representation subspaces.
[0015] In the present invention, by adding a global information fusion module to the YOLOv5 network model, the model's ability to perceive temperature changes in the measured part is improved, and the detection accuracy of the measured abnormal points is further improved.
[0016] In an optional embodiment, the loss function of the improved YOLOv5 network model is expressed by the following formula:
[0017] L=L CIOU +w×L penalty ×Focal Loss
[0018] Focal Loss = -α(1-p) γ ylog(p)-(1-a) γ (1-y)log(1-p)
[0019] Where, L CIOU represents the complete intersection-over-union loss function, w represents the weight coefficient, L penalty Represents the penalty loss function, Focal Loss represents, α and γ represent adjustment parameters, p represents the predicted category probability, and y represents the true category label.
[0020] In an optional embodiment, after the improved YOLOv5 network model is trained using the training samples to obtain a defect or fault detection model, the method further includes: evaluating the defect or fault detection model using recall rate, accuracy rate, average precision mean, average cross-well ratio, and area under the precision-recall curve.
[0021] In the present invention, the recall rate, accuracy rate, average precision mean, average cross-well ratio and area under the precision-recall curve are used to evaluate the defect or fault detection model, thereby clarifying the performance of the defect or fault detection model.
[0022] In an optional embodiment, the infrared image is obtained by using active excitation infrared imaging zoom camera infrared sensing technology and single-line and multi-line laser phase-locked thermal imaging, wherein the working principle of single-line and multi-line laser phase-locked thermal imaging includes: using an excitation unit to modulate a continuous wave laser beam into a pulsed laser beam, and using a cylindrical lens to convert the shape of the pulsed laser beam from point to linear; sending a control signal to the galvanometer scanner to guide the linear laser beam to the surface of the target to be measured; using the thermal waves generated by the linear laser beam along the required excitation line to scan the surface of the target to be measured horizontally and vertically to detect randomly oriented defects or faults; or, the infrared image is obtained by using active excitation infrared imaging zoom camera infrared sensing technology and multi-point array laser phase-locked thermal imaging, wherein the working principle of multi-point array laser phase-locked thermal imaging includes: using a fixed or adjustable multi-point pulsed laser beam to generate thermal waves at multiple points on the surface of the target to be measured; using a high-speed infrared camera to measure the corresponding thermal response, and performing real-time online detection during the manufacturing process of the target to be measured.
[0023] In an optional embodiment, the infrared image is enhanced using histogram equalization, including: converting the infrared image into a grayscale image; statistically analyzing the histogram of the grayscale image, calculating the number of times each pixel value appears, normalizing the number of times to obtain the frequency of each pixel value; accumulating the frequency of each pixel value to obtain a cumulative distribution function; mapping each pixel value in the grayscale image to a new pixel value according to the cumulative distribution function; and replacing the pixel values in the infrared image with the mapped pixel values.
[0024] In an optional embodiment, the working principle of the YOLOv5 network model includes: using the Mosaic data enhancement algorithm to combine multiple input infrared images into one infrared image according to a preset ratio; using the BottleneckCSP structure and CBS structure in the YOLOv5 network to extract the features of the input image, and designing an aggregation strategy through multiple upsampling, splicing, and dot product to perform multi-scale feature fusion on the extracted features, and pass these features to the prediction module, wherein the prediction module uses CIOU_LOSS as the loss function of the Bounding Box; using the prediction module to detect the position and category of the target based on the acquired features to obtain a prediction result.
[0025] In an optional embodiment, the working principle of the improved YOLOv5 network model includes: using a Mosaic data enhancement algorithm to combine multiple input infrared images into an infrared image according to a preset ratio; using an improved backbone network to extract features from the combined infrared image, wherein the feature extraction includes performing convolution operations, normalization operations, nonlinear processing based on activation functions, self-attention weight calculation based on a shifted window multi-head self-attention mechanism of a Swin Transformer, and pooling, splicing and mapping based on a fast spatial pyramid pooling module; using a neck network to upsample and fuse the extracted features; using an improved head network to perform global pooling operations on the features processed by the neck network, learning and fusing global features based on a global information fusion module, and extracting and weighted sum calculations of local features to obtain a new feature map, and the new feature map is used to predict the target category of each position to obtain a prediction result.
[0026] In an optional embodiment, the self-attention weight of the shifted window multi-head self-attention mechanism based on Swin Transformer is calculated using the following formula:
[0027]
[0028] Among them, Q, K, V represent the query matrix, key matrix and value matrix respectively, d k is the dimension of the key;
[0029] The computational complexity of the multi-head self-attention mechanism is calculated using the following formula:
[0030] Ω(W-MSA)=4hwC 2 +2M 2 HkDJ
[0031] Where h×w is the number of input image segmentation patches, C is the number of channels of the input image, and M is the size of a single window.
[0032] In an optional embodiment, the learning and fusion of global features based on the global information fusion module includes: taking the output of global pooling as global context information, fusing it with the original feature map, and jointly calculating the global features and local features through the self-attention mechanism to obtain the self-attention weight of each local position relative to the global context.
[0033] In a second aspect, the present invention provides a defect or fault detection device based on an improved YOLOv5, the device comprising: a sample construction module for acquiring infrared images containing defect or fault parts and non-defect or fault parts to construct training samples; a model construction module for constructing an improved YOLOv5 network model based on the shifted window multi-head self-attention mechanism of Swin Transformer; a model training module for training the improved YOLOv5 network model using the training samples to obtain a defect or fault detection model; a defect or fault detection module for inputting the infrared image of the part to be detected into the defect or fault detection model to obtain a defect or fault detection result.
[0034] In a third aspect, the present invention provides a computer device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to thereby execute the defect or fault detection method based on the improved YOLOv5 of the above-mentioned first aspect or any corresponding embodiment thereof.
[0035] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the defect or fault detection method based on the improved YOLOv5 of the above-mentioned first aspect or any corresponding embodiment thereof.
[0036] In a fifth aspect, the present invention provides a computer program product, comprising computer instructions, which are used to enable a computer to execute the defect or fault detection method based on the improved YOLOv5 of the above-mentioned first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0038] Figure 1 1 is a flow chart of a defect or fault detection method based on an improved YOLOv5 according to an embodiment of the present invention;
[0039] Figure 2 1. A schematic diagram of cracks in components used in a smart electricity (water, gas) meter according to an embodiment of the present invention;
[0040] Figure 3This is a schematic diagram of solder overflow in a component of a smart electricity (water, gas) meter according to an embodiment of the present invention;
[0041] Figure 4 1 is a schematic diagram of an image showing a situation where a convex hull of a component is missing in a smart electricity (water, gas) meter according to an embodiment of the present invention;
[0042] Figure 5 2 is a schematic structural diagram of an improved YOLOv5 network model according to an embodiment of the present invention;
[0043] Figure 6 2 is a schematic diagram of the W-MSA module structure according to an embodiment of the present invention;
[0044] Figure 7 is a schematic diagram of window displacement of a SW-MSA module according to an embodiment of the present invention;
[0045] Figure 8 is a schematic diagram of a confusion matrix of a detection result according to an embodiment of the present invention;
[0046] Figure 9 1 is a schematic diagram of the detection results of infrared thermal imaging of a PCB according to an embodiment of the present invention (there are crack defects or faults in the box);
[0047] Figure 10 is a schematic diagram of detection results including delamination, convex hull, void, crack, and microcrack in 2D and 3D transient amplitude images of a PCB according to an embodiment of the present invention;
[0048] Figure 11 1 is a schematic diagram of detection results including convex hulls, voids, cracks, and microcracks in amplitude and phase images of PCB lock-in thermal imaging according to an embodiment of the present invention;
[0049] Figure 12 1 is a structural block diagram of a defect or fault detection device based on an improved YOLOv5 according to an embodiment of the present invention;
[0050] Figure 13 Schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0051] As described in the background technology, traditional welding quality inspection methods rely on manual inspection and empirical judgment, and have problems such as low inspection efficiency, poor accuracy and strong subjectivity. Since thermal infrared images are highly sensitive to product changes and the temperature distribution of the parts to be tested, they can effectively display product changes and temperature changes of the parts to be tested. Therefore, they can be used to detect abnormal points such as cracks, solder overflow, bulges, separation, delamination, voids, microcracks, breakdown damage and micropores. At the same time, with the rapid development of deep learning technology, image recognition technology based on deep learning has made breakthrough progress in many fields. Especially in the field of target detection, there is still room for improvement in the accuracy of general deep learning target detection methods for small targets such as abnormal points to be tested in infrared images.
[0052] In view of this, the present embodiment provides a defect or fault detection method based on an improved YOLOv5, which adopts infrared images as training samples and adds a shifted window multi-head self-attention mechanism based on Swin Transformer to the YOLOv5 network model to construct an improved YOLOv5 network model; thereby, the improved network model is able to better capture complex patterns and detailed features, and thus the trained defect or fault detection module is used for detection, which not only solves the problems of low efficiency, poor accuracy and strong subjectivity of traditional manual welding quality detection, but also has the advantages of short training time and high computational efficiency compared to the Swin-Transformer model, and improves the accuracy of small target detection compared to the traditional YOLOv5 model.
[0053] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present invention.
[0054] According to an embodiment of the present invention, an embodiment of a defect or fault detection method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0055] In this embodiment, a defect or fault detection method based on an improved YOLOv5 is provided. Figure 1 is a flow chart of a defect or fault detection method according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:
[0056] Step S101, acquire infrared images containing defective or faulty parts and parts without defective or faulty parts, and construct training samples. Specifically, the defective or faulty parts can be defects or faults on various products, such as poor soldering (including crystal oscillators, chips, etc.) status, missing protrusions, cracks, voids, solder overflow, etc. on the PCB of smart water, electricity, and gas meters; chip defects or faults such as missing balls, cold solder joints, surface contaminants, internal contamination, aging of packaging materials, changes in solder composition, electrical short circuits, abnormal resistance, loose packaging, component displacement, abnormal solder balls in key areas, defects or faults in different areas (such as defects or faults in the upper left corner area, defects or faults in the center area), mounting offset, welding short circuits, potential cracks, potential cold solder joints, and other defects or faults and material and structural failures. When constructing training samples, infrared images of the corresponding parts with and without defects or faults can be acquired to form training samples.
[0057] When acquiring infrared images, infrared image acquisition devices known in the related art can be used, such as handheld infrared image acquisition devices. This embodiment does not limit the specific infrared image acquisition device employed. An active excitation infrared imaging zoom camera can be used to acquire infrared images. The heat source selected by the camera can be pulse-selective thermography, lock-in thermography, ultrasonically stimulated vibration thermography, or eddy current thermography. For example, single-line laser lock-in thermography can be used to detect defects or faults in power transformers, cables, welded water, electricity, and gas meters, circuit breakers, conductive wire systems, circuit boards, PCBs, and welds that cause open circuits or short circuits. Multi-line laser lock-in thermography can be used to detect defects or faults in composite materials and multi-layer components in large-area multi-layer welds that cause open circuits or short circuits. Fixed or adjustable multi-point array laser lock-in thermography can be used for efficient large-scale inspection of welded water, electricity, and gas meters, circuit breakers, conductive wire systems, circuit boards, PCBs, and welds involving multiple components.
[0058] The operating principle of single-line and multi-line laser lock-in thermal imaging is as follows: a continuous-wave laser beam is modulated into a pulsed laser beam by an excitation unit, and a cylindrical lens transforms the pulsed laser beam shape from a point to a linear shape. A control unit then sends a control signal to a galvanometer scanner to direct the line laser beam onto the target surface. The line laser beam then generates thermal waves along the desired excitation line, scanning the target surface horizontally and vertically, effectively detecting defects or faults such as randomly oriented cracks. The integration of active excitation infrared imaging zoom camera infrared sensing technology with single-line and multi-line laser lock-in thermal imaging significantly improves thermal imaging sensitivity and resolution. Thermal imaging sensitivity is increased by two orders of magnitude to approximately 102μK, while the resolution of surface and subsurface defects or faults is reduced to 4.8μm. Furthermore, the use of an 808nm optical tube enables detection of micro-hole defects or faults as small as 1.35mm in depth and delamination defects or faults as deep as 430μm in complex multi-layer PCB structures.
[0059] Multi-point array laser lock-in thermal imaging operates as follows: a fixed or adjustable multi-point pulsed laser beam is used to simultaneously generate thermal waves at multiple points on the surface of a target semiconductor chip. A high-speed infrared camera is used to measure the corresponding thermal response, enabling real-time, in-line inspection during the semiconductor chip manufacturing process. The integration of active excitation infrared imaging zoom camera infrared sensing technology with multi-point array laser lock-in thermal imaging significantly improves thermal imaging sensitivity and resolution. Thermal imaging sensitivity has increased by two orders of magnitude, reaching approximately 110μK, while the resolution of surface defects or faults has been reduced to 4μm. Using an 808nm optical tube, it can detect micro-hole defects or faults as small as 1.38mm deep in complex multi-layer PCB structures, as well as delamination defects or faults as deep as 438μm.
[0060] Step S102: Build an improved YOLOv5 network model based on the Shifted Windows Multi-Head Self-Attention mechanism of the Swin Transformer. Shifted Windows Multi-Head Self-Attention (SW-MSA) is a self-attention mechanism used in the Swin Transformer for image processing. It divides the image into multiple small windows and independently calculates self-attention within each window, which reduces computational complexity and improves efficiency. The improved YOLOv5 network, named YOLOv5-WMA, combines the advantages of the YOLOv5 model: strong generalization, high integration, and ease of deployment and training, with the powerful feature extraction capabilities, strong adaptability, and strong robustness of the Swin Transformer. YOLOv5-WMA allows the model to learn different attention distributions in different representation subspaces, which not only enhances the model's expressiveness, making it better at capturing complex patterns and detailed features, but also enhances the model's perception of fine-grained features. This enables the YOLOv5-WMA model to perform excellently in the task of detecting anomaly points in thermal infrared images, making it suitable for detection tasks with low image contrast and complex detection targets.
[0061] Step S103: Using the training samples to train the improved YOLOv5 network model to obtain a defect or fault detection model. Specifically, during training, the infrared images in the training samples are used as input, and the model output is compared with the actual defect or fault results to adjust the model parameters, thereby obtaining a trained network model, i.e., a defect or fault detection model.
[0062] In step S104, the infrared image of the part to be inspected is input into the defect or fault detection model to obtain a defect or fault detection result. Specifically, when performing defect or fault detection, an infrared image of the part to be inspected is first acquired, and then input into the defect or fault detection model to obtain a defect or fault detection result.
[0063] In this embodiment, a defect or fault detection method based on an improved YOLOv5 is provided, and the method includes the following steps:
[0064] Step S201 : Acquire infrared images containing defective or faulty parts and parts without defective or faulty parts to construct training samples.
[0065] Specifically, the above step S201 includes:
[0066] Step S2011: Acquire infrared images of areas containing defects or faults and areas without defects or faults. The defects or faults may be abnormal points to be detected. When acquiring infrared images, a handheld thermal infrared camera may be used to capture infrared images of the areas to be detected. The captured images can be categorized as those containing abnormal points and those with a normal convex hull of the areas to be detected without abnormal points.
[0067] Step S2012: Enhance the infrared image using histogram equalization. Using histogram equalization to enhance the infrared image can improve the effect of detecting small targets such as abnormal points to be detected.
[0068] Specifically, the histogram equalization process is implemented using the following process: First, the captured infrared image is converted into a grayscale image. Then, the histogram of the grayscale image is counted. The number of times each pixel value appears is calculated, and then they are normalized to obtain the frequency of each pixel value. The next step is to calculate the cumulative distribution function (CDF). CDF is the integral of the frequency distribution function, which represents the probability of each pixel value appearing in the original grayscale image. After that, the equalized pixel value is calculated, and each pixel value in the original grayscale image is mapped to a new pixel value, so that the equalized histogram is approximately a uniformly distributed histogram. The pixel values in the original image are replaced with the mapped pixel values to enhance the contrast of the image.
[0069] Step S2013, use preset labels to label the enhanced infrared image, and construct training samples based on the labeled image. In this embodiment, the image data labeling tool Labelimg can be used to label the infrared image to achieve the identification of different defects or fault types. Specifically, the preset labels include three types of abnormal point labels to be tested and no abnormal point labels. The three types of abnormal point labels to be tested include cracks, solder overflow, and convex hull loss; among them, Figure 2 The figure shows the crack situation of components used in smart electricity (water, gas) meters. Figure 3 The figure shows the image diagram of the overflow of solder of components used in smart electricity (water, gas) meters; Figure 4 The figure shows an image diagram of the missing convex hull of components used in smart electricity (water, gas) meters.
[0070] Step S202: construct an improved YOLOv5 network model based on the shifted window multi-head self-attention mechanism of Swin Transformer.
[0071] Specifically, the above step S202 includes:
[0072] Step S2021: Obtain a YOLOv5 network model, where the YOLOv5 network model includes a backbone network, a neck network, and a head network.
[0073] In step S2022, a shifted window multi-head self-attention mechanism based on Swin Transformer is added to the backbone network of the YOLOv5 network model to obtain an improved backbone network. The Swin Transformer-based shifted window multi-head self-attention mechanism is used to divide the image into multiple windows, calculate self-attention weights within each window, and implement a shift mechanism on each window to calculate attention weights between different windows.
[0074] In step S2023, a global information fusion module is added to the head network of the YOLOv5 network model to obtain an improved head network. The global information fusion module is used to learn and fuse global features of the feature map output by the improved neck network using a global self-attention mechanism.
[0075] Step S2024: construct an improved YOLOv5 network model based on the improved backbone network, the neck network, and the improved head network.
[0076] Specifically, if Figure 5 As shown in the figure, the YOLOv5 network model consists of four parts, including input module 1 (Input), basic module 2 (Backbone, backbone network), neck module 3 (Neck, neck network) and prediction module 4 (Head, head network). These modules are specifically Figure 5 The ABCDE boxes shown are composed of some basic modules. Box A is the CBS structure, which consists of a convolutional layer (Conv), batch normalization (BN) and an activation function (SiLU), and is responsible for preliminary feature extraction and data processing. Boxes B and C are BottleNeck structures, which are composed of two CBS structures connected in different ways. The residual connection and feature compression improve the feature expression ability and efficiency of the model. Through feature sharing and fusion, the model's perception of the target features to be detected is further enhanced. Box D is the CSP structure (Cross Stage Partial). This module effectively utilizes shallow and deep features by fusing features from different layers, thereby improving the ability to capture details and global information. Box E is the SPPF structure (SpatialPyramid Pooling-Fast), which is used to perform multi-scale pooling on feature maps to extract global features at different scales.
[0077] Based on the above modules, the YOLOv5 network model operates as follows: First, the Mosaic data augmentation algorithm is applied to the input infrared image. This algorithm combines multiple input images into a single image at a specific ratio. This enriches the dataset, enhances model robustness, strengthens the batch normalization layer, improves the network's small object detection performance, and trains the network to recognize objects within a smaller range. The Bottleneck CSP and CBS structures in the YOLOv5 network serve as the backbone of the entire network, extracting features from the input image. The neck module, which includes the SPPF structure, employs an aggregation strategy using multiple upsampling, concatenation, and dot products to perform multi-scale feature fusion on the extracted features and pass these features to the prediction module. The prediction module uses CIOU_LOSS as the bounding box loss function and is primarily responsible for the final regression prediction. It uses the feature maps extracted by the basic modules to detect the location and category of the object. The network's final output is the prediction result, which includes information such as the category of each object and its corresponding bounding box coordinates.
[0078] In an optional embodiment, taking an infrared image of a specific size as an example, the processing process of the improved YOLOv5 network model is described:
[0079] The input module inputs multiple infrared images of size 640×640×3 (width×height×number of channels). The Mosaic data augmentation algorithm is used to stitch the multiple infrared images together, and the size of the stitched image is kept at 640×640, which helps the model learn richer features.
[0080] After the input module completes processing, the image enters the Backbone module for feature extraction. The image undergoes a 6×6 convolution operation in the Conv_6×6 module. This module uses a 6×6 convolution kernel with a stride of 1, meaning the convolution window moves only one pixel each time. To ensure the output image is the same size as the input image, an edge padding method is used to pad the edges of the output feature map with pixels, ensuring that the width and height of the output feature map match those of the input image. This module uses 32 convolution kernels, each extracting a different feature from the input image. The resulting output feature map has a size of 640×640×32.
[0081] The feature map output by the Conv_6×6 module is fed into the CBS module. After a 3×3 convolution operation, the feature map size is halved and the number of channels is doubled, becoming 320×320×64. Batch normalization (BN) is then performed on the feature map, normalizing the convolution output features for a more uniform distribution. The SiLU activation function is then used to weight each output feature value using a sigmoid function, preserving positive values and compressing negative values, thereby increasing the model's nonlinear expressiveness. The CBS module enhances the feature map's expressiveness.
[0082] When the 320×320×64 feature map output by the CBS module is input to the C3_1_×3 module, it first undergoes a 1×1 convolution operation, increasing the number of channels from 64 to 128. It then undergoes a 3×3 convolution operation, compressing the number of channels from 128 back to 64 while maintaining the spatial size of 320×320. The feature map then undergoes batch normalization to equalize the feature distribution across channels, followed by a SiLU activation function to introduce nonlinearity. The activated feature map is then fused with the original input feature map via a residual connection, preserving the original information and accelerating training. The final output feature map maintains its size of 320×320×64, but boasts enhanced feature representation, capturing richer local and global information. After three passes through the CBS+C3_1_×N modules, the feature map's feature representation is fully enhanced, resulting in an image size of 20×20×1024.
[0083] The SW-MSA module is the W-MSA module (Windows Multi-head Self-Attention) with the addition of a window displacement mechanism. Figure 6 As shown in the figure, the W-MSA module implements a localized application of the multi-head self-attention mechanism by defining a fixed-size window in the image. Unlike the traditional global self-attention mechanism, the image input to the W-MSA module is evenly divided into several different windows, and each pixel in the window will only calculate the self-attention weight within each window. The features within each window are linearly transformed to obtain the query (Q), key (K), and value (V) matrices. The calculated self-attention weight is:
[0084]
[0085] Among them, d kis the dimension of the key. At the same time, the W-MSA module follows the multi-head mechanism, dividing the features into multiple heads. Each head independently performs the above self-attention calculation to capture features in different subspaces. The outputs of all heads are concatenated in the channel dimension and then mapped back to the original dimension through a linear layer. Allowing the model to learn different attention distributions in different representation subspaces not only enhances the model's expressive power, making it better at capturing complex patterns and detailed features, but also enhances the model's perception of fine-grained features, making it suitable for detection tasks such as small target detection in infrared images with low image contrast and complex detection targets. This not only significantly reduces computational complexity but also improves processing efficiency.
[0086] The specific computational complexity of the W-MSA module is:
[0087] Ω(W-MSA)=4hwC 2 +2M 2 HkDJ
[0088] Where h×w is the number of input image patches, C is the number of channels of the input image, and M is the size of a single window. Considering that the W-MSA module only performs self-attention calculations within each divided window, the SW-MSA module is introduced to enhance the model's ability to extract features between windows. The displacement mechanism in the SW-MSA module is shown in the figure below. Figure 7 As shown in the figure, the window originally located in the upper left corner is shifted one pixel to the lower right corner through the displacement mechanism, so that the pixels in the four windows that were originally unconnected can perform self-attention calculations, realizing information exchange between different windows.
[0089] After three iterations of the CBS+C3_1_×N module, the output feature image (size 20×20×1024) is fed into the SW-MSA module. It is then divided into four windows of 5×5 size. Self-attention weights are calculated within each window to enhance the correlation of features within the local region. A shift mechanism is then applied to each window, and attention weights are calculated across windows. This allows the model to fully learn features across windows, enhancing its ability to learn global features and improving its receptive field and contextual awareness. After the SW-MSA module, the feature map size remains unchanged at 20×20×1024, but its feature representation is significantly enhanced, enabling better detection of complex information and small object feature points in the image.
[0090] When the feature maps processed by the SW-MSA module are input into the SPPF module (Fast Spatial Pyramid Pooling), they first undergo 1×1, 5×5, and 9×9 max pooling operations to extract global features at different scales. Each pooling operation captures important features across different receptive fields. These feature maps are then concatenated along the channel dimension, enhancing the model's ability to perceive objects of varying sizes. The concatenated feature maps are then mapped back to 1024 channels via a linear convolutional layer. The output feature map maintains its size at 20×20×1024, but its feature representation is significantly enhanced, containing richer multi-scale global information.
[0091] The feature map is then input into the Neck part and subjected to upsampling and CBS modules to achieve multi-scale feature fusion. Features are extracted and fused from feature maps of different resolutions, so that the final feature map has the richest feature information.
[0092] After processing through the Neck component, feature maps of three different sizes—20×20×1024, 40×40×512, and 80×80×256—are imported into the Head module. To further enhance the ability to capture global information during the detection of anomalies in infrared images, the Head module incorporates a global information fusion module within the YOLOv5 Head. This module uses a global self-attention mechanism to learn and fuse global features from the feature maps, thereby enhancing the model's ability to perceive temperature changes in the target area and further improving the accuracy of anomaly detection.
[0093] Specifically, in the Head module, a global pooling operation is applied to the input feature map to extract global statistical features and generate a feature vector representing the overall information. The output of the global feature pooling is then used as the global context information and fused with the original feature map. The global and local features are jointly calculated through the self-attention mechanism to obtain the self-attention weight of each local position relative to the global context. Specifically, the feature map is mapped into a query (Q), key (K), and value (V) matrix through a linear layer, and the self-attention weight of each position is calculated. The result of the self-attention calculation is fused into the original feature map, so that the local features contain global temperature and anomaly distribution information.
[0094] Finally, the fused feature map is input into the Conv2d (two-dimensional convolutional layer) module of the Head part. Conv2d applies multiple convolution kernels to slide on the input feature map to extract local features. Each convolution kernel calculates the weighted sum within the local area of the feature map and outputs a new feature map to identify important features such as the edge, shape, and texture of the target in the image. It is used to predict the target category at each position and accurately identify the bounding box coordinates of each target to be detected. The feature map processed by the global information fusion module can more accurately reflect the temperature distribution of the test area and the location of abnormal points, thereby significantly improving the overall perception ability and detection accuracy of the model in infrared welding image detection.
[0095] Step S203: Use the training sample to train the improved YOLOv5 network model to obtain a defect or fault detection model. Figure 1 Step S103 of the illustrated embodiment will not be described in detail here.
[0096] Step S204: Input the infrared image of the part to be inspected into the defect or fault detection model to obtain the defect or fault detection result. Figure 1 Step S104 of the illustrated embodiment will not be described in detail here.
[0097] In this embodiment, a defect or fault detection method based on an improved YOLOv5 is provided, and the method includes the following steps:
[0098] Step S301: Acquire infrared images containing defective or faulty parts and parts without defective or faulty parts to construct training samples. Figure 1 Step S101 of the illustrated embodiment will not be described in detail here.
[0099] Step S302: Build an improved YOLOv5 network model based on the Swin Transformer's shifted window multi-head self-attention mechanism; see Figure 1 Step S102 of the illustrated embodiment will not be described in detail here.
[0100] In step S303, the improved YOLOv5 network model is trained using the training samples to obtain a defect or fault detection model. Specifically, when training the model, the training parameters can be set to Batch-Size (batch size) of 16, imgSize of 640, epochs of 300, learning rate (lr0, lrf) of 0.01, weight decay (weight_decay) of 0.0003, and IoU threshold of 0.35. The AdamW optimizer is also used for parameter adjustment.
[0101] The loss function of the improved YOLOv5 network model is expressed as follows:
[0102] L=L CIOU +w×L penalty ×Focal Loss
[0103] Focal Loss = -α(1-p) γ ylog(p)-(1-a)γ(1-y)log(1-p)
[0104] Where, L CIOU represents the complete intersection-over-union loss function, w represents the weight coefficient, L penalty Represents the penalty loss function, Focal Loss represents, α and γ represent adjustment parameters, p represents the predicted category probability, and y represents the true category label.
[0105] In the process of model training, the above loss function can be used to adjust the parameters in the model. The weight coefficient in the above loss function can be determined according to the ratio of the target size and the image size. The target size S obj The area size can be determined by the product of the width and height of the target; the image size S img It can be the total area determined by the image width and height. The penalty loss function is a customized small target detection penalty term, and the specific calculation method can be determined based on the specific characteristics of the small target being detected.
[0106] Step S304 : Evaluate the defect or fault detection model using recall rate, accuracy rate, average precision mean, average cross-well ratio, and area under the precision-recall curve.
[0107] Specifically, the recall rate R (Recall) represents the ratio of the predicted correct detection frames to the actual true frames among all the detection frames predicted by the model. The recall rate is calculated using the following formula:
[0108]
[0109] The accuracy P (Precision) represents the proportion of correct detection frames (positive samples) among all the detection frames predicted by the model. The accuracy is calculated using the following formula:
[0110]
[0111] In the mAP-50 (average precision) model, AP measures the detection performance of a specific category. Different confidence levels and IoU (intersection over union) thresholds correspond to different precision and recall rates. The area of the two-dimensional curve formed by these two rates is the AP value. The average of the AP values across different categories is the mAP, which measures the detection performance of multiple categories.
[0112] The mean intersection-over-well (mloU) represents the average of the intersection-over-union (IU) of multiple small objects. For a single target bounding box prediction result and the true box, let their intersection area be I and their union area be U, then the IU is A series of IoU values are calculated for the bounding boxes of all small targets. Assuming there are m small targets, the mlOU calculation formula is:
[0113] The area under the precision-recall curve (AUPRC) is similar to the calculation of AP, but here there are no restrictions on conditions such as the IoU threshold. The precision-recall curve (PR curve) is directly plotted based on the precision values corresponding to the model prediction results at different recall rates. The area under the curve is then calculated using numerical calculation methods such as integration to obtain the AUPRC.
[0114] Step S305: Input the infrared image of the part to be detected into the defect or fault detection model to obtain the defect or fault detection result. Figure 1 Step S104 of the illustrated embodiment will not be described in detail here.
[0115] As a specific application example of the embodiment of the present invention, Figure 5 As shown, the defect or fault detection method can be implemented using the following process:
[0116] 1. Based on imaging principles such as pulsed thermography, phase-locked thermography, ultrasonically stimulated vibration thermography, and eddy current thermography, infrared images containing defective or faulty parts and parts without defective or faulty parts are obtained to construct training samples.
[0117] To acquire an image, the product is placed on a cross-adjustment stage, and infrared images are captured using an infrared imager or other imaging device. Defects or faults specifically include solder defects or faults on the PCB of smart electricity, water, and gas meters, as well as defects or faults on various chips (such as micron-level multi-layer chips, metering chips, main control chips (MCU chips), communication chips, security encryption chips (ESAM chips), clock chips, memory chips, and RF front-end chips). Solder faults on PCBs include lw (cracks), qlwy (solder overflow), tbqs (missing convex hulls), and tbzc (normal convex hulls). Defects or faults on various chips include lq (missing solder ball), xh (cold solder joint), mqxs (missing protrusion), qtxb (protrusion deformation), lk (crack), kd (void), bmwr (surface contaminants), nbwr (internal contamination), fbcl (aging of packaging materials), hldf (change in solder composition), dqdl (electrical short circuit), dzcy (abnormal resistance), fbss (loose packaging), bjwy (component displacement), gjqy (abnormal solder ball in critical area), zjsq (defect or fault in the upper left corner area), zxqy (defect or fault in the center area), tzpy (mounting offset), hjdl (soldering short circuit), qzlk (potential crack), qxh (potential cold solder joint), etc.
[0118] When collecting infrared images, an active excitation infrared imaging zoom camera can be used, and the infrared image can be collected according to the following imaging requirements and imaging parameters:
[0119] The infrared detection area includes near infrared (0.75um-1.5um), mid-infrared (1.5um-20um) and far infrared (20um-1000um). At the same time, professional-grade high-speed and high-resolution imaging chips are used. For example, the imaging chip has flexible pixel reading and fast readout capabilities, supporting high frame rates of 200 frames / s and multiple resolution requirements. The lens uses an optical lens with high transmittance to reduce light loss and make full use of weak light for imaging. The lens used at the same time can be adapted to a variety of focal lengths and apertures, with the focal length ranging from wide angle (such as 2mm) to telephoto (such as 120mm), which facilitates the realization of zoom function. For example, a lens based on phase detection autofocus technology can be used.
[0120] In addition, the lens has multiple modes. In black and white mode, the resolution must be at least 20 megapixels, with 2448×2048 being the most common. In color mode, the resolution must be no less than 4K, or 3840×2160. Near-infrared enhanced mode is often used for special inspection needs, such as inspecting the internal structure of chips, with a resolution generally at 1920×1080. The signal-to-noise ratio must be above 60dB.
[0121] 2. Build an improved YOLOv5 network model based on the Swin Transformer's shift window multi-head self-attention mechanism.
[0122] 3. Divide the training samples into a training set and a test set. Use the training set to train the improved YOLOv5 network model to obtain a defect or fault detection model. Use the test set to test the performance of the defect or fault detection model.
[0123] Specifically, this example uses 3161 infrared images of the test areas of smart electricity (water, gas) meter sampling components, taken in real production environments, as training samples. The training samples are divided into a training set and a test set. The defect or fault detection model is tested using the test set. The test results are shown in Table 1 below. The table shows that the recall (R) rate reaches 0.833, the precision (P) rate reaches 0.916, and the mean average precision (mAP-50) reaches 0.929.
[0124] Table 1
[0125]
[0126] In addition, 593 infrared images were used as the test set for testing, and the confusion matrix of the test results is as follows Figure 8 As shown in the figure, the false detection rate of crack (lw) is only 0.07, the false detection rate of solder overflow (qlwy) is 0.02, the false detection rate of normal convex hull (tbzc) is 0.08, and the false detection rate of missing convex hull (tbqs) is 0.12.
[0127] 4. Input the infrared image of the part to be inspected into the defect or fault detection model to obtain the defect or fault detection result.
[0128] Specifically, the infrared images to be tested in this embodiment include 50 infrared images of intelligent electricity (water, gas) meter PCB boards, high-speed boards used for computing power, high-speed computers and high-end servers, blind hole boards used for mobile phones and notebooks, aluminum substrates used for power supplies and LED lighting, copper substrates used for precision communications and high-frequency circuits, radar, satellite communications and high-power density electronic products, and 690 infrared images of power system automatic control circuit boards. By inputting these infrared images into the defect or fault detection model, various defects or faults such as layer separation, delamination, breakdown damage and micropores that occur during processing and use are detected. Among them, Figure 9 The following is the test result of infrared thermal imaging of PCB (there are crack defects or faults in the box); Figure 10 The following are the detection results of delamination, convex hull, void, crack and micro crack in the 2D (left side) and 3D (right side) transient amplitude images of PCB; Figure 11Shown are the detection results of convex hulls, voids, cracks, and microcracks in the amplitude image (upper half of the figure) and phase image (lower half of the figure) of the PCB board's phase-locked thermal imaging.
[0129] The defect or fault detection method provided in this embodiment provides a basis for identifying fault anomalies, repairing solder joints, and improving the prevention process of cold solder joints, and provides a method for predicting and preventing defects during production and use. Furthermore, by detecting poor solder joints caused by changes in material properties, it can extend service life or resolve potential reliability issues, ensuring the health of the chip throughout its entire lifecycle.
[0130] This embodiment also provides a defect or fault detection device based on an improved YOLOv5, which is used to implement the above-mentioned embodiments and preferred embodiments. Details that have already been described will not be repeated here. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.
[0131] This embodiment provides a defect or fault detection device based on an improved YOLOv5, such as Figure 12 As shown, including:
[0132] The sample construction module 121 is used to obtain infrared images containing defective or faulty parts and parts without defective or faulty parts to construct training samples;
[0133] A model building module 122 is used to build an improved YOLOv5 network model based on the shift window multi-head self-attention mechanism of Swin Transformer;
[0134] A model training module 123 is configured to train the improved YOLOv5 network model using the training samples to obtain a defect or fault detection model;
[0135] The defect or fault detection module 124 is used to input the infrared image of the part to be detected into the defect or fault detection model to obtain a defect or fault detection result.
[0136] The further functional description of each of the above modules is the same as that of the above corresponding embodiments and will not be repeated here.
[0137] The embodiment of the present invention also provides a computer device having the above Figure 12 The defect or fault detection device based on the improved YOLOv5 is shown.
[0138] See also Figure 13 , Figure 13 is a structural diagram of a computer device provided by an optional embodiment of the present invention, such as Figure 13 As shown, the computer device includes: one or more processors 10, memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components utilize different buses to communicate with each other and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in the memory or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Equally, multiple computer devices can be connected, and each device provides part of the necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 13 A processor 10 is taken as an example.
[0139] The processor 10 may be a central processing unit, a network processor, or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.
[0140] The memory 20 stores instructions that can be executed by at least one processor 10, so as to enable at least one processor 10 to execute the method shown in the above embodiment.
[0141] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created based on the use of a computer device for displaying a small program landing page, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0142] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid-state drive; the memory 20 may also include a combination of the above types of memory.
[0143] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or a communication network.
[0144] The embodiment of the present invention also provides a computer-readable storage medium. The above-mentioned method according to the embodiment of the present invention can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.
[0145] A portion of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the form in which the computer program instruction exists in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc. Accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium that can be accessed by the computer.
[0146] Although the embodiments of the present invention have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention. Such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A defect or fault detection method based on improved YOLOv5, characterized in that: The method comprises: Acquire infrared images containing defective or faulty parts and parts without defective or faulty parts to construct training samples; An improved YOLOv5 network model is constructed based on the Swin Transformer's shifted window multi-head self-attention mechanism; Using the training samples to train the improved YOLOv5 network model to obtain a defect or fault detection model; The infrared image of the part to be detected is input into the defect or fault detection model to obtain the defect or fault detection result.
2. The method according to claim 1, characterized in that Acquire infrared images containing defects or faults and areas without defects or faults, and construct training samples, including: Acquire an infrared image containing defective or faulty parts and parts without defective or faulty parts; Enhancing the infrared image by using histogram equalization; The enhanced infrared images are annotated with preset labels, and training samples are constructed based on the annotated images.
3. The method according to claim 2, characterized in that The preset labels include three types of welding abnormal point labels and non-abnormal point labels. The three types of welding abnormal point labels include cracks, solder overflow and convex hull missing; the defect or fault detection results include cracks, solder overflow, convex hull missing and no abnormality.
4. The method according to claim 1, wherein The shifted window multi-head self-attention mechanism based on Swin Transformer is used to divide the image into multiple windows, calculate self-attention in each window, and implement a shift mechanism for each window to calculate the attention weights between different windows.
5. The method according to claim 1, wherein An improved YOLOv5 network model is constructed based on the Swin Transformer's shifted window multi-head self-attention mechanism, including: Obtain a YOLOv5 network model, where the YOLOv5 network model includes a backbone network, a neck network, and a head network; A shifted window multi-head self-attention mechanism based on Swin Transformer is added to the backbone network of the YOLOv5 network model to obtain an improved backbone network; A global information fusion module is added to the head network of the YOLOv5 network model to obtain an improved head network, wherein the global information fusion module is used to learn and fuse global features of the feature map output by the improved neck network using a global self-attention mechanism; An improved YOLOv5 network model is constructed based on the improved backbone network, neck network and improved head network. The improved YOLOv5 network model learns different attention distributions in different representation subspaces.
6. The method according to claim 1, characterized in that The loss function of the improved YOLOv5 network model is expressed as follows: L=L CIOU +w×L penalty ×FocalLoss Focal Loss1-α(1-p) γ ylog(p)-(1-a) γ (1-y)log(1-p) Where, L CIOU represents the complete intersection-over-union loss function, w represents the weight coefficient, L penalty Represents the penalty loss function, Focal Loss represents, α and γ represent adjustment parameters, p represents the predicted category probability, and y represents the true category label.
7. The method according to claim 1, characterized in that After training the improved YOLOv5 network model using the training samples to obtain a defect or fault detection model, the method further includes: The defect or fault detection model is evaluated using recall rate, accuracy rate, average precision mean, average cross-well ratio and area under the precision-recall curve.
8. The method according to claim 1, characterized in that The infrared image is obtained by using active excitation infrared imaging zoom camera infrared sensing technology and single-line and multi-line laser phase-locked thermal imaging. Among them, the working principles of single-line and multi-line laser lock-in thermal imaging include: An excitation unit is used to modulate a continuous wave laser beam into a pulsed laser beam, and a cylindrical lens is used to convert the shape of the pulsed laser beam from a point shape to a linear shape; Send control signals to the galvanometer scanner to guide the linear laser beam to the target surface to be measured; The thermal waves generated by a line laser beam along the desired excitation line are used to scan the target surface horizontally and vertically to detect randomly oriented defects or faults. Alternatively, the infrared image is obtained by using active excitation infrared imaging zoom camera infrared sensing technology and multi-point array laser phase-locked thermal imaging Among them, the working principles of multi-point array laser lock-in thermal imaging include: Use fixed or adjustable multi-point pulsed laser beam to generate thermal waves at multiple points on the surface of the target to be measured; A high-speed infrared camera is used to measure the corresponding thermal response and perform real-time online detection during the manufacturing process of the target to be tested.
9. The method according to claim 2, characterized in that The infrared image is enhanced by using histogram equalization, including: converting the infrared image into a grayscale image; Counting the histogram of the grayscale image, calculating the number of times each pixel value appears, and normalizing the number to obtain the frequency of each pixel value; The frequency of each pixel value is accumulated to obtain the cumulative distribution function; mapping each pixel value in the grayscale image to a new pixel value according to the cumulative distribution function; The pixel values in the infrared image are replaced with the mapped pixel values.
10. The method according to claim 5, characterized in that The working principle of the YOLOv5 network model includes: The Mosaic data enhancement algorithm is used to combine multiple input infrared images into one infrared image according to the preset ratio; The BottleneckCSP and CBS structures in the YOLOv5 network are used to extract features from the input image. An aggregation strategy is designed through multiple upsampling, splicing, and dot product to perform multi-scale feature fusion on the extracted features. These features are then passed to the prediction module, which uses CIOU_LOSS as the Bounding Box loss function. The prediction module is used to detect the location and category of the target based on the acquired features to obtain the prediction result.
11. The method according to claim 10, characterized in that The working principles of the improved YOLOv5 network model include: The Mosaic data enhancement algorithm is used to combine multiple input infrared images into one infrared image according to the preset ratio; The improved backbone network is used to extract features from the combined infrared images. The feature extraction includes convolution operations, normalization operations, nonlinear processing based on activation functions, self-attention weight calculation based on the shifted window multi-head self-attention mechanism of the Swin Transformer, and pooling, splicing, and mapping based on the fast spatial pyramid pooling module. The neck network is used to upsample and fuse the extracted features; The improved head network is used to perform global pooling operations on the features processed by the neck network, learn and fuse global features based on the global information fusion module, and extract and weight and calculate local features to obtain a new feature map. The new feature map is used to predict the target category at each position to obtain the prediction result.
12. The method according to claim 11, characterized in that The self-attention weight of the shifted window multi-head self-attention mechanism based on Swin Transformer is calculated using the following formula: Among them, Q, K, V represent the query matrix, key matrix and value matrix respectively, d k is the dimension of the key; The computational complexity of the multi-head self-attention mechanism is calculated using the following formula: Ω(W-MSA)=4hwC 2 +2M 2 hwC Where h×w is the number of input image segmentation patches, C is the number of channels of the input image, and M is the size of a single window.
13. The method according to claim 11, characterized in that The learning and fusion of global features based on the global information fusion module include: The output of global pooling is used as the global context information and fused with the original feature map. The global features and local features are jointly calculated through the self-attention mechanism to obtain the self-attention weight of each local position relative to the global context.
14. A defect or fault detection device based on improved YOLOv5, characterized in that: The device comprises: A sample construction module is used to obtain infrared images containing defective or faulty parts and parts without defective or faulty parts to construct training samples; Model building module, used to build an improved YOLOv5 network model based on the Swin Transformer's shifted window multi-head self-attention mechanism; A model training module, configured to train the improved YOLOv5 network model using the training samples to obtain a defect or fault detection model; The defect or fault detection module is used to input the infrared image of the part to be detected into the defect or fault detection model to obtain the defect or fault detection result.