Electronic device classification detection method based on multi-scale feature fusion
By adopting multi-scale feature fusion technology and improved YOLOv10 model in the electronic device recognition model, combined with the EMA attention mechanism, VanillaBlock module and Shape-IoU loss function, the accuracy and speed problems of the electronic component recognition model in morphological differences and complex environments are solved, and efficient and accurate electronic device classification detection is achieved.
Patent Information
- Application Number
- CN202411753302.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-02
- Publication Date
- 2025-05-13
AI Technical Summary
When the existing electronic component identification model recognizes complex and diverse electronic component, it is difficult to adapt to the morphological differences of different devices and the complex detection environment, resulting in poor detection accuracy and speed.
The electronic device classification detection method based on multi-scale feature fusion is adopted. The improved YOLOv10 model combines the EMA multi-scale attention mechanism, VanillaBlock module and Shape-IoU border loss function to improve the feature extraction ability and detection accuracy of the model.
It realizes accurate classification and real-time identification of electronic devices, significantly improves detection accuracy and speed, reduces misjudgment and misjudgment, and improves detection efficiency and overall reliability.
Smart Images

Figure CN119992152A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and in particular to an electronic device classification and detection method based on multi-scale feature fusion. Background Art
[0002] With the rapid development of electronic technology, basic electronic components such as capacitors, inductors, light-emitting diodes, resistors and transistors are increasingly widely used, and the accuracy of their classification and detection is directly related to the performance and reliability of the equipment. However, traditional detection methods mainly rely on manual visual inspection and basic testing equipment, which have problems such as low efficiency and prone to errors, and are difficult to meet the needs of modern large-scale production.
[0003] In order to meet the urgent market demand for high-quality and efficient detection, the research and development of electronic device classification detection methods has become crucial. Especially on the production line, the need to handle massive electronic components has prompted researchers to explore more efficient and automated detection systems. As an advanced target detection algorithm, YOLOv10 uses computer vision and deep learning technology to significantly improve detection efficiency and accuracy, thereby effectively solving the shortcomings of traditional methods.
[0004] Capacitors, inductors, light-emitting diodes, resistors, and transistors each have unique physical properties and structural forms. Capacitors are generally cylindrical or rectangular, while inductors are commonly seen in different winding forms; light-emitting diodes are thin and long metal wires, resistors have various appearances, and transistors may have a variety of packaging forms. These diversities increase the complexity of automatic classification and detection, but YOLOv10, with its efficient feature extraction capabilities, can quickly identify and classify these different types of electronic components by processing device images.
[0005] The multi-scale feature fusion technology used by YOLOv10 enables the model to identify specific devices in complex backgrounds and adapt to different lighting conditions and background interference. Through training on a large amount of labeled data, YOLOv10 can learn the unique characteristics of each device and achieve efficient classification. In addition, the CIoU loss function is used to optimize the border positioning, which further improves the accuracy of the model in object detection, ensuring that classification tasks can be completed efficiently and accurately in practical applications.
[0006] Nevertheless, the size differences, inconsistent placement angles and environmental factors of electronic devices still have a certain impact on the detection effect. Therefore, in order to improve the intelligence level of the detection system, a more comprehensive monitoring system can be built by combining the improved multi-scale feature fusion model. In addition, the use of Internet of Things technology to feed back the detection results to the production line in real time can achieve dynamic adjustment and optimization, further improving the overall production efficiency and product quality.
[0007] Traditional detection methods have insufficient recognition accuracy when facing complex and diverse electronic components such as capacitors, inductors, and light-emitting diodes, and it is difficult to meet real-time requirements in mass production. To this end, a detection method that can efficiently extract multi-scale features, handle small target detection, and ensure high accuracy while reducing the amount of calculation is needed to cope with the morphological differences of different devices and complex detection environments. To avoid the technical problem that the network has poor detection accuracy and speed when identifying and classifying electronic components due to poor model detection performance, a classification and detection method for electronic components based on multi-scale feature fusion is proposed. Summary of the invention
[0008] The purpose of the present invention is to solve the problem that the existing models and methods for electronic component identification cannot adapt well to the morphological differences of different devices and the complex detection environment when used for complex and diverse electronic component identification, resulting in poor detection accuracy and speed when identifying and classifying electronic components. A method for electronic component classification and detection based on multi-scale feature fusion is proposed.
[0009] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0010] A method for classifying and detecting electronic devices based on multi-scale feature fusion comprises the following steps:
[0011] Step S1: obtaining original images of various common and typical electronic devices, integrating them with open source data sets, and screening and sorting the images;
[0012] Step S2: pre-process the original image using the canny edge operator and blur processing method to obtain several data set images;
[0013] Step S3: Use Labelimg to annotate the target data set, annotate the collected electronic device images, divide the data set according to a certain ratio, and finally obtain the electronic component data set, and divide the data set into a training set, a validation set, and a test set;
[0014] Step S4: Obtain an improved YOLOv10 model, use the model to recognize the acquired electronic device image, and use precision (Precision), recall rate (Recall), mean average precision (mAP) and frames per second (FPS) as evaluation indicators.
[0015] In step S1, a high-definition camera is used to take original images of five common and typical electronic devices in the electronic laboratory of the affiliated unit, and the images are integrated with open source data sets, and the images are screened and sorted.
[0016] In step S2, the obtained data set images are divided into five types: capacitor, inductor, light emitting diode, resistor, and transistor;
[0017] In step S3, the collected electronic device images are annotated, and the defects are divided into five categories: capacitor, inductor, light-emitting diode, resistor, and transistor, among which the capacitor is "Capacitor", the inductor is "Inductor", the light-emitting diode is "Led", the resistor is "Resistor", and the transistor is "Transistor".
[0018] In step S4, the improved YOLOv10 model obtained is specifically:
[0019] Initialize the input image and input it to the first Conv module, the output of the first Conv module is connected to the input of the second Conv module, the output of the second Conv module is connected to the input of the first C2f module, the output of the first C2f module is connected to the input of the third Conv module, the output of the third Conv module is connected to the input of the second C2f module, the output of the second C2f module is connected to the input of the first CSDown module, the output of the first CSDown module is connected to the input of the third C2f module, the output of the third C2f module is connected to the input of the second SCDown module, the output of the second SCDown module is connected to the input of the first C2fCIB module, the output of the first C2fCIB module is connected to the input of the SPPF module, and the output of the SPPF module is connected to the input of the EMA module;
[0020] The output of the EMA module is connected to the input of the first Upsample module and the input of the third VanillaBlock module, the output of the first Upsample module is connected to the input of the first Concat module, the output of the first Concat module is connected to the input of the fourth C2f module, the output of the fourth C2f module is connected to the input of the second Upsample module, the output of the second Upsample module is connected to the input of the second Concat module, the output of the second Concat module is connected to the input of the fifth C2f module, and the output of the fifth C2f module is connected to the input of the first VanillaBlock module;
[0021] The output of the first VanillaBlock module is connected to the input of the fourth Conv module and the input of the first v10Detect module, the output of the fourth Conv module is connected to the input of the third Concat module, the output of the third Concat module is connected to the input of the sixth C2f module, the output of the sixth C2f module is connected to the input of the second VanillaBlock module, and the output of the second VanillaBlock module is connected to the input of the second v10Detect module and the input of the third CSDown module;
[0022] The output of the third CSDown module is connected to the input of the fourth Concat module, the output of the fourth Concat module is connected to the input of the second C2fCIB module, the output of the second C2fCIB module is connected to the input of the third VanillaBlock module, and the output of the third VanillaBlock module is connected to the input of the third v10Detect module;
[0023] The output of the first v10Detect module, the output of the second v10Detect module, and the output of the third v10Detect module constitute the output of the total feature.
[0024] The EMA module is specifically:
[0025] The input features are fed into the Groups module, and the output of the Groups module is connected to the input of the first Re-weight module, the input of the second Re-weight module, the input of the first X Avg Pool module, the input of the Y Avg Pool module, and the input of the Conv module;
[0026] The output of the first X Avg Pool module and the output of the Y Avg Pool module are connected to the input of the second X Avg Pool module, the output of the second X Avg Pool module is connected to the input of the first Sigmoid module and the input of the second Sigmoid module, the output of the first Sigmoid module and the output of the second Sigmoid module are connected to the input of the second Re-weight module, the output of the second Re-weight module is connected to the input of the GroupNorm module, and the output of the GroupNorm module is connected to the input of the first Avg Pool module and the input of the first Matmul module;
[0027] The output of the first Avg Pool module is connected to the input of the first Softmax module, and the output of the first Softmax module is connected to the input of the second Matmul module;
[0028] The output of the Conv module is connected to the input of the first Matmul module and the input of the second Avg Pool module, the output of the second Avg Pool module is connected to the input of the second Softmax module, the output of the second Softmax module and the output of the GroupNorm module are connected to the input of the first Matmul module, the output of the first Matmul module and the output of the second Matmul module are connected to the input of the splicing module, the output of the splicing module is connected to the input of the third Sigmiod module, the output of the third Sigmiod module is connected to the input of the second Re-weight module, and the output of the second Re-weight module is used to output the output features of the EMA module.
[0029] The VanillaBlock modules are:
[0030] The input features of the VanillaBlock module are input to the 4×4 convolution module, the output of the 4×4 convolution module is connected to the input of the first 1×1 convolution module, the output of the first 1×1 convolution module is connected to the input of the second 1×1 convolution module, the output of the second 1×1 convolution module is connected to the input of the third 1×1 convolution module, the output of the third 1×1 convolution module is connected to the input of the global average pooling layer, the output of the global average pooling layer is connected to the input of the fully connected layer, and the output of the fully connected layer is used to output the output features of the VanillaBlock module.
[0031] The Shape-IoU loss function is used as the border loss function of the improved YOLOv10 model; specifically:
[0032] Intersection over Union (IoU) is a measure of the accuracy of detecting corresponding objects in a specific dataset. IoU is a simple measurement standard that can be used to measure any task that produces a predicted bounding box in the output. IoU is calculated by dividing the intersection of the predicted box (A) and the true box (B) by the union of the two, as shown in Formula 1.
[0033]
[0034] Shape-IoU evaluates the impact of the shape and scale of the bounding box on the regression results by calculating the width (ww) and height (hh) of the bounding box and the weight coefficient. This method makes up for the deficiency of ignoring the shape and scale factors of the bounding box itself in traditional methods, thereby improving the accuracy of detection.
[0035]
[0036]
[0037] Distance in Shape-IoU shape The main focus is on the shape and scale of the bounding box, and the accuracy of object detection is improved by considering these factors.
[0038]
[0039] Ω in Shape-IoU shape The shape and scale of the bounding box are taken into account when calculating the loss, making the bounding box regression more accurate. Traditional bounding box regression methods mainly focus on the geometric relationship between the predicted box and the true box (GT box), such as position, shape, and angle, but ignore the impact of the shape and scale of the bounding box itself on the regression results. Shape-IoU introduces Ω shape To make up for this deficiency, the loss function can more comprehensively consider the characteristics of the bounding box.
[0040]
[0041]
[0042]
[0043] where w gt ,h gt Represent the width and height of the real bounding box respectively, v represents the minimum closure area of the predicted box and the real box, (xc, yc), They represent the center coordinates of the predicted box and the real box respectively, t is the scale factor, and its value range is between [0, 1.5]. ww and wh are the weight coefficients of the real box in the horizontal and vertical directions respectively, and their values are positively correlated with the width and height of the real box. 2 represents the relative distance between the centers of the two. ww and hh are the weight coefficients in the horizontal and vertical directions respectively, and their values are related to the shape of the GT box. scale is the scale factor, and its value is related to the scale of the target in the electronic device dataset. For smaller targets, the larger the scale value, the corresponding bounding box regression loss is as follows:
[0044] L Shape-IoU =1-IoU+distance shape +0.5×Ω shape (8).
[0045] Compared with the prior art, the present invention has the following technical effects:
[0046] The technical effects of the present invention are:
[0047] 1) The electronic device classification model of the present invention surpasses the existing detection model in terms of accuracy and speed; it can realize the accurate classification of electronic devices, and has the ability to be applied to the real-time classification and identification of electronic devices at the factory end, helping the staff to make quick judgments and avoid omissions and misoperations, which would cause immeasurable losses;
[0048] 2) The improved YOLOv10s model of the present invention is significantly superior to the existing detection models in terms of classification accuracy and speed of electronic devices, and can achieve accurate identification and real-time classification detection of a variety of electronic components. This method effectively reduces misjudgments and missed judgments during the detection process, thereby improving detection efficiency and overall reliability;
[0049] 3) By introducing the EMA multi-scale attention mechanism, the present invention can maintain high-precision detection under complex backgrounds and changing lighting conditions. This attention mechanism enhances the model's ability to capture multi-scale features of electronic components, making it more adaptable in actual application environments.
[0050] 4) After the present invention adopts the VanillaBlock module, the model calculation efficiency is improved, reducing the burden of parameters and calculation. The module design is concise and effective, so that the model can meet the real-time requirements of the factory while maintaining high accuracy, and achieve efficient and stable detection performance;
[0051] 5) The present invention uses Shape-IoU as the border regression loss function, which effectively improves the accuracy of small target detection. Shape-IoU takes into account the shape and scale factors of the bounding box, so that the present invention has higher accuracy when processing electronic components of different sizes and shapes, and is particularly suitable for the detection of small components such as capacitors and resistors;
[0052] 6) The improved YOLOv10s model of the present invention shows excellent real-time performance in actual detection applications, and can perform electronic device classification detection at a high frame rate on the production line. This feature ensures that the detection system can adapt to the fast production environment and improves the overall production efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] The present invention will be further described below in conjunction with the accompanying drawings and embodiments:
[0054] Figure 1 It is a schematic diagram of the network structure of the present invention;
[0055] Figure 2 is a flow chart of the present invention;
[0056] Figure 3 This is a schematic diagram of the structure of the multi-scale feature fusion EMA attention module in the present invention;
[0057] Figure 4 This is a schematic diagram of the structure of VanillaBlock, the core module of VanillaNet in the present invention;
[0058] Figure 5 This is an example diagram of the “Capacitor” detection result in an embodiment of the present invention;
[0059] Figure 6 This is an example diagram of the “Inductor” detection result in an embodiment of the present invention;
[0060] Figure 7 This is an example diagram of the detection result of "light emitting diode" (Led) in an embodiment of the present invention;
[0061] Figure 8 This is an example diagram of the “Resistor” detection result in an embodiment of the present invention;
[0062] Fig. 9 This is an example diagram of the “Transistor” detection results in an embodiment of the present invention. DETAILED DESCRIPTION
[0063] A method for classifying and detecting electronic devices based on multi-scale feature fusion includes the following steps:
[0064] S1. Use a high-definition camera to take original images of five common and typical electronic devices in the electronic laboratory of the affiliated unit, integrate them with open source datasets, and screen and organize the images.
[0065] Use high-definition cameras to capture images of various electronic components in electronic laboratories to build a high-quality dataset. In this process, refer to existing open source datasets to ensure a more complete integration and expansion effect. At the same time, carefully screen the images to ensure the diversity and accuracy of the dataset. Through this series of steps, the aim is to create an electronic component image library containing rich information for research and development, and to promote further research and application in the field of electronics;
[0066] S2, using canny edge operator and fuzzy processing methods to pre-process the original image, and obtain 2426 data set images, which divide typical electronic components into five types: capacitor, inductor, light-emitting diode, resistor, and transistor (Capacitor, Inductor, Led, Resistor, Transistor);
[0067] S3. Label the dataset using Labelimg. Label the collected electronic device images. The defects are divided into five categories: capacitor, inductor, light-emitting diode, resistor, and transistor. The capacitor is "Capacitor", the inductor is "Inductor", the light-emitting diode is "Led", the resistor is "Resistor", and the transistor is "Transistor". The dataset is divided according to a certain ratio, and finally the electronic component dataset is obtained. It is divided into training set, validation set, and test set, of which there are 2103 training sets, 226 validation sets, and 97 test sets.
[0068] S4. Add the efficient multi-scale attention module (EMA network module) for cross-space learning to the 10th layer of the YOLOv10s neck network (the first ordinary convolution module - Conv module is recorded as the 0th layer, and so on), as follows Figure 1 shown.
[0069] S5. Add the core module of the VanillaNet network, the VanillaBlock module, to the 17th, 20th, and 23rd layers of the neck network. Figure 1 As shown in the figure, the core module VanillaBlock in the VanillaNet network module is added to the 16th, 19th, and 22nd layers of the network to compare the accuracy of the model.
[0070] S6. Use the Shape-IoU border loss function as the new border loss function of YOLOv10s to replace the original CIoU border loss function and build an improved YOLOv10s network structure;
[0071] S7. Perform ablation tests on each improved model. The test contents are as follows: 1) Add the EMA multi-scale feature fusion module only in the 10th layer. 2) Add the EMA module, and add the VanillaBlock module to the 16th, 19th, and 22nd layers of the neck. 3) Add the EMA module, and add the VanillaBlock module to the 17th, 20th, and 23rd layers of the neck.
[0072] S8. Input the training set into the improved YOLOv10 model, combine the model to identify the acquired electronic device images, and use precision (Precision), recall rate (Recall), mean average precision (mAP) and frames per second (FPS) as evaluation indicators.
[0073] S9. Use the trained model best.pt to perform end-to-end accurate prediction and detection on images, and deeply analyze the advantages of the model for better promotion.
[0074] The improved YOLOv10 model includes Backbone part (backbone network part), feature fusion Neck part (neck network) and Head part (detection head part). The Backbone part is the part of the model used to extract image features, which includes ordinary convolution module (Conv convolution module), EMA multi-scale feature fusion module, SCDown module and SPPF module; the Neck part is composed of upsampling module, splicing module, C2f module and CSDown module, C2fCIB module, VanillaBlock module and ordinary convolution module (Conv module). The Head part is composed of three end-to-end detection heads (v10detect modules).
[0075] The improved YOLOv10 model includes Backbone module, feature fusion Neck module, and detection head module. Backbone is the part of this model used to extract image features, including GSConv module, SCDown module, SPPF module, SE attention mechanism module, and C2f module; Feature fusion Neck is the part of this model used to connect the Backbone module and the detection head module, including ordinary convolution module, Upsample layer, C2f module, and C2fCIB module; the Upsample layer is used for upsampling; the detection head module is the part of this model used to predict the target bounding box and category probability, including the Shape-IoU border loss function.
[0076] The structure of this model is as follows: initialize the input image input to the input of the first Conv module, the output of the first Conv module is connected to the input of the second Conv module, the output of the second Conv module is connected to the input of the first C2f module, the output of the first C2f module is connected to the input of the third Conv module, the output of the third Conv module is connected to the input of the second C2f module, the output of the second C2f module is connected to the input of the first CSDown module, the output of the first CSDown module is connected to the input of the third C2f module, the output of the third C2f module is connected to the input of the second SCDown module, the output of the second SCDown module is connected to the input of the first C2fCIB module, the output of the first C2fCIB module is connected to the input of the SPPF module, the output of the SPPF module is connected to the input of the EMA module, the output of the EMA module is connected to the input of the first Upsample module, the output of the first Upsample module is connected to the input of the first Concat module, the output of the first Concat module is connected to the input of the fourth C2f module, the output of the fourth C2f module is connected to the input of the second Upsample module, and the second Upsamp The output of the 1e module is connected to the input of the second Concat module, the output of the second Concat module is connected to the input of the fifth C2f module, the output of the fifth C2f module is connected to the input of the first VanillaBlock module, the output of the first VanillaBlock module is connected to the input of the fourth Conv module and the input of the first v10Detect module, the output of the fourth Conv module is connected to the input of the third Concat module, the output of the third Concat module is connected to the input of the sixth C2f module, the output of the sixth C2f module is connected to the input of the second VanillaBlock module, the output of the second VanillaBlock module is connected to the input of the second v10Detect module and the input of the third CSDown module, the output of the third CSDown module is connected to the input of the fourth Concat module, the output of the fourth Concat module is connected to the input of the second C2fCIB module, the output of the second C2fCIB module is connected to the input of the third VanillaBlock module, and the output of the third VanillaBlock module is connected to the input of the third v10Detect module. The output of the first v10Detect module, the output of the second v10Detect module, and the output of the third v10Detect module constitute the output of the total feature.
[0077] Specifically, step S1 includes:
[0078] S1.1: Use a high-definition camera to photograph five types of basic electronic components in the electronics laboratory to obtain the original data set.
[0079] S1.2: Select the best open source images and put them in a folder together with the original dataset.
[0080] Specifically, step S2 includes:
[0081] The data set was preprocessed to reduce the inaccuracy and randomness of some data set images. The data preprocessing methods included blur processing and canny edge operator, and finally 2426 images were obtained.
[0082] Specifically, step S3 includes:
[0083] S3.1: Labeling is performed in the Anaconda environment. The labelimg image annotation tool is used to annotate the collected electronic device classification images. The defects are divided into five categories: capacitors, inductors, light-emitting diodes, resistors, and transistors. The capacitor components are labeled as "Capacitor", the inductor components are labeled as "Inductor", the light-emitting diode is "Led", the resistor is "Resistor", and the transistor is "Transistor". The PascalVOC format data label (XML tag) is obtained, and then it is converted into a txt file suitable for training through the data format conversion code written in Python. The file includes most of the data, including the x-coordinate, y-coordinate, width, height and other annotation information of the center point of the bounding box.
[0084] S3.2: After the labeling work is completed, the data set is divided into a training set, a validation set, and a test set of the electronic component image classification detection data set, and 2103 training sets, 226 validation sets, and 97 test sets are obtained;
[0085] Specifically, step S4 includes:
[0086] S4: Add the EMA multi-scale attention module to the YOLOv10s backbone network and add it to the end of the backbone network (layer 10).
[0087] The EMA cross-space learning multi-scale attention module and the backbone network fusion can simultaneously capture the channel and spatial information of electronic device images, and effectively improve the feature representation capabilities of different electronic devices without adding too many parameters and computational costs. The EMA module reshapes some channels into batch dimensions and groups the channel dimensions into multiple sub-features, so that the spatial semantic features are well distributed within each feature group, improving the feature expression capabilities.
[0088] The core feature of the multi-scale attention mechanism of EMA cross-space learning is that it efficiently aggregates multi-scale information of electronic devices and improves the ability of feature extraction through cross-space learning methods. This mechanism reshapes part of the channel dimension into a batch dimension and groups the channel dimension into multiple sub-features to retain information on each channel and reduce computational overhead. The EMA module recalibrates the channel weights in each parallel branch by encoding global information and captures the pixel-level relationship of the electronic device image through cross-dimensional interactions.
[0089] Specifically, the operation process of EMA multi-scale attention includes the following steps:
[0090] 1) Grouping: The channels of the input feature map are divided into multiple groups to reduce the complex calculations generated by identifying electronic components. Assume that the input feature map has a shape of C×H×W, where C is the number of channels, and H and W are the height and width of the feature map. Through grouping, the feature map is divided into G groups, and the size of each group of feature maps is C / G×H×W.
[0091] 2) Global pooling in the X and Y directions: The electronic device feature map is average pooled in the X (width) and Y (height) directions, respectively, to obtain two feature maps of sizes C / G×H×1 and C / G×1×W, representing the global features 1 in the X and Y directions.
[0092] 3) Concat+Conv(1×1): After concatenating the two electronic device feature maps, feature fusion is performed through 1×1 convolution to output a new feature map with a shape of C / G×H×W to further capture the cross-spatial information of the electronic component feature map.
[0093] 4) Sigmoid reweighting: A weight coefficient is generated through the Sigmoid function to re-weight the feature map, enhancing the areas with important features of typical electronic devices while suppressing unimportant areas.
[0094] 5) Cross-spatial Learning: Normalize each set of electronic device feature maps to keep each set of features at the same scale. Generate a cross-spatial attention matrix through global average pooling and softmax function, perform matrix multiplication for interactive learning, and finally restore the output shape to a C×H×W feature map.
[0095] The benefits of EMA multi-scale attention are mainly reflected in the following aspects:
[0096] 1) Improve efficiency: Through grouping and cross-space learning, the computational overhead of electronic device image processing is reduced and the operating efficiency of the model is improved.
[0097] 2) Enhanced feature representation: The EMA module can better capture the multi-scale features of electronic devices and enhance the model’s ability to distinguish electronic devices of different sizes and shapes.
[0098] 3) Optimize feature extraction: Through reweighting and cross-space learning, EMA can more accurately extract and utilize electronic device feature information, improving the performance of the model.
[0099] The structure of the EMA multi-scale feature fusion module is as follows: the input feature information is connected to the input of the Groups module, the output of the Groups module is connected to the input of the first Re-weight module, the input of the second Re-weight module, the input of the first X AvgPool module, the input of the Y Avg Pool module and the input of the Conv module, the output of the first X Avg Pool module and the output of the Y Avg Pool module are connected to the input of the second X Avg Pool module, the output of the second X Avg Pool module is connected to the input of the first Sigmoid module and the input of the second Sigmoid module, the output of the first Sigmoid module and the output of the second Sigmoid module are connected to the input of the second Re-weight module, the output of the second Re-weight module is connected to the input of the GroupNorm module, the output of the GroupNorm module is connected to the input of the first Avg Pool module and the input of the first Matmul module, the output of the first Avg Pool module is connected to the input of the first Softmax module, the output of the first Softmax module is connected to the input of the second Matmul module, the output of the Conv module is connected to the input of the first Matmul module and the input of the second AvgPool module, and the output of the second Avg The output of the Pool module is connected to the input of the second Softmax module, the output of the second Softmax module and the output of the GroupNorm module are connected to the input of the first Matmul module, the output of the first Matmul module and the output of the second Matmul module are connected to the input of the ☉ module, the output of the ☉ module is connected to the input of the third Sigmiod module, the output of the third Sigmiod module is connected to the input of the second Re-weight module, and the output of the second Re-weight module completes the whole process.
[0100] S5: The core module of VanillaNet network, VanillaBlock, is embedded into the 17th, 20th, and 23rd layers of the neck part of the model to evaluate the quality of the model. A reference experiment is also set up: VanillaBlock is embedded into the 16th, 19th, and 22nd layers of the neck part of the model. The reliability of the model is verified through comparative experiments.
[0101] The architecture of VanillaNet consists of three main parts: stem block, main body and fully connected layers. The main body usually consists of four stages, each of which consists of stacked identical blocks. After each stage, the channel of features is expanded, while the height and width are reduced. VanillaNet avoids excessive depth, shortcuts and complex operations such as self-attention mechanism, making the network structure concise and powerful. Each layer is carefully designed to be compact and intuitive, and the nonlinear activation function is pruned after training to restore the original architecture. The main function of VanillaBlock, the core module of VanillaNet network, is to improve the accuracy and efficiency of electronic device classification. VanillaBlock consists of a series of stacked Vanilla layers, each of which contains some convolution and pooling operations to further extract and integrate electronic device feature information. Through this design, VanillaBlock is able to improve the accuracy of electronic device classification detection while maintaining efficient performance.
[0102] The structure of the VanillaBlock module is as follows: the output of the module with a channel number of 224×224×3 is connected to the input of the first convolution kernel module, the output of the first convolution kernel module is connected to the input of the module with a channel number of 56×56×512, the output of the module with a channel number of 56×56×512 is connected to the input of the second convolution kernel module, the output of the second convolution kernel module is connected to the input of the module with a channel number of 28×28×1024, the output of the module with a channel number of 28×28×1024 is connected to the input of the third convolution kernel module, the output of the third convolution kernel module is connected to the input of the module with a channel number of 14×14×2048, the output of the module with a channel number of 14×14×2048 is connected to the input of the fourth convolution kernel module, the output of the fourth convolution kernel module is connected to the input of the module with a channel number of 4096, the output of the module with a channel number of 4096 is connected to the input of the module with a channel number of 1000, and the output of the module with a channel number of 1000 is the output of the whole process.
[0103] Specifically, VanillaBlock achieves its role in the following ways:
[0104] VanillaBlock has efficient computing performance. VanillaBlock uses lightweight convolution modules and optimized loss functions to reduce the amount of calculation and the number of parameters generated by the electronic device classification algorithm, thereby improving computing efficiency. In addition, VanillaBlock has excellent feature extraction capabilities. VanillaBlock enhances the nonlinearity and feature expression capabilities of the electronic device classification algorithm network by serializing activation functions and multi-layer convolution operations, enabling the electronic device classification algorithm model to better identify and classify targets. VanillaNet uses a multi-level pyramid structure, and through multi-level pyramid feature fusion, the features of electronic device images with different resolutions are effectively fused, which not only maintains the detailed information of different electronic devices, but also improves the global understanding ability.
[0105] VanillaNet removes the complex residual and attention modules and achieves efficient electronic device classification and detection performance through simple convolution calculations. This design not only simplifies the network structure, but also reduces computing resource consumption, making the model easier to deploy and maintain.
[0106] S6: Use the Shape-IoU loss function as the border loss function of the model proposed in the present invention.
[0107] The Shape-IoU border loss function is used as the new border loss function of YOLOv10s to replace the original CIoU border loss function.
[0108] The main idea of Shape-IoU is that the bounding box regression loss is an important part of the positioning branch of the electronic device classification detection detector, and most calculation methods focus on the geometric relationship between the GT box and the predicted box. By using the relative position and relative shape between the bounding boxes to calculate the loss, the inherent attributes of the bounding box itself, such as its shape and scale, are not considered. The Shape-IoU method can calculate the loss by focusing on the shape and scale of the bounding box itself, making the bounding box regression more accurate. The definition of Shape-IoU is as follows:
[0109]
[0110]
[0111]
[0112]
[0113]
[0114]
[0115]
[0116] Where wgt, hgt represent the width and height of the real bounding box, respectively, v represents the minimum closure area of the predicted box and the real box, (xc, yc), Represent the center point coordinates of the predicted box and the real box respectively, t is the scale factor, the value range is between [0, 1.5], ww and wh are the weight coefficients of the real box in the horizontal and vertical directions respectively, and their values are positively correlated with the width and height of the real box. c2 represents the relative distance between the centers of the two. ww and hh are the weight coefficients in the horizontal and vertical directions respectively, and their values are related to the shape of the GT box. scale is the scale factor, and its value is related to the scale of the target in the electronic device dataset. For smaller targets, the larger the scale value. The corresponding bounding box regression loss is as follows:
[0117] L Shape-IoU =1-IoU+distance shape +0.5×Ω shape (8)
[0118] For small target datasets such as electronic components, the scale value is 1.5. Compared with CIoU, Shape_IoU can better consider the impact of the inherent attributes of the bounding box, such as its own shape and scale, on the bounding box regression, and can solve the problem that for regression samples with the same shape, when the regression sample deviation and shape deviation are the same and not all 0, the IoU value of the smaller regression sample is more significantly affected by the shape of the real annotation box than the larger regression sample, thereby improving the model performance and better completing the task of electronic device classification and detection.
[0119] Shape-IoU has the following advantages:
[0120] 1) Considering the shape and scale of the bounding box: Existing bounding box regression methods usually only consider the geometric relationship between the GT box and the predicted box, ignoring the impact of the shape and scale of the bounding box itself on the regression results. Shape-IoU makes the regression more accurate by introducing the shape and scale factors of the bounding box.
[0121] 2) Improve detection performance: Shape-IoU has been shown to achieve advanced performance in different detection tasks, especially when dealing with small object detection.
[0122] In summary, Shape-IoU significantly improves the accuracy of object detection by considering the shape and scale of the bounding box, especially in the small object detection task, and is an effective improvement method.
[0123] Specifically, step S7 includes:
[0124] S7: The electronic device image data set obtained in S3 is put into the improved YOLOv10 model reconstructed by the present invention for training, and the number of training times is 200. The training set is applied to the improved YOLOv10 model for training, and finally the detection results are generated and the best.pt file is obtained. The original model is trained and analyzed by comparing the comprehensive performance indicators.
[0125] Specifically, step S8 includes:
[0126] S8: In order to evaluate the model, several key indicators are used. The improved YOLOv10s model constructed by the present invention is evaluated by using recall (Recall, Rec), mean average precision (mean average precision, mAP) and frames per second (frames per second, FPS). Among them, the recall (recall, Rec) reflects the ratio between the number of relevant documents successfully retrieved and the number of relevant documents missed, specifically, the proportion of actual relevant documents retrieved. The mean average precision (mean average precision, mAP) is a summary of the average precision (AP) values of all categories. The number of frames processed per second (frames per second, FPS) refers to the number of image frames that can be processed per second during the video processing process. The higher the FPS, the smoother the video performance.
[0127] Specifically, step S9 includes:
[0128] The trained model is used to identify and locate images of electronic components collected on-site in the electronic laboratory of the affiliated unit, and can be accurately classified.
[0129] Through the ablation experiment shown in Table 1, it can be seen that the EMA attention mechanism module is added to the trunk and the VanillaBlock modules are added to the 17th, 20th, 23rd layers and the 16th, 19th, and 22nd layers of the neck respectively. Compared with the baseline model and other improved methods, the model of the present invention is much higher than other detection models in terms of accuracy and speed, and can meet the requirements of real-time classification detection.
[0130] The following is a comparison table of the results with the original model
[0131] Table 1 Comparative test results
[0132]
[0133] The platform used for the test of the present invention is Pycharm11.0.15, the model framework is Pytorch1.13.1, the CPU is 15vCPU Intel(R)Xeon(R)Platinum 8358P CPU@2.60GHz, and the GPU is RTX3090, 24GB.
[0134] When setting hyperparameters, the number of training rounds was set to 200, the optimizer used was SGD, and the batch size was set to 32.
[0135] In summary, the improved YOLOv10 model constructed by the present invention can bring significant advantages and significance to the application in the identification of electronic devices. First, compared with the previous generation model, YOLOv10 has higher detection accuracy and faster processing speed, and can identify multiple electronic devices in real time, which is crucial for modern manufacturing and quality control. In real-time monitoring on the production line, this model can quickly classify and reduce human errors, thereby improving the overall quality and reliability of the product. Secondly, the multi-target detection capability of this model enables it to work efficiently in complex environments, greatly enhancing its flexibility in practical applications. In addition, YOLO v10 supports end-to-end detection, speeds up the detection speed, and can quickly train the model according to the characteristics of specific electronic devices, reducing development costs and time. In summary, this model not only improves the recognition efficiency of electronic devices, but also provides strong technical support for the automation and intelligent development of the industry, thereby ensuring the high standards and high quality of the products while improving production efficiency, and has important economic and social significance.
Claims
1. A method for electronic device classification and detection based on multi-scale feature fusion, characterized in that: The following steps are involved: Step S1: obtaining original images of various common and typical electronic devices, integrating them with open source data sets, and screening and sorting the images; Step S2: pre-process the original image using the canny edge operator and blur processing method to obtain several data set images; Step S3: Use Labelimg to annotate the target data set, annotate the collected electronic device images, divide the data set according to a certain ratio, and finally obtain the electronic component data set, and divide the data set into a training set, a validation set, and a test set; Step S4: Obtain an improved YOLOv10 model, use the model to recognize the acquired electronic device image, and use precision Precision, recall rate Recall, mean average precision mAP and frames per second FPS as evaluation indicators.
2. The method according to claim 1, characterized in that In step S1, a high-definition camera is used to take original images of five common and typical electronic devices in the electronic laboratory of the affiliated unit, and the images are integrated with open source data sets, and the images are screened and sorted.
3. The method according to claim 1, characterized in that: In step S2, the obtained data set images are divided into five types: capacitor, inductor, light emitting diode, resistor, and transistor; In step S3, the collected electronic device images are annotated, and the defects are divided into five categories: capacitor, inductor, light-emitting diode, resistor, and transistor, where capacitor is "Capacitor", inductor is "Inductor", light-emitting diode is "Led", resistor is "Resistor", and transistor is "Transistor".
4. The method according to claim 1, characterized in that: In step S4, the improved YOLOv10 model obtained is specifically: Initialize the input image and input it to the first Conv module, the output of the first Conv module is connected to the input of the second Conv module, the output of the second Conv module is connected to the input of the first C2f module, the output of the first C2f module is connected to the input of the third Conv module, the output of the third Conv module is connected to the input of the second C2f module, the output of the second C2f module is connected to the input of the first CSDown module, the output of the first CSDown module is connected to the input of the third C2f module, the output of the third C2f module is connected to the input of the second SCDown module, the output of the second SCDown module is connected to the input of the first C2fCIB module, the output of the first C2fCIB module is connected to the input of the SPPF module, and the output of the SPPF module is connected to the input of the EMA module; The output of the EMA module is connected to the input of the first Upsample module and the input of the third VanillaBlock module, the output of the first Upsample module is connected to the input of the first Concat module, the output of the first Concat module is connected to the input of the fourth C2f module, the output of the fourth C2f module is connected to the input of the second Upsample module, the output of the second Upsample module is connected to the input of the second Concat module, the output of the second Concat module is connected to the input of the fifth C2f module, and the output of the fifth C2f module is connected to the input of the first VanillaBlock module; The output of the first VanillaBlock module is connected to the input of the fourth Conv module and the input of the first v10Detect module, the output of the fourth Conv module is connected to the input of the third Concat module, the output of the third Concat module is connected to the input of the sixth C2f module, the output of the sixth C2f module is connected to the input of the second VanillaBlock module, and the output of the second VanillaBlock module is connected to the input of the second v10Detect module and the input of the third CSDown module; The output of the third CSDown module is connected to the input of the fourth Concat module, the output of the fourth Concat module is connected to the input of the second C2fCIB module, the output of the second C2fCIB module is connected to the input of the third VanillaBlock module, and the output of the third VanillaBlock module is connected to the input of the third v10Detect module; The output of the first v10Detect module, the output of the second v10Detect module, and the output of the third v10Detect module constitute the output of the total feature.
5. The method according to claim 4, characterized in that The EMA module is specifically: The input features are fed into the Groups module, and the output of the Groups module is connected to the input of the first Re-weight module, the input of the second Re-weight module, the input of the first X Avg Pool module, the input of the Y Avg Pool module, and the input of the Conv module; The output of the first X Avg Pool module and the output of the Y Avg Pool module are connected to the input of the second X Avg Pool module, the output of the second X Avg Pool module is connected to the input of the first Sigmoid module and the input of the second Sigmoid module, the output of the first Sigmoid module and the output of the second Sigmoid module are connected to the input of the second Re-weight module, the output of the second Re-weight module is connected to the input of the GroupNorm module, and the output of the GroupNorm module is connected to the input of the first AvgPool module and the input of the first Matmul module; The output of the first Avg Pool module is connected to the input of the first Softmax module, and the output of the first Softmax module is connected to the input of the second Matmul module; The output of the Conv module is connected to the input of the first Matmul module and the input of the second Avg Pool module, the output of the second AvgPool module is connected to the input of the second Softmax module, the output of the second Softmax module and the output of the GroupNorm module are connected to the input of the first Matmul module, the output of the first Matmul module and the output of the second Matmul module are connected to the input of the splicing module, the output of the splicing module is connected to the input of the third Sigmiod module, the output of the third Sigmiod module is connected to the input of the second Re-weight module, and the output of the second Re-weight module is used to output the output features of the EMA module.
6. The method according to claim 4, characterized in that The VanillaBlock modules are: The input features of the VanillaBlock module are input to the 4×4 convolution module, the output of the 4×4 convolution module is connected to the input of the first 1×1 convolution module, the output of the first 1×1 convolution module is connected to the input of the second 1×1 convolution module, the output of the second 1×1 convolution module is connected to the input of the third 1×1 convolution module, the output of the third 1×1 convolution module is connected to the input of the global average pooling layer, the output of the global average pooling layer is connected to the input of the fully connected layer, and the output of the fully connected layer is used to output the output features of the VanillaBlock module.
7. The method according to any one of claims 4 to 6, characterized in that The Shape-IoU loss function is used as the border loss function of the improved YOLOv10 model; specifically: 1) Divide the intersection of the predicted box A and the true box B by the union of the two, as shown in Formula 1: 2) The influence of the shape and scale of the bounding box on the regression results is evaluated by calculating the width ww and height hh of the bounding box and the weight coefficient, specifically: 3) Pay attention to the shape and scale of the bounding box. The accuracy of object detection can be improved by considering these factors. Specifically: 4) By introducing Ω shape To make up for this deficiency, the loss function can more comprehensively consider the characteristics of the bounding box, specifically: where w gt ,h gt Represent the width and height of the real bounding box respectively, v represents the minimum closure area of the predicted box and the real box, (xc, yc), They represent the center coordinates of the predicted box and the real box respectively, t is the scale factor, ww and wh are the weight coefficients of the real box in the horizontal and vertical directions respectively, and their values are positively correlated with the width and height of the real box, c 2 represents the relative distance between the centers of the two. ww and hh are the weight coefficients in the horizontal and vertical directions respectively, and their values are related to the shape of the GT box. scale is the scale factor, and its value is related to the scale of the target in the electronic device dataset. For smaller targets, the larger the scale value, the corresponding bounding box regression loss is as follows: L Shape-IoU =1-IoU+distance shape +0.5×Ω shape (8)。
Citation Information
Cited By
Lightweight silkworm cocoon classification detection algorithm based on improved YOLOv8 network
CN120563950A
A lightweight silkworm cocoon classification and detection method based on improved YOLOv8 network
CN120563950B
Bionic digging and sorting device for potatoes
CN121128419A