Method and device for appearance inspection of electronic control components based on deep learning
By constructing a feature pyramid and optimizing the deep learning method of feature representation, the problems of low accuracy and high missed detection rate in traditional electronic control component appearance inspection are solved, and high-precision and low missed detection rate appearance inspection of electronic control components is achieved.
Patent Information
- Application Number
- CN202510689116.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-05-27
AI Technical Summary
Traditional appearance inspection methods for electronic control components have low detection accuracy and high missed detection rates when the background texture is complex or the defect is highly similar to the background. They also have limited adaptability to different defect scales and complex shapes, and cannot meet the high-precision requirements of industrial production.
A deep learning-based appearance inspection method for electronic control components is adopted. By constructing a feature pyramid, dimension importance evaluation and regional significance analysis are performed. Combined with multiple pooling processing and skip connections, feature representation is optimized to build an appearance inspection model.
It significantly improves detection accuracy and robustness, reduces missed detection and false detection rates, and can effectively handle defect detection of electronic control components under complex backgrounds.
Smart Images

Figure CN120198439B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of machine vision technology, and in particular to a method and device for detecting the appearance of electronic control components based on deep learning. Background Art
[0002] With the widespread application of electronic control components in modern industry, ensuring their appearance quality has become a crucial step in the production process. Traditional methods for electronic control component appearance inspection mostly rely on manual visual inspection or simple image processing techniques such as threshold segmentation and edge detection. While these methods can detect relatively obvious defects to a certain extent, due to their reliance on the difference between the background and the defect, they often suffer from low detection accuracy and high missed detection rates when the background texture is complex or highly similar to the defect. Existing machine vision inspection methods mostly use frameworks based on feature extraction and classification algorithms. However, these methods have difficulty effectively distinguishing defects from normal backgrounds when dealing with situations where the similarity between the background and the defect is high, resulting in unstable detection results. In addition, traditional inspection methods have limited adaptability to different defect scales and complex shapes, and cannot meet the high-precision requirements for electronic control component appearance inspection in industrial production. Summary of the Invention
[0003] The present application provides a method and device for appearance inspection of electronic control components based on deep learning, which are used to improve the detection accuracy of automated inspection of electronic control components.
[0004] In a first aspect, an embodiment of the present application provides a method for detecting the appearance of an electronically controlled component based on deep learning, the method comprising:
[0005] Obtaining a training image set and a verification image set of the electronic control component;
[0006] constructing a feature pyramid based on the training image set, performing dimension importance evaluation and regional saliency analysis on the feature pyramid to obtain enhanced feature representation;
[0007] Performing multiple pooling processes on the enhanced feature representation to obtain multiple groups of pooled feature representations;
[0008] Determining a high-order representation and a low-order representation according to each group of pooled feature representations, performing dynamic weight integration on the high-order representation and the low-order representation and introducing a skip connection to obtain an optimized feature representation;
[0009] Based on the optimized feature representation and the verification image set, an appearance detection model is obtained, and appearance detection of the electronic control component is performed using the appearance detection model.
[0010] In a second aspect, an embodiment of the present application provides an electronic control component appearance inspection device based on deep learning, the device comprising:
[0011] An image acquisition module, used to acquire a training image set and a verification image set of the electronic control component;
[0012] A feature extraction module is used to construct a feature pyramid based on the training image set, perform dimension importance evaluation and regional significance analysis on the feature pyramid, and obtain enhanced feature representation;
[0013] A feature pooling module, configured to perform multiple pooling processes on the enhanced feature representation to obtain multiple groups of pooled feature representations;
[0014] a feature optimization module, configured to determine a high-order representation and a low-order representation according to each group of the pooled feature representations, perform dynamic weight integration on the high-order representation and the low-order representation, and introduce skip connections to obtain an optimized feature representation;
[0015] A model generation module is used to obtain an appearance detection model based on the optimized feature representation and the verification image set, and perform appearance detection on the electronic control component through the appearance detection model.
[0016] In a third aspect, an embodiment of the present application provides an electronic device, the electronic device including a memory and a processor;
[0017] The memory is used to store computer programs;
[0018] The processor is used to execute the computer program and implement the deep learning-based appearance detection method for electronic control components as described in any one of the embodiments of the present application when executing the computer program.
[0019] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the processor implements the deep learning-based appearance detection method for electronic control components as described in any one of the embodiments of the present application.
[0020] An embodiment of the present application provides a method for appearance detection of electronic control components based on deep learning, the method comprising: obtaining a training image set and a verification image set of the electronic control component; constructing a feature pyramid based on the training image set, performing dimensional importance evaluation and regional significance analysis on the feature pyramid, and obtaining an enhanced feature representation; performing multiple pooling processes on the enhanced feature representation to obtain multiple groups of pooled feature representations; determining a high-order representation and a low-order representation based on each group of pooled feature representations, performing dynamic weight integration on the high-order representation and the low-order representation and introducing jump connections to obtain an optimized feature representation; obtaining an appearance detection model based on the optimized feature representation and the verification image set, and performing appearance detection on the electronic control component using the appearance detection model. In the above method, macro and micro image information is extracted by constructing a feature pyramid, and the feature representation of important areas is strengthened through dimensional importance evaluation and regional saliency analysis. After adopting multiple pooling processing, features can be extracted from different scales, and the adaptability to defects of different sizes and shapes is enhanced. The dynamic weight integration and jump connection of high-order and low-order features are then used to optimize the feature expression, so that the model can better maintain accuracy when processing complex backgrounds, integrate macro and micro features, significantly improve detection accuracy and robustness, and reduce missed detection and false detection rates. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0022] Figure 1 A schematic flow chart of a method for detecting the appearance of an electronically controlled component based on deep learning provided in an embodiment of the present application;
[0023] Figure 2 A schematic block diagram of a deep learning-based appearance inspection device for electronically controlled components provided in an embodiment of the present application. DETAILED DESCRIPTION
[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0025] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, combined, or partially merged, so the actual execution order may vary depending on the actual situation.
[0026] It should also be understood that the terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0027] It should be further understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0028] See also Figure 1 , Figure 1 This is a schematic flow chart of a method for detecting the appearance of an electric control component based on deep learning provided by an embodiment of the present application. Figure 1 As shown, the specific steps of the electric control component appearance detection method based on deep learning include: S101-S105.
[0029] S101: Acquire a training image set and a verification image set of an electronic control component.
[0030] For example, an industrial camera is used to capture multi-view images of electronic control components under standard lighting conditions, ensuring image clarity and contrast meet subsequent processing requirements. The original images undergo resolution normalization and resizing to 640×640 pixels. Brightness normalization and color correction are then performed to eliminate interference caused by ambient lighting variations. Data augmentation techniques are used to increase the diversity of training samples, including random rotation (±15°), horizontal / vertical flipping, brightness and contrast adjustment (±20%), and Gaussian noise addition. The enhanced images are professionally annotated, accurately labeling defect location coordinates, type, severity, and other information, generating a standard XML annotation file. The processed images are divided into a training set and a validation set in an 8:2 ratio. The training set is used for model learning, while the validation set is used to evaluate model performance. During dataset construction, a balanced distribution of defect samples is ensured, and rare defect types are oversampled to avoid biasing model training towards high-frequency defect types, thereby improving the model's ability to detect low-probability defects.
[0031] S102. Construct a feature pyramid based on the training image set, perform dimension importance evaluation and regional saliency analysis on the feature pyramid, and obtain enhanced feature representation.
[0032] For example, based on the improved YOLOv5 backbone network, a three-layer feature pyramid structure is constructed, corresponding to the 1 / 8, 1 / 16, and 1 / 32 scale layers, respectively, to capture the multi-scale defect features of electronic control components. A dual global image attention module is designed, which includes a channel attention mechanism and a spatial attention mechanism. In the channel attention part, the global average pooling and maximum pooling values are calculated for each feature map channel. After connection, a channel weight vector is generated through a fully connected layer. This weight vector quantifies the contribution of different channels to defect detection. The spatial attention part calculates the importance of each pixel in the spatial dimension of the feature map, generates a two-dimensional spatial attention map, and highlights the location information of the defect area. The dimension importance evaluation process uses adaptive parameters to automatically adjust the fusion ratio of channel and spatial attention to generate a weighted feature map. The cross-layer connection design realizes information interaction between high- and low-level features. The low-level layer retains rich edge, texture and other detailed information, while the high-level layer contains abstract semantic information. The two complement each other to improve the defect detection performance of electronic control components. After being enhanced by the global image attention module, the multi-scale enhanced feature representation effectively improves the network's ability to perceive small defects and distinguish defects that are highly integrated with the background, laying the foundation for subsequent feature optimization.
[0033] S103: Perform multiple pooling processes on the enhanced feature representation to obtain multiple groups of pooled feature representations.
[0034] For example, multiple pooling processes are performed on the enhanced feature representation to replace the traditional maximum pooling operation and reduce the loss of feature information. Three parallel pooling branches are designed, with pooling kernel sizes of 5×5, 9×9, and 13×13, respectively, to ensure coverage of the receptive fields of defects of different scales. A weighted average pooling strategy is adopted to avoid the loss of defect edge information caused by traditional maximum pooling, calculate the exponentially weighted average of the pixel values in the window, and retain more texture details. Independent pooling parameters are designed for each parallel pooling branch to achieve multi-scale expression of defect features. A trainable feature selection mechanism is used to automatically adjust the weight coefficients of different pooling branches and optimize the multi-scale feature fusion process. The features after parallel processing are normalized to eliminate the scale differences introduced by different pooling methods and ensure the consistency of feature distribution. Three sets of pooled feature representations are output, corresponding to the defect representation under different receptive fields, forming multiple sets of pooled feature representations, providing a diversified feature basis for subsequent feature optimization, and effectively dealing with various complex defects on the surface of electronic control components.
[0035] S104. Determine a high-order representation and a low-order representation based on each group of pooled feature representations, dynamically weight the high-order representation and the low-order representation, and introduce skip connections to obtain an optimized feature representation.
[0036] For example, each set of pooled feature representations is divided into a main path and a side path. The main path features are processed through a series of convolutional layers and nonlinear activation functions to extract deep semantic information and form a high-order representation. The side path retains the original feature information through identity mapping to form a low-order representation. 1×1 convolution is applied to the main path to reduce the number of channels, 3×3 convolution is applied to extract features, and then 1×1 convolution is applied to restore the number of channels, forming a "bottleneck" structure that reduces computational complexity while maintaining expressive power. A channel attention mechanism is designed for both high-order and low-order representations to calculate the importance of each channel feature and generate a dynamic weight coefficient matrix. Based on the weight coefficients, a weighted fusion of the high-order and low-order representations is performed to highlight key feature channels. A residual connection mechanism is introduced to perform element-wise addition of the fused features to the original pooled features, effectively alleviating the vanishing gradient problem and improving the stability of deep network training. During the feature fusion process, a gating mechanism is used to dynamically adjust the contribution ratio of high- and low-order features, adaptively adjust the feature representation strategy for different types of defects, and output high-quality optimized feature representation with rich defect texture information and semantic information, providing high-discrimination feature representation for subsequent detection, effectively improving the accuracy of defect detection of electronic control components and reducing the missed detection rate.
[0037] S105 : Based on the optimized feature representation and the verification image set, an appearance inspection model is obtained, and appearance inspection of the electronic control component is performed using the appearance inspection model.
[0038] Based on optimized feature representation, an appearance inspection network was constructed, with three detection branches designed to detect objects of different scales. An adaptive anchor box generation mechanism was designed to analyze the size distribution of defects in the training data and dynamically adjust the size and aspect ratio of the anchor boxes, improving the detection network's adaptability to defects of varying scales. An optimized loss function was designed, consisting of a bounding box regression loss, a classification loss, and a confidence loss. The bounding box regression loss employed CIOU loss to improve localization accuracy, while the classification loss employed focal loss to mitigate sample imbalance. The model was evaluated using a validation set to select optimal model parameters and calculate performance metrics such as mean average prediction (mAP), recall, and precision. A complete inspection pipeline was constructed for the actual inspection phase, including image preprocessing, model inference, and post-processing. Non-maximum suppression was used to merge overlapping detection boxes, and a confidence threshold filtering mechanism was designed to filter out low-confidence predictions, balancing detection accuracy and recall. A visualization module was developed for real-world production environments, displaying inspection results in real time, including defect location, type, and confidence level. Inspection result storage and statistical analysis capabilities were implemented to record historical data and generate analysis reports, providing data support for production line quality management. This step completes the entire process from optimized feature representation to actual defect detection, forming an end-to-end electronic control component appearance inspection system that meets the real-time inspection needs in modern electronic control production and achieves high-precision, low-missed detection results.
[0039] An embodiment of the present application provides a method for appearance detection of electronic control components based on deep learning, the method comprising: obtaining a training image set and a verification image set of the electronic control component; constructing a feature pyramid based on the training image set, performing dimensional importance evaluation and regional significance analysis on the feature pyramid, and obtaining an enhanced feature representation; performing multiple pooling processes on the enhanced feature representation to obtain multiple groups of pooled feature representations; determining a high-order representation and a low-order representation based on each group of pooled feature representations, performing dynamic weight integration on the high-order representation and the low-order representation and introducing jump connections to obtain an optimized feature representation; obtaining an appearance detection model based on the optimized feature representation and the verification image set, and performing appearance detection on the electronic control component using the appearance detection model. In the above method, macro and micro image information is extracted by constructing a feature pyramid, and the feature representation of important areas is strengthened through dimensional importance evaluation and regional saliency analysis. After adopting multiple pooling processing, features can be extracted from different scales, and the adaptability to defects of different sizes and shapes is enhanced. The dynamic weight integration and jump connection of high-order and low-order features are then used to optimize the feature expression, so that the model can better maintain accuracy when processing complex backgrounds, integrate macro and micro features, significantly improve detection accuracy and robustness, and reduce missed detection and false detection rates.
[0040] In order to more clearly introduce the technical solution of the present application, the technical solution of the present application will be introduced through specific embodiments below. It should be noted that the specific embodiments are used to expand the technical solution of the present application, but are not intended to limit the present application.
[0041] In some embodiments, a feature pyramid is constructed based on a training image set, and dimension importance evaluation and regional saliency analysis are performed on the feature pyramid to obtain an enhanced feature representation, including: S1021-S1025.
[0042] S1021. Construct a feature pyramid having a first scale layer, a second scale layer, and a third scale layer according to the training image set.
[0043] For example, a modified YOLOv5 backbone network structure is used as the basis for feature extraction, with the training image set fed into the network for forward propagation. The first scaling layer is set to 1 / 8 the input image resolution, with a feature map size of 80×80 pixels and 128 channels. This layer primarily captures subtle surface defects on electronic control components. The second scaling layer is set to 1 / 16 the input image resolution, with a feature map size of 40×40 pixels and 256 channels, focusing on detecting medium-sized defects. The third scaling layer is set to 1 / 32 the input image resolution, with a feature map size of 20×20 pixels and 512 channels, for identifying large defects. Upsampling and downsampling operations establish information flow paths between layers, ensuring the complementary fusion of multi-scale features.
[0044] S1022. Add a channel attention mechanism to each scale layer of the feature pyramid to obtain the channel feature representation of each scale layer.
[0045] Exemplarily, a channel attention module is designed for each scale layer of the feature pyramid, and adaptive weights are assigned to each channel of the feature map. First, global average pooling and global maximum pooling operations are applied to the input feature map to obtain description vectors of two channel dimensions, with the length of each vector equal to the number of channels of the input feature map. These two vectors are concatenated and processed through a two-layer fully connected network. The number of neurons in the middle layer is 1 / 16 of the number of input channels, and the ReLU activation function is used. The last layer uses the Sigmoid activation function to output normalized channel weight coefficients with a value range of 0 to 1. The weight coefficients are multiplied by each channel of the original feature map to obtain a weighted channel feature representation, which can enhance key channels and suppress invalid channels.
[0046] S1023. Calculate the spatial dependency of the channel feature representation to obtain a spatial feature representation.
[0047] For example, a spatial attention module is constructed based on the channel feature representation to capture global dependencies in the spatial dimension of the feature map. The maximum and average values are calculated along the channel dimension, respectively, to obtain two two-dimensional spatial feature maps with the same height and width as the original feature map. After concatenating these two feature maps in the channel dimension, they are processed through a 7×7 convolutional layer with an output channel of 1 and a padding setting of 3 to maintain the spatial dimension unchanged. The convolution output is mapped to the range of 0-1 using a Sigmoid function to obtain a spatial attention weight map, where higher weight values indicate more important locations. The spatial attention weight map is element-wise multiplied with the original channel feature representation to obtain a feature representation weighted in the spatial dimension, which effectively highlights defect areas and suppresses background interference.
[0048] S1024: Adaptively fuse the channel feature representation and the spatial feature representation to generate an enhanced feature representation for each scale layer.
[0049] For example, an adaptive fusion module is designed to integrate the advantages of channel feature representation and spatial feature representation. First, the weight ratio of the two feature representations is controlled by a trainable fusion parameter alpha, and the parameter value range is 0 to 1. The fusion process adopts a weighted summation method: enhanced feature representation = alpha × channel feature representation + (1-alpha) × spatial feature representation. The fusion parameter alpha is automatically learned through backpropagation. The alpha values of different scale layers may be different to adapt to the expression requirements of features at different levels. In addition, a residual connection is introduced in the fusion process to add the original features to the enhanced features to ensure smooth information flow and prevent the loss of useful features. The generated enhanced feature representation retains important information in both the channel dimension and the spatial dimension, improving the expressive power of the features.
[0050] S1025. Perform cross-layer connections on the enhanced feature representation to obtain a strengthened feature representation.
[0051] For example, cross-layer connections are constructed between layers of different scales to achieve interaction and complementarity of multi-level feature information. Using a feature pyramid network (FPN) structure, high-level features (third scale layer) are upsampled by a factor of 2 after adjusting the number of channels through 1×1 convolution, and then element-wise fused with middle-level features (second scale layer); similarly, the fused middle-level features are upsampled again and fused with low-level features (first scale layer). At the same time, a bottom-up path is implemented, where low-level features are downsampled through 3×3 convolution with a stride of 2 and then fused with middle-level features; the fused middle-level features are downsampled again and fused with high-level features, forming a feature network with bidirectional information flow. The low-level features retain detailed texture information, while the high-level features contain abstract semantic information. Cross-layer connections ensure that this complementary information is fully utilized, outputting an enhanced and strengthened feature representation.
[0052] This approach utilizes a dual attention mechanism (channel attention and spatial attention) to highlight key features and suppress irrelevant background. An adaptive fusion mechanism achieves the optimal combination of the two attentional approaches, and cross-layer connections ensure the complementary fusion of multi-scale features. This design significantly improves the network's ability to perceive small defects on the surfaces of electronically controlled components and enhances its ability to distinguish defects that are highly integrated with the background. This provides high-quality feature representation for subsequent detection, thereby improving overall detection accuracy and reducing missed detection rates.
[0053] In some embodiments, the enhanced feature representation is subjected to multiple pooling processes to obtain multiple groups of pooled feature representations, including: S1031-S1032.
[0054] S1031. Construct three parallel soft pooling branches, and set a first pooling core, a second pooling core, and a third pooling core respectively. The size of the first pooling core is 5×5, the size of the second pooling core is 9×9, and the size of the third pooling core is 13×13.
[0055] S1032. Pass the enhanced feature representation through the first pooling core, the second pooling core, and the third pooling core respectively to obtain a first pooling feature representation, a second pooling feature representation, and a third pooling feature representation.
[0056] For example, soft pooling is an adaptive pooling method that is more flexible than traditional maximum pooling and average pooling, and can dynamically assign weights based on feature importance. The three pooling branches use pooling kernels of 5×5, 9×9, and 13×13 sizes, respectively, to form feature representations with different receptive fields. The 5×5 pooling kernel focuses on local detail features, the 9×9 pooling kernel captures medium-scale texture information, and the 13×13 pooling kernel extracts large-scale semantic features. Through parallel processing, each pooling kernel independently processes the enhanced feature representation, generating first, second, and third pooled feature representations with different scale perception capabilities. By setting soft pooling kernels of different sizes for parallel feature extraction, effective multi-scale feature fusion is achieved, enhancing the model's adaptability to objects of different sizes. The adaptive nature of soft pooling ensures the preservation of important feature information while reducing redundant information, improving the discriminability and robustness of the feature representation.
[0057] In some embodiments, a high-order representation and a low-order representation are determined based on each group of pooled feature representations, and the high-order representation and the low-order representation are dynamically weighted and skip connections are introduced to obtain an optimized feature representation, including: S1041-S1046.
[0058] S1041 : Perform path separation on the first pooled feature representation, the second pooled feature representation, and the third pooled feature representation, respectively, and divide them into a main path and a bypass path.
[0059] For example, a dual-path separation strategy is adopted for each set of pooled feature representations, dividing the feature flow into two independent information processing channels: the main path and the bypass path. The main path is responsible for extracting deep semantic features and maintaining the main information flow of the features; the bypass path maintains the integrity of the original features and serves as an information supplement channel. In specific implementation, the first pooled feature representation, the second pooled feature representation, and the third pooled feature representation are copied into two copies, one of which is input into the main path for deep feature extraction, and the other is input into the bypass path to maintain the original expression of the features. This dual-path structural design can achieve deep feature extraction while maintaining the integrity of the original features.
[0060] S1042. Perform convolution transformation and nonlinear activation on the pooled feature representation of the trunk path to obtain a trunk feature representation.
[0061] For example, the backbone path uses multiple layers of convolutional transformations and nonlinear activation operations to extract deep features. First, a 3×3 convolution kernel is used to extract spatial features from the feature map, with a stride of 1 and padding using the SAME mode to maintain the feature map size. The convolutional layer is followed by BatchNormalization for feature normalization. The ReLU nonlinear activation function is then used to increase the nonlinearity of the feature representation, with an activation threshold of 0.
[0062] S1043. Perform identity mapping on the pooled feature representation of the bypass path to obtain a bypass feature representation.
[0063] For example, the bypass path uses identity mapping to maintain the integrity of the original features. Identity mapping directly copies the input features to the output without any transformation, maintaining identical feature dimensions and values. This direct feature transfer mechanism avoids information loss associated with deep networks and provides pristine and complete feature information for subsequent low-level feature extraction. The bypass feature representation preserves the fundamental feature patterns of the pooled features, effectively complementing the backbone features to construct a multi-scale feature representation.
[0064] S1044. Extract high-order representation and low-order representation according to the backbone feature representation and the bypass feature representation, respectively, wherein the backbone feature representation is processed by a densely connected convolutional network to obtain a high-order representation, and the bypass feature representation is processed by linear projection to obtain a low-order representation.
[0065] For example, the backbone feature representation is processed using a densely connected convolutional network. The network consists of multiple dense blocks, each of which uses a cross-layer feature reuse mechanism. The features of the previous layer are directly passed to all subsequent layers through connections. In specific implementation, each dense block contains 4 convolutional layers with a growth rate of 32. Intermediate features are fused through channel splicing to obtain a high-level representation. For the bypass feature representation, a 1×1 convolution is used for linear projection to adjust the feature dimension while maintaining the basic expression of the feature to obtain a low-level representation.
[0066] S1045. Calculate inter-channel adaptive coefficients for the high-order representation and the low-order representation to generate a weighted coefficient matrix, and dynamically integrate the high-order representation and the low-order representation according to the weighted coefficient matrix to obtain an integrated feature representation.
[0067] For example, a channel attention module is designed to calculate the channel importance between high-order and low-order representations. First, a global average pooling is performed on the features to obtain channel descriptors. The correlation between channels is learned using a two-layer fully connected network. The number of neurons in the middle layer is set to 1 / 16 of the number of channels, and a ReLU activation function is used. Finally, the output is normalized to the range of 0-1 using a Sigmoid function to obtain a weighted coefficient matrix. Based on this matrix, a weighted summation of the high-order and low-order representations is performed to achieve adaptive feature fusion, resulting in an integrated feature representation that balances deep semantic information and basic visual features.
[0068] S1046. Establish a skip connection between the integrated feature representation and the corresponding pooled feature representation, and obtain an optimized feature representation with a multi-scale receptive field through element-level addition operation.
[0069] For example, a skip connection mechanism is used to complement the information of the integrated feature representation with the original pooled feature representation. Specifically, the integrated feature representation is element-wise added to the pooled feature representation of the corresponding scale to ensure feature dimensionality matching. Through skip connections, detailed features from lower layers can be directly transferred to higher layers, complementing the semantic features and obtaining an optimized feature representation that combines both local details and global semantic information.
[0070] In this technical solution, the separation of the main and bypass paths ensures deep feature extraction without losing original information. The dual feature extraction mechanism of dense connections and linear projections achieves the complementarity of high-order semantic features and low-order basic features. Adaptive feature fusion is achieved through a channel-wise attention mechanism, and skip connections are introduced to establish multi-scale feature connections. The resulting optimized feature representation has a rich multi-scale receptive field, capable of simultaneously expressing local details and global semantic information, significantly improving feature expression capabilities.
[0071] In some embodiments, a skip connection is established between the integrated feature representation and the corresponding pooled feature representation, and an optimized feature representation with a multi-scale receptive field is obtained through an element-level addition operation, including: S461-S463.
[0072] S461. Perform channel grouping operation on the integrated feature representation, divide the feature channels into multiple subgroups, calculate the inter-channel correlation matrix for each subgroup, reorganize the features within the subgroup according to the correlation matrix, and obtain a reconstructed feature representation.
[0073] Exemplarily, a channel grouping operation is performed on the integrated feature representation, and the entire feature channel is evenly divided into 8 subgroups according to a preset number, each of which contains adjacent feature channels. An N×N dimensional inter-channel correlation matrix is calculated for each subgroup, where N is the number of channels in the subgroup. The correlation coefficient is obtained by calculating the dot product between any two channel feature maps and performing L2 normalization. The correlation coefficient range is controlled between [-1, 1]. Channel reorganization weights are generated based on the correlation matrix. Channel features with weights greater than 0.7 are reorganized and aggregated, channel features with weights less than 0.3 are suppressed, and channel features with weights between 0.3 and 0.7 remain unchanged, generating a reconstructed feature representation with higher discriminative ability.
[0074] S462. Perform feature alignment on the reconstructed feature representation and the corresponding pooled feature representation, perform spatial transformation on the reconstructed feature through learnable affine transformation parameters, and scale normalize the corresponding pooled feature representation to obtain an aligned feature representation pair.
[0075] For example, a feature alignment process precisely matches the reconstructed feature representation with the pooled feature representation in spatial dimensions. A learnable affine transformation is applied to the reconstructed features, with the transformation parameters predicted by a high-order attention network. The network input is the global statistics of the reconstructed features. A grid sampling strategy is used during the affine transformation to generate a 16×16 regular sampling grid. The sampling points are then resampled using bicubic interpolation after applying the transformation matrix to the sampling points, with an accuracy of 0.01 pixel. The pooled feature representation is scale-normalized. Using the average amplitude of the reconstructed features as a reference, the channel-wise scaling of the pooled feature representation is dynamically adjusted within the range of [0.8, 1.2]. The scaled pooled features and the transformed reconstructed features maintain consistent feature response patterns in spatial dimensions, resulting in an aligned feature representation pair. This bidirectional feature alignment strategy effectively eliminates differences in geometry and scale between the two feature representations, laying the foundation for subsequent feature fusion.
[0076] In some embodiments, spatial transformation of the reconstructed features is performed using learnable affine transformation parameters, including: S4621-S4623.
[0077] S4621. Calculate the spatial feature distribution statistics for the reconstructed feature representation, construct a feature covariance matrix by extracting the mean and variance of each channel of the feature map, perform eigenvalue decomposition on the covariance matrix, determine the main direction and scale parameters of the feature based on the main eigenvector, and obtain a spatial feature descriptor.
[0078] For example, the spatial feature distribution statistics of the reconstructed feature representation are calculated. First, the spatial mean and standard deviation are calculated for each feature channel to form a first-order statistical feature description vector. A 128×128-dimensional feature covariance matrix is further constructed. The matrix elements represent the correlation between the feature responses of any two spatial positions. The calculation of the covariance matrix is accelerated by fast Fourier transform, reducing the computational complexity from O(n²) to O(n log n). Eigenvalue decomposition is performed on the covariance matrix, and the first 8 main eigenvectors are extracted as the main directions of the features. The corresponding eigenvalues are used as scale parameters. The size of the eigenvalue reflects the degree of change of the feature in this direction. A 16-dimensional spatial feature descriptor is constructed based on the main direction and scale parameters. The numerical range of each element in the descriptor is controlled between [-1, 1]. The descriptor contains the main spatial distribution information of the feature and provides an accurate geometric reference for subsequent affine transformation.
[0079] S4622. Construct a six-parameter affine transformation matrix based on the spatial feature descriptor, decompose the six-parameter affine transformation matrix into a rotation component, a translation component, and a scaling component, set a learnable parameter for each component, update the learnable parameter through back propagation, and obtain an optimized affine transformation matrix.
[0080] For example, a six-parameter affine transformation matrix is constructed based on the spatial feature descriptor, with the six parameters corresponding to horizontal scaling, vertical scaling, horizontal shearing, vertical shearing, horizontal translation, and vertical translation, respectively. The affine transformation matrix is decomposed into a 2×2 rotation and scaling matrix and a 2×1 translation vector. The rotation and scaling matrix controls the shape transformation of the feature, with the rotation angle range limited to [-30°, 30°] and the scaling ratio limited to [0.75, 1.25]. The translation vector controls the position offset of the feature, with the offset distance not exceeding 10% of the feature map size. A learnable parameter with a dynamic learning rate is set for each transformation component, with the initial learning rate set to 0.01. During training, the learning rate is adaptively adjusted according to the gradient change of the loss function. The transformation parameters are optimized and updated using the backpropagation algorithm, with the step size of each update not exceeding 5% of the current parameter value to ensure the stability of the transformation. After approximately 200 iterative optimizations, the optimal affine transformation matrix is obtained, which accurately captures the geometric deformation requirements of the feature.
[0081] S4623. Apply an optimized affine transformation matrix to the reconstructed feature representation, spatially rearrange the features through a bilinear interpolation resampling method, adjust the spatial position relationship and geometric shape of the features, and obtain the feature representation after spatial transformation.
[0082] Exemplarily, the optimized affine transformation matrix is applied to the reconstructed feature representation, and spatial transformation is performed by establishing a mapping relationship from the target coordinates to the source coordinates. The bilinear interpolation resampling method is used in the transformation process. For each target position, the corresponding floating-point coordinates in the source feature map are calculated, and then interpolated from the values at the four nearest integer coordinates of the source feature. The interpolation weight is inversely proportional to the pixel distance. In order to improve computational efficiency, 16-bit floating-point numbers are used to represent the intermediate calculation results, which reduces the amount of calculation by about 40% while ensuring accuracy. The spatial transformation process adaptively adjusts the sampling density according to the local texture complexity. A higher density of sampling points is used in texture-rich areas, and the sampling density is reduced in texture-flat areas, effectively balancing the resampling quality and computational overhead. After spatial rearrangement, the geometric shape and spatial distribution of the reconstructed features are precisely adjusted to form a feature representation that is highly matched with the target pooling features in spatial structure.
[0083] S463. Adaptively fuse the aligned feature representations, calculate the fusion weights through a trainable gating unit, perform weighted combination of features according to the fusion weights, and aggregate features through a residual connection structure to obtain an optimized feature representation with a multi-scale receptive field.
[0084] For example, adaptive feature fusion is performed on the aligned feature representations. First, a gating unit is constructed using a three-layer fully connected neural network. The network input is the channel statistics of the two features. The number of hidden layer nodes is 512, and the activation function is LeakyReLU. The number of nodes in the output layer is equal to the number of feature channels. The activation function is the Sigmoid function, and the output value range is [0, 1], representing the fusion weight of each channel. The fusion weight determines the contribution ratio of the pooled and transformed features on different channels. Weight values close to 1 favor the pooled features, while weight values close to 0 favor the transformed features. The two features are element-wise weighted summed according to the fusion weight, with the weight coefficient accurate to 0.001. The fused features are then aggregated with the original integrated features using a residual connection structure with a weight factor set to 0.2. This preserves the discriminative information of the fused features while maintaining the basic expressive power of the original features. This process generates an optimized feature representation with a multi-scale receptive field, which combines fine-grained local texture information with large-scale global semantic information.
[0085] In some embodiments, an appearance detection model is obtained based on the optimized feature representation and the verification image set, including: S1051-S102.
[0086] S1051. Construct an appearance detection network based on the optimized feature map.
[0087] S1052. After training and verification using a verification image set, a trained appearance detection model is obtained.
[0088] Exemplarily, an appearance detection network is constructed based on the optimized feature maps obtained in the previous steps. After construction, the network is trained, with backpropagation continuously optimizing network parameters. Simultaneously, model performance is regularly evaluated using a validation image set containing a variety of typical electrical control appearance defect samples. Based on the test results from the validation set, training strategies and network parameters can be adjusted promptly to produce a stable appearance detection model. This combined training and validation approach effectively improves the model's generalization capabilities, ensuring high detection accuracy and low missed detection rates in practical applications, providing reliable technical support for electrical control appearance defect detection.
[0089] See also Figure 2 , Figure 2 This is a schematic block diagram of a deep learning-based electronic control component appearance inspection device provided in an embodiment of the present application. This deep learning-based electronic control component appearance inspection device 200 is used to perform the aforementioned deep learning-based electronic control component appearance inspection method. This deep learning-based electronic control component appearance inspection device 200 can be configured in a server.
[0090] Among them, the server can be an independent server, a server cluster, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0091] like Figure 2 As shown, the electric control component appearance inspection device 200 based on deep learning includes: an image acquisition module 201, a feature extraction module 202, a feature pooling module 203, a feature optimization module 204 and a model generation module 205. The electric control component appearance inspection device 200 based on deep learning is used to perform the electric control component appearance inspection method based on deep learning as described in any one of the embodiments of the present application.
[0092] The image acquisition module 201 is used to acquire a training image set and a verification image set of the electronic control component.
[0093] The feature extraction module 202 is used to construct a feature pyramid based on the training image set, perform dimension importance evaluation and regional significance analysis on the feature pyramid, and obtain enhanced feature representation.
[0094] The feature pooling module 203 is used to perform multiple pooling processes on the enhanced feature representation to obtain multiple groups of pooled feature representations.
[0095] The feature optimization module 204 is used to determine a high-order representation and a low-order representation according to each group of pooled feature representations, perform dynamic weight integration on the high-order representation and the low-order representation, and introduce skip connections to obtain an optimized feature representation.
[0096] The model generation module 205 is used to obtain an appearance inspection model based on the optimized feature representation and the verification image set, and perform appearance inspection on the electronic control component using the appearance inspection model.
[0097] An embodiment of the present application provides an electronic device, which includes a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and implement a deep learning-based appearance detection method for an electronic control component as described in any one of the embodiments of the present application when executing the computer program.
[0098] An embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the processor implements a method for detecting the appearance of an electronically controlled component based on deep learning, such as any one of the embodiments of the present application.
[0099] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A method for appearance inspection of electronic control components based on deep learning, characterized in that: include: Obtaining a training image set and a verification image set of the electronic control component; constructing a feature pyramid based on the training image set, performing dimension importance evaluation and regional saliency analysis on the feature pyramid to obtain enhanced feature representation; Construct three parallel soft pooling branches, set the first pooling core, the second pooling core, and the third pooling core respectively, wherein the size of the first pooling core is 5×5, the size of the second pooling core is 9×9, and the size of the third pooling core is 13×13; the enhanced feature representation is passed through the first pooling core, the second pooling core, and the third pooling core respectively to obtain the first pooling feature representation, the second pooling feature representation, and the third pooling feature representation; The first pooled feature representation, the second pooled feature representation, and the third pooled feature representation are respectively subjected to path separation and divided into a trunk path and a bypass path; the pooled feature representation of the trunk path is subjected to convolution transformation and nonlinear activation to obtain a trunk feature representation; the pooled feature representation of the bypass path is subjected to identity mapping to obtain a bypass feature representation; a high-order representation and a low-order representation are respectively extracted according to the trunk feature representation and the bypass feature representation, wherein the trunk feature representation is processed by a densely connected convolutional network to obtain the high-order representation, and the bypass feature representation is subjected to linear projection to obtain the low-order representation; inter-channel adaptive coefficients are calculated for the high-order representation and the low-order representation to generate a weighting coefficient matrix, and the high-order representation and the low-order representation are dynamically weighted integrated according to the weighting coefficient matrix to obtain an integrated feature representation; a skip connection is established between the integrated feature representation and the corresponding pooled feature representation, and an optimized feature representation with a multi-scale receptive field is obtained through an element-level addition operation; Based on the optimized feature representation and the verification image set, an appearance detection model is obtained, and appearance detection of the electronic control component is performed using the appearance detection model.
2. The method for appearance inspection of electronic control components based on deep learning according to claim 1, characterized in that: The step of constructing a feature pyramid based on the training image set, performing dimension importance evaluation and regional significance analysis on the feature pyramid to obtain enhanced feature representation includes: Constructing a feature pyramid having a first scale layer, a second scale layer, and a third scale layer according to the training image set; Adding a channel attention mechanism to each scale layer of the feature pyramid to obtain a channel feature representation of each scale layer; Calculating spatial dependencies on the channel feature representation to obtain a spatial feature representation; Adaptively fusing the channel feature representation and the spatial feature representation to generate an enhanced feature representation for each scale layer; Perform cross-layer connections on the enhanced feature representation to obtain a strengthened feature representation.
3. The method for appearance inspection of electronic control components based on deep learning according to claim 1, characterized in that: The step of establishing a skip connection between the integrated feature representation and the corresponding pooled feature representation and obtaining an optimized feature representation with a multi-scale receptive field through element-level addition operation includes: performing a channel grouping operation on the integrated feature representation to divide the feature channels into a plurality of subgroups, calculating an inter-channel correlation matrix for each subgroup, and recombining the features within the subgroup according to the correlation matrix to obtain a reconstructed feature representation; Performing feature alignment on the reconstructed feature representation and the corresponding pooled feature representation, performing spatial transformation on the reconstructed feature using learnable affine transformation parameters, and performing scale normalization on the corresponding pooled feature representation to obtain an aligned feature representation pair; Adaptive feature fusion is performed on the aligned feature representation pairs, fusion weights are calculated through a trainable gating unit, features are weightedly combined according to the fusion weights, and feature aggregation is performed through a residual connection structure to obtain an optimized feature representation with a multi-scale receptive field.
4. The method for detecting the appearance of an electronic control component based on deep learning according to claim 3, wherein: The spatial transformation of the reconstructed features by using learnable affine transformation parameters includes: Calculating spatial feature distribution statistics for the reconstructed feature representation, constructing a feature covariance matrix by extracting the mean and variance of each channel of the feature map, performing eigenvalue decomposition on the covariance matrix, determining the main direction and scale parameters of the feature based on the main eigenvector, and obtaining a spatial feature descriptor; Constructing a six-parameter affine transformation matrix based on the spatial feature descriptor, decomposing the six-parameter affine transformation matrix into a rotation component, a translation component, and a scaling component, setting a learnable parameter for each component, and updating the learnable parameter through back propagation to obtain an optimized affine transformation matrix; The optimized affine transformation matrix is applied to the reconstructed feature representation, and the features are spatially rearranged through a bilinear interpolation resampling method to adjust the spatial position relationship and geometric shape of the features to obtain a feature representation after spatial transformation.
5. The method for detecting appearance of electronic control components based on deep learning according to claim 1, wherein: The obtaining of an appearance detection model based on the optimized feature representation and the verification image set includes: constructing an appearance detection network based on the optimized feature representation; After training and verification using the verification image set, a trained appearance detection model is obtained.
6. A device for detecting the appearance of electronic control components based on deep learning, characterized in that: The electric control component appearance inspection device based on deep learning is used to perform the electric control component appearance inspection method based on deep learning according to any one of claims 1 to 5, and the electric control component appearance inspection device based on deep learning includes: An image acquisition module, used to acquire a training image set and a verification image set of the electronic control component; A feature extraction module is used to construct a feature pyramid based on the training image set, perform dimension importance evaluation and regional significance analysis on the feature pyramid, and obtain enhanced feature representation; A feature pooling module is used to construct three parallel soft pooling branches, respectively setting a first pooling core, a second pooling core, and a third pooling core, wherein the size of the first pooling core is 5×5, the size of the second pooling core is 9×9, and the size of the third pooling core is 13×13; the enhanced feature representation is passed through the first pooling core, the second pooling core, and the third pooling core respectively to obtain a first pooled feature representation, a second pooled feature representation, and a third pooled feature representation; A feature optimization module is used to perform path separation on the first pooled feature representation, the second pooled feature representation, and the third pooled feature representation, respectively, and divide them into a trunk path and a bypass path; perform convolution transformation and nonlinear activation on the pooled feature representation of the trunk path to obtain a trunk feature representation; perform identity mapping on the pooled feature representation of the bypass path to obtain a bypass feature representation; extract a high-order representation and a low-order representation based on the trunk feature representation and the bypass feature representation, respectively, wherein the trunk feature representation is processed by a densely connected convolutional network to obtain the high-order representation, and the bypass feature representation is processed by linear projection to obtain the low-order representation; perform inter-channel adaptive coefficient calculation on the high-order representation and the low-order representation to generate a weighting coefficient matrix, and perform dynamic weight integration on the high-order representation and the low-order representation according to the weighting coefficient matrix to obtain an integrated feature representation; establish a skip connection between the integrated feature representation and the corresponding pooled feature representation, and obtain an optimized feature representation with a multi-scale receptive field through element-level addition operation; A model generation module is used to obtain an appearance detection model based on the optimized feature representation and the verification image set, and perform appearance detection on the electronic control component through the appearance detection model.
Citation Information
Patent Citations
Small sample steel defect detection method based on attention feature pyramid mechanism
CN116958073A
Steel surface defect detection algorithm based on improved YOLOv5
CN118674697A