Vision-based automobile plastic part surface flaw identification method

Through the deep learning model based on the generative adversarial network, the existing automotive plastic parts defect detection methods have been solved, and high-precision identification and adaptability to various defect types have been achieved, and production efficiency and product quality have been improved.

CN120125969AActive Publication Date: 2025-06-10JIANGSU LV NENG AUTO PARTS SCI & TECH CO LTD

Patent Information

Application Number
CN202510218936.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-10
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

The existing automotive plastic parts defect detection methods have low accuracy and limited generalization capabilities, making it difficult to deal with complex and variable lighting conditions and defect forms.

Method used

A deep learning model based on a generative adversarial network is adopted to acquire and preprocess the surface images of historical automobile plastic parts, create a data set, and train it through the generator network, discriminator network and classification module to achieve high-precision recognition of various defect types.

Benefits of technology

It realizes high-precision identification of various defect types on the surface of automotive plastic parts, adapts to changes in different lighting conditions and shooting angles, improves production efficiency and product quality, and has high scalability and flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125969A_ABST
    Figure CN120125969A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of industrial intelligent detection, and discloses an automobile plastic part surface flaw identification method based on vision. According to the method, a historical automobile plastic part surface image is obtained, a data set is established, and after preprocessing, labeling and division, a deep learning model based on a generative adversarial network is established for training, so that an automobile plastic part surface defect recognition model is obtained. The model can collect the surface image of the automobile plastic part in real time, can identify various flaw types such as scratches, bubbles, material shortage and weld lines, and is suitable for the change of different illumination conditions and shooting angles. Meanwhile, a dynamic threshold judgment mechanism and a model explanation module are adopted, flaw problems are found and alarmed in time, and the production efficiency and the product quality are improved. The method has high expandability and flexibility, and provides powerful support for intelligent transformation of the automobile plastic part manufacturing industry.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of industrial intelligent detection, and particularly relates to a method for identifying surface defects of automotive plastic parts based on vision. Background Art

[0002] During the manufacturing process of automotive plastic parts, surface defects seriously affect product quality and customer satisfaction. Traditional manual inspection methods are not only time-consuming and laborious, but also easily affected by subjective factors, resulting in difficulties in ensuring inspection efficiency and accuracy. With the rapid development of machine vision technology, using deep learning models for defect detection has become a research hotspot. However, existing defect detection methods often target specific defect types, have limited generalization ability, and are difficult to handle complex and variable lighting conditions and defect morphologies. Therefore, there is an urgent need for an intelligent detection technology that can efficiently identify various surface defect types of automotive plastic parts. Summary of the Invention

[0003] Aiming at the above-mentioned technical deficiencies, the purpose of the present invention is to provide a method and system for detecting the surface quality of brake discs, which solves the problems of low accuracy and limited generalization ability in the existing defect detection methods for automotive plastic parts.

[0004] To solve the above technical problems, the present invention adopts the following technical solutions: In the first aspect, the present invention provides a method for identifying surface defects of automotive plastic parts based on vision, and the method includes: Step S100: Obtain historical surface images of automotive plastic parts, preprocess the images, and establish a surface dataset of automotive plastic parts; Step S200: Label the surface dataset of automotive plastic parts and divide it into a training set and a validation set; Step S300: Build a deep learning model based on a generative adversarial network; Step S400: Input the training set into the deep learning model based on the generative adversarial network for training, and use the validation set for verification during the training process to obtain a surface defect recognition model for automotive plastic parts; Step S500: Real-time collect surface images of automotive plastic parts, and input them into the surface defect recognition model for automotive plastic parts to obtain the surface defect recognition result of automotive plastic parts.

[0005] Preferably, in a possible implementation manner of the first aspect, the historical surface images of automotive plastic parts are collected by a high-resolution linear array CCD camera, including surface images of plastic parts taken under different lighting conditions and at multiple angles, and covering sample data within at least a 12-month production cycle.

[0006] Preferably, in a possible implementation manner of the first aspect, the preprocessing includes: The bilateral filtering algorithm is used for image denoising processing; The image contrast is enhanced by histogram equalization; The image is scaled to a resolution of 512×512 pixels.

[0007] Preferably, in a possible implementation manner of the first aspect, LabelImg is used to label the preprocessed surface dataset of automotive plastic parts, and the labeling types include four types of defects: scratches, bubbles, material shortage, and weld lines.

[0008] Preferably, in a possible implementation manner of the first aspect, the deep learning model based on the generative adversarial network includes: The generator network, adopting the U-Net architecture, contains 8 layers of downsampling-upsampling structures, which are used to gradually extract and restore the surface image features of automotive plastic parts. Each layer integrates a residual block and a self-attention mechanism module. Among them, the 3rd to 5th layers embed Transformer encoder units to capture long-range dependencies and learn the feature associations of cross-region defects. The output end uses a spectral normalization convolutional layer to generate a 512×512 high-resolution feature map; The discriminator network constructs a multi-scale pyramid architecture, containing 3 parallel convolutional streams, which respectively process the original resolution, 1 / 2 downsampled, and 1 / 4 downsampled images to capture different-scale defects on the surface of automotive plastic parts. Each branch adopts depthwise separable convolution and adaptive instance normalization techniques to adapt to different image feature distributions, and the end integrates multi-scale image features through a gated fusion module; The classification module contains a parallel attention mechanism layer and a fully connected layer. Among them, the attention mechanism layer adopts a dynamic multi-head self-attention mechanism, sets 12 attention heads and integrates a channel attention module to analyze the defect features on the surface of automotive plastic parts from different angles. The fully connected layer adopts a mixture of experts system architecture, containing 32 expert networks and a learnable gating controller to judge the defect types.

[0009] Preferably, in a possible implementation manner of the first aspect, the generator network specifically includes: The input layer uses a 4-channel tensor with positional encoding to mark the defect position features; The downsampling stage contains 4 residual blocks, and each residual block is composed of 2 3×3 spectral normalization convolutional layers, with a window self-attention module based on Swin Transformer inserted in the middle; The bottleneck layer integrates 3 cascaded Transformer encoding units, and each unit contains 12 attention heads and a 2048-dimensional feed-forward network; The upsampling stage adopts sub-pixel convolution technology and cooperates with adaptive instance normalization for multi-level image feature fusion; The output layer uses the tanh activation function and combines a channel attention gating mechanism to generate a synthetic image with enhanced local details.

[0010] Preferably, in a possible implementation manner of the first aspect, the discriminator network specifically includes: The multi-scale input branch adopts an adaptive resampling technique to dynamically adjust the resolution of the input pictures of the surface of automotive plastic parts for each branch; Each convolutional stream contains 5 stages. In each stage, a depthwise separable convolution and adaptive instance normalization combination is adopted, and the LeakyReLU activation function is used to extract multi-scale features of defects of different sizes and shapes on the surface of automotive plastic parts; The feature fusion module adopts a topological structure generated by differentiable neural architecture search, integrates a cross-scale feature interaction mechanism, and fuses multi-scale features of defects of different sizes and shapes on the surface of automotive plastic parts; The discriminative head contains a region focusing mechanism that applies a 3-fold weight to the key region through a saliency map generated in real time; The training process uses the Wasserstein loss function, integrates gradient penalty and spectral regularization constraints, and improves the training stability and convergence speed of the discriminator network.

[0011] Preferably, in a possible implementation manner of the first aspect, the classification module specifically includes: The attention mechanism layer adopts a spatio-temporal joint attention architecture, which contains 12 dynamically configured attention heads, among which 4 heads are dedicated to spatial dimension interaction, 6 heads handle the relationship between channels, and 2 heads analyze cross-layer feature associations to mine the associations between different channel features and defect types; The channel attention module integrates a dynamic channel compression mechanism, and the compression rate is automatically adjusted according to the input features to highlight the features related to defects; The mixture of experts system of the fully connected layer contains 32 heterogeneous expert networks, among which 16 experts adopt a 3-layer MLP structure to process the local features of defects, 8 experts adopt a graph convolutional network to classify the defect features, and 8 experts adopt a time series analysis module to analyze the change trend of defects; The output layer adopts a dynamic weight fusion technique to integrate the outputs of the attention branch and the fully connected branch through a learnable gating parameter α; During training, label smoothing regularization and contrastive learning auxiliary loss are introduced, and the classification decision threshold is dynamically adjusted according to the performance of the validation set.

[0012] Preferably, in a possible implementation manner of the first aspect, the step S400 specifically includes: Set training parameters, including an Adam optimizer with an initial learning rate of 0.002, a batch size of 32, and a maximum of 200 training epochs, and configure an early stopping mechanism to terminate training when the validation set loss has not decreased for 10 consecutive epochs; Adopt an alternating training strategy. In the first round, freeze the generator network parameters and train the discriminator network using the Wasserstein loss function. In the second round, freeze the discriminator network parameters and optimize the generator network through a combination of adversarial loss, L1 reconstruction loss, and perceptual loss; After every 5 rounds of training, start the validation set evaluation, and synchronously calculate three metrics: the FID metric between the generated image and the real image, the defect classification accuracy, and the intersection over union (IoU). When the FID is lower than 15 and the IoU is higher than 0.85, trigger the saving of the model snapshot; After training is completed, compress the generative adversarial network into a lightweight recognition model through knowledge distillation technology, retaining more than 97% of the accuracy of the original model.

[0013] Preferably, in a possible implementation manner of the first aspect, the step S500 specifically includes: Use a high-resolution line array CCD camera deployed on the production line to collect 2048×2048 original resolution images at a rate of 1 frame per second, and perform preprocessing operations in real time; Input the preprocessed images into the automotive plastic part surface defect recognition model. The generator network outputs a 512×512 feature map with a defect heat map, and the discriminator network generates multi-scale confidence scores. The classification module generates structured data containing defect types, position coordinates, and confidence levels after synthesizing the two types of outputs; Adopt a dynamic threshold determination mechanism. When the comprehensive confidence level exceeds the dynamic decision threshold, trigger an audible and visual alarm and record the defect image and defect coordinate information; Visualize the decision basis through the model interpretation module, use the class activation mapping technology to highlight the defect area, and generate an interactive detection report containing the timestamp, production line station, and defect statistical analysis.

[0014] The beneficial effects of the present invention are as follows: Establish a dataset for the surface of automotive plastic parts, and use a generative adversarial network deep learning model for training, achieving high-precision recognition of defects on the surface of automotive plastic parts.

[0015] This method can not only effectively identify various defect types such as scratches, bubbles, material shortages, and weld lines, but also adapt to changes in different lighting conditions and shooting angles. At the same time, by collecting and processing image data in real time, this method can promptly detect and alarm defect problems, improving production efficiency and product quality. In addition, this method also has high scalability and flexibility, and can be optimized and upgraded according to actual needs, providing strong support for the intelligent transformation of the automotive plastic part manufacturing industry. Description of the Drawings

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0017] Figure 1 This application provides a flowchart of a method for identifying surface defects of automotive plastic parts based on vision. Detailed implementation manners

[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0019] Embodiment 1: As Figure 1 shown, the present invention provides a method for identifying surface defects of automotive plastic parts based on vision, including: Step S100: Obtain historical surface images of automotive plastic parts, preprocess the images, and establish a surface dataset of automotive plastic parts.

[0020] In this embodiment, a Basler ace acA5472-5gm linear array CCD camera is selected to collect historical surface images of automotive plastic parts. This camera has a resolution of 20 million pixels and a frame rate of 5 fps. The linear array CCD camera is deployed 50 centimeters away from the plastic parts on the automotive plastic part production line, and the angle of the camera forms a 90-degree angle with the transmission direction of the plastic parts. The collected images are transmitted and stored in the control center in real time using the TCP / IP transmission protocol. The historical surface images of automotive plastic parts include images taken under different lighting conditions and at multiple angles, and cover a production cycle of at least 12 months, completely including automotive plastic parts of different batches and different production stages.

[0021] After completing the image acquisition and data set establishment, the data set is preprocessed. First, the preprocess_image function is defined to implement the specific preprocessing process. In this function, the cv2.bilateralFilter function is first used to perform bilateral filtering and noise reduction on the input image. By setting parameters such as the filter kernel diameter, the color space filter standard deviation, and the coordinate space filter standard deviation, the image edge information is retained while removing noise. Then, it is determined whether the image is grayscale or color. If it is a grayscale image, the cv2.equalizeHist function is directly used for histogram equalization; if it is a color image, it is first converted from the BGR color space to the YCrCb color space, and the Y channel is histogram equalized to enhance the contrast. Then the processed Y channel is merged with the unprocessed Cr and Cb channels and converted back to the BGR color space. Finally, the cv2.resize function is used to scale the image to a resolution of 512×512 pixels and return the processed image.

[0022] Step S200: annotating the automobile plastic parts surface dataset and dividing it into a training set and a validation set.

[0023] In this embodiment, after obtaining the pre-processed automobile plastic parts surface dataset from the control center, the data is first sorted and checked to ensure that the file naming is standardized and consistent for easy management and search. At the same time, the integrity and quality of each image are carefully checked, and damaged, blurred or abnormally pre-processed images are removed from the dataset to ensure that the data used for annotation and training is of high quality and available.

[0024] Then formulate the marking specifications. For scratches, draw polygons along the edges to accurately outline the shape and range, and clearly mark them as "scratches" in the marking properties; for bubbles, select them completely with a rectangular frame, set the marking property to "bubble", and record relevant information such as size; for missing material, mark the area with a polygon, set the marking property to "missing material", and mark features such as area; for weld lines, use broken lines or polygons to mark the direction and range, and mark the marking properties as "weld line".

[0025] Then select LabelImg as the annotation tool, and annotate each image strictly according to the established annotation specifications during the annotation process. Finally, divide the data set. After completing the review and correction of the annotated data, randomly divide the annotated data set into a training set and a validation set in a ratio of 8:2, using the train_test_split function in the Python sklearn library. After the division is completed, store the training set and the validation set in different folders, and record the image file name and corresponding annotation information contained in each data set.

[0026] Step S300: Build a deep learning model based on a generative adversarial network.

[0027] In this embodiment, the deep learning model based on the generative adversarial network includes a generator network, a discriminator network, and a classification module.

[0028] The input layer of the generator network uses a 4-channel tensor with positional encoding to mark the defect location features. This helps the model clearly know the possible defect locations in the image during subsequent processing, so as to more specifically extract and process the features of these locations, providing a location information basis for accurately identifying four types of defects: scratches, bubbles, material shortage, and weld lines.

[0029] The overall architecture adopts the U-Net architecture, which contains 8 layers of downsampling-upsampling structures for gradually extracting and restoring the surface image features of automotive plastic parts. The downsampling process can continuously compress the image size while extracting the high-level semantic features in the image, and the upsampling restores these high-level features to the original image size, thus achieving a comprehensive capture of the surface image features of automotive plastic parts and providing rich feature information for identifying different types of defects.

[0030] The downsampling stage contains 4 residual blocks, each of which consists of 2 3×3 spectral normalization convolutional layers, with a window self-attention module based on Swin Transformer inserted in the middle. Spectral normalization enhances the stability of model training, ensuring the smoothness of parameter updates during training and avoiding problems such as gradient vanishing or explosion. The window self-attention module of Swin Transformer enables the model to pay more careful attention to and process the image features within the local window, which helps to capture the fine features of defects such as scratches and bubbles in the local area.

[0031] The bottleneck layer integrates 3 cascaded Transformer encoding units, each of which contains 12 attention heads and a 2048-dimensional feed-forward network. Through this structure, the features extracted in the downsampling stage are further deeply processed and fused, enabling the capture of more complex and abstract feature relationships, which helps to distinguish different types of defects, such as defects with relatively complex feature patterns like material shortage and weld lines.

[0032] The upsampling stage uses sub-pixel convolution technology, combined with adaptive instance normalization for multi-level image feature fusion. Sub-pixel convolution technology can effectively improve the image resolution and gradually restore the spatial detail information lost during the downsampling process. Adaptive instance normalization performs adaptive normalization processing according to different feature maps, enabling better fusion of features at different levels, thereby generating a synthetic image containing richer details, which is beneficial for accurately identifying the detailed features of various defects.

[0033] The output layer uses the tanh activation function and combines the channel attention gating mechanism to generate a synthetic image with enhanced local details. The tanh activation function maps the output values to the range [-1, 1], which helps the model learn more reasonable feature representations. The channel attention gating mechanism further enhances the model's ability to capture defect features by highlighting the features related to the surface defects of automotive plastic parts, enabling the generated synthetic image to more clearly present the details of defects such as scratches and bubbles, providing more favorable image data for subsequent defect recognition.

[0034] Specifically, a generator model Generator based on the generative adversarial network is first defined. This model inherits from torch.nn.Module and is used to generate synthetic images with enhanced local details. Its input is a 4-channel tensor with positional encoding, and the output is a 3-channel image.

[0035] In the model initialization part, the __init__ method receives two parameters, input_channel and output_channels, which are 4 and 3 respectively, representing the number of input and output channels. Subsequently, the downsampling stage, bottleneck layer, upsampling stage, output layer, and channel attention gating mechanism are defined.

[0036] The downsampling stage consists of 4 modules, stored in down_blocks of type nn.ModuleList. Each module contains a residual block and a Swin Transformer. The residual block consists of 2 3×3 spectral normalization convolutional layers, with a window self-attention module based on Swin Transformer inserted in the middle. Spectral normalization is used to enhance the stability of model training. The input channel number of the Swin Transformer is the output channel number of the residual block, and the channel number doubles after each module is processed. The bottleneck layer is also stored using nn.ModuleList and contains 3 cascaded Transformer encoding units, whose input channel number is the output channel number of the last module in the downsampling stage. The upsampling stage contains 4 modules, stored in up_blocks. Each module first performs sub-pixel convolution through nn.PixelShuffle to increase the image resolution, and then sequentially undergoes instance normalization, 3x3 convolution, and instance normalization operations. The channel number is halved after each module is processed. The output layer consists of a 3x3 convolutional layer and a tanh activation function, which converts the output of the upsampling stage into the final 3-channel image. The channel attention gating mechanism consists of adaptive average pooling, 1x1 convolution, ReLU activation function, 1x1 convolution, and Sigmoid activation function in sequence, and is used to highlight the features related to the surface defects of automotive plastic parts.

[0037] The forward method defines the forward propagation process of the model. First, during downsampling, the input data passes through the modules in down_blocks in sequence. At the same time, the output of each module is stored in the skip_connections list, and average pooling is used to downsample the feature map. Then, the input data passes through 3 Transformer encoding units in the bottleneck layer. In the upsampling stage, the skip_connections list is reversed and passes through the modules in up_blocks in sequence. At the same time, the custom _adaptive_instance_norm method is used for adaptive instance normalization and feature fusion, and the currently upsampled feature map is concatenated with the feature map stored during the previous downsampling. After that, an attention map is generated through the channel attention gating mechanism and multiplied with the feature map to highlight important features. Finally, the final synthesized image is obtained through the output layer. The _adaptive_instance_norm method implements adaptive instance normalization. By calculating the mean and standard deviation of the content feature map and the style feature map, the content feature map is normalized and rescaled to have similar statistical characteristics to the style feature map.

[0038] The discriminator network first constructs a multi-scale pyramid architecture. The multi-scale input branches adopt adaptive resampling technology to dynamically adjust the resolution of the input images of the automotive plastic part surfaces for each branch. This enables the model to analyze the images of the automotive plastic part surfaces from different resolution perspectives and capture defects of different scales, whether they are tiny scratches, bubbles, or larger-scale material shortages, weld lines and other defects can be effectively noticed.

[0039] The parallel convolutional streams consist of 3 parallel convolutional streams, which process the original resolution, 1 / 2 downsampled, and 1 / 4 downsampled images respectively to capture different-scale defects on the automotive plastic part surfaces. Each convolutional stream contains 5 stages, and each stage adopts a combination of depthwise separable convolution and adaptive instance normalization, and the LeakyReLU activation function is used. Depthwise separable convolution can effectively extract multi-scale features of defects of different sizes and shapes while reducing the computational amount; adaptive instance normalization normalizes according to different feature maps to enhance the distinguishability of features; the LeakyReLU activation function introduces non-linearity while avoiding the information loss problem caused by the ReLU function completely setting to zero in the negative half-axis, which helps to better extract and retain the feature information related to defects.

[0040] At the end of the gated fusion module, multi-scale image features are integrated through the gated fusion module. The feature fusion module adopts the topological structure generated by differentiable neural architecture search, integrates the cross-scale feature interaction mechanism, and fuses the multi-scale features of different sizes and shapes of defects on the surface of automotive plastic parts. This fusion method can effectively integrate different-scale features extracted by different convolutional streams, enabling the model to comprehensively consider defect information at different scales and more accurately judge whether there are various defects such as scratches, bubbles, material shortages, and weld lines in the image.

[0041] The discriminative head contains a region focusing mechanism that applies a 3-fold weight to the key region through the saliency map generated in real time. It adopts the Wasserstein loss function, integrates gradient penalty and spectral regularization constraints, and improves the training stability and convergence speed of the discriminator network. The region focusing mechanism enables the model to pay more attention to the key regions where defects may exist in the image, enhancing the discriminative ability for defects; the Wasserstein loss function combined with gradient penalty and spectral regularization constraints helps optimize the training process of the discriminator, enabling it to converge to a better discriminative effect more stably and quickly, thus accurately identifying various defects on the surface of automotive plastic parts.

[0042] Specifically, first define the DifferentiableNASModule class, which inherits from nn.Module. The role of this class is the topological structure generated by differentiable neural architecture search. During initialization, it receives the number of input channels in_channels as a parameter, creates two fully connected layers fc1 and fc2, and a ReLU activation function layer relu. fc1 maps the number of input channels to in_channels / / 2, and fc2 then maps it back to in_channels. In the forward method, the input tensor x is first reshaped into a two-dimensional tensor, processed sequentially by fc1, relu, and fc2, and then the output is reshaped back into a four-dimensional tensor for subsequent operations with other feature maps.

[0043] Next, the core Discriminator class is defined, also inheriting from nn.Module. In the initialization part, multi-scale input branches are created, including three parallel convolutional streams: scale_original, scale_half, and scale_quarter, which are used to process images at the original resolution, 1 / 2 downsampling, and 1 / 4 downsampling, respectively. These three convolutional streams are created by calling the _create_conv_stream method. Each convolutional stream contains 5 stages, and each stage consists of a depthwise separable convolution, a 1x1 convolution, an instance normalization layer, and a LeakyReLU activation function. In this way, image features at different scales can be effectively extracted. At the same time, a gated fusion module gate_fusion is created, using the previously defined DifferentiableNASModule, with the input channel number being the sum of the final output channel numbers of the three convolutional streams. Finally, a discriminator head discriminator_head is created, which is a 3x3 convolutional layer used to map the feature map to an output of one channel to obtain the final discrimination result. The _create_conv_stream method is used to create each convolutional stream, receiving the input channel number in_channels as a parameter, and constructing the layer structure of each stage through 5 loops. In each loop, the channel number is first doubled using a depthwise separable convolution, then the channel is adjusted using a 1x1 convolution, followed by instance normalization and LeakyReLU activation, while updating the input channel number. Finally, all the layers are combined into an nn.Sequential object and returned.

[0044] In the forward method, the forward propagation process of the discriminator is implemented. First, the adaptive resampling technique is used for the input image x to obtain x_original at the original resolution, x_half downsampled by 1 / 2, and x_quarter downsampled by 1 / 4 respectively. Then they are respectively input into the corresponding convolutional streams to obtain feature maps output_original, output_half, and output_quarter at different scales. For subsequent feature fusion, the downsampled feature maps are restored to the same size as the feature map at the original resolution through bilinear interpolation. Next, these three feature maps are concatenated in the channel dimension to obtain concatenated_features. After that, the concatenated feature map is input into the gated fusion module gate_fusion to obtain gated_features, which is expanded to the same shape as the concatenated feature map, and then the two are multiplied element-wise to obtain the fused feature map fused_features. For cross-scale feature interaction, the fused feature map is reshaped and summed in the specified dimension to obtain final_features. Finally, the final feature map is input into the discriminator head discriminator_head to obtain the discrimination result and return it.

[0045] The attention mechanism layer in the classification module adopts a spatio-temporal joint attention architecture, which includes 12 dynamically configured attention heads. Among them, 4 heads are dedicated to spatial dimension interaction, 6 heads handle the relationships between channels, and 2 heads analyze cross-layer feature associations to mine the associations between different channel features and defect types. Through this configuration, the image features can be analyzed and focused from different dimensions. The spatial dimension interaction heads help capture the features of defects in the spatial position, the heads for relationships between channels can analyze the connections between different channel features, and the cross-layer feature association heads can integrate feature information at different levels, so as to more comprehensively mine the feature patterns related to the four types of defects, namely scratches, bubbles, material shortage, and weld lines, providing a basis for accurate classification.

[0046] The channel attention module integrates a dynamic channel compression mechanism, and the compression rate is automatically adjusted according to the input features to highlight the features related to defects. This mechanism automatically adjusts the channel compression ratio according to the characteristics of the input features, removes some channel information that contributes less to defect classification, and at the same time highlights the key channel features related to defects, further improving the pertinence and effectiveness of the features, enabling the model to make more accurate judgments based on the features closely related to defects during classification.

[0047] Hybrid expert system with fully connected layers: It contains 32 heterogeneous expert networks, of which 16 experts use a 3-layer MLP structure to process local features of defects. They can deeply analyze and process the features of defects in local areas, capture subtle local feature differences, and help distinguish the differences in local manifestations of different types of defects; 8 experts use graph convolutional networks to classify defect features, and use the advantages of graph convolutional networks in processing graph-structured data to explore the potential relationship between defect features and improve classification accuracy; 8 experts use time series analysis modules to analyze defect change trends. For some defects that may change over time or during the production process, such as material shortages, there may be certain change patterns in the continuous production process. By analyzing this change trend, the defect type can be judged more accurately.

[0048] The output layer uses dynamic weight fusion technology to integrate the outputs of the attention branch and the fully connected branch through a learnable gating parameter α. Label smoothing regularization and contrastive learning auxiliary loss are introduced during training, and the classification decision threshold is dynamically adjusted according to the performance of the validation set. The dynamic weight fusion technology can automatically adjust the weight contribution of the attention branch and the fully connected branch in the final classification result according to different input features, giving full play to the advantages of the two branches; label smoothing regularization can avoid overfitting of the model during training and improve the generalization ability of the model; contrastive learning auxiliary loss further enhances the model's ability to distinguish different defect types by comparing the feature differences between different samples; the classification decision threshold is dynamically adjusted according to the performance of the validation set, which enables the model to find an optimal classification decision boundary in different data sets and task scenarios, thereby improving the classification accuracy.

[0049] Specifically, the TemporalSpatialAttention class implements the spatiotemporal joint attention architecture. At initialization, three different types of attention heads are defined, namely 4 heads for spatial dimension interactions, 6 heads for processing inter-channel relationships, and 2 heads for analyzing cross-layer feature associations. Then, 1x1 convolutional layers for query, key, and value are created to convert the input feature map into the corresponding tensor. Finally, there is a 1x1 convolutional layer for output projection to convert the result of the attention calculation back to the same number of channels as the input. In the forward method, the batch size, number of channels, height, and width of the input feature map are first obtained. Then the input feature map is passed through the query, key, and value convolutional layers, reshaped, and transposed. Then the attention score is calculated, divided by the square root of the number of channels for scaling, and then passed through the softmax function to obtain the attention probability. Finally, the attention probability is multiplied with the value tensor, and after transposition and reshape adjustment, the final attention output is obtained through the output projection convolutional layer.

[0050] The DynamicChannelCompression class implements a dynamic channel compression mechanism. In the forward method, a scalar value is obtained by calculating the mean of the input feature map across all dimensions, and the torch.clamp function is used to limit this mean between 0.1 and 0.9 as the compression rate. The new number of channels is calculated based on the compression rate. If the new number of channels is greater than 0, the number of channels of the input feature map is cropped to the new number of channels, and finally the cropped feature map is returned.

[0051] The MixedExpertSystem class implements a mixed expert system. During initialization, three different types of expert networks are created. Among them, 16 experts with a 3-layer MLP structure are used to process local defect features; 8 expert graph convolutional networks simulated by fully connected layers are used to classify defect features; 8 LSTM-based time series analysis experts are used to analyze the trend of defect changes. In the forward method, the input feature map is first flattened and then input into different types of expert networks respectively to obtain the outputs of each expert. For the LSTM expert, the input is reshaped, input into the LSTM, and the output of the last time step is taken. Finally, the outputs of all experts are stacked together as the final output of the mixed expert system.

[0052] The ClassificationModule class, during initialization, includes the previously defined attention mechanism layer (TemporalSpatialAttention), the channel compression module (DynamicChannelCompression), and the mixed expert system (MixedExpertSystem). At the same time, a learnable gating parameter alpha is defined for dynamic weight fusion. Two fully connected layers are created, one for processing the output of the attention branch and the other for final classification. In the forward method, the input feature map first passes through the attention mechanism layer to obtain the attention output, then undergoes channel cropping through the channel compression module, and is converted into a one-dimensional vector through global average pooling. Finally, the output of the attention branch is obtained through the fully connected layer. At the same time, the input feature map is also input into the mixed expert system to obtain the output of the fully connected branch. Then, the learnable gating parameter alpha is used to perform dynamic weight fusion on the outputs of the two branches. Finally, the fused output passes through the final fully connected layer to obtain the classification result and returns it.

[0053] Step S400: Input the training set into the deep learning model based on the generative adversarial network for training. The training process is verified using the validation set to obtain a surface defect recognition model for automotive plastic parts.

[0054] In this embodiment, the training set is input into a deep learning model based on a generative adversarial network for training. During the training process, a validation set is used for verification to obtain a surface defect recognition model for automotive plastic parts. For the training parameters, the Adam optimizer is used, with its initial learning rate set to 0.002, the batch size set to 32, and the maximum number of training epochs set to 200. At the same time, to avoid overfitting of the model, an early stopping mechanism is configured. When the loss of the validation set has not decreased for 10 consecutive epochs, the training process will be terminated.

[0055] The training process adopts an alternating training strategy. In the first round of training, the parameters of the generator network are frozen so that they are not updated during this round of training. At this time, the Wasserstein loss function is used to train the discriminator network. The Wasserstein loss function can effectively measure the difference between the generated data and the real data and can avoid the problem of vanishing gradients.

[0056] In the second round of training, the parameters of the discriminator network are frozen and the generator network is optimized. During the optimization process, it is carried out by combining the adversarial loss, the L1 reconstruction loss, and the perceptual loss. The adversarial loss is used to measure the difference between the images generated by the generator and the real images in the eyes of the discriminator, prompting the generator to generate more realistic images; the L1 reconstruction loss focuses on the pixel-level difference between the generated images and the real images, which helps to improve the details and quality of the images; the perceptual loss measures the similarity between the generated images and the real images from the feature level, making the generated images closer to the real images in terms of semantic features.

[0057] During the training process, after every 5 rounds of training, the validation set evaluation is started. During the evaluation process, three important metrics are calculated synchronously, namely the FID (Frechet Inception Distance) metric between the generated images and the real images, the defect classification accuracy rate, and the intersection over union (IoU). The FID metric is used to measure the difference in feature distribution between the generated images and the real images. The lower the value, the closer the generated images are to the real images; the defect classification accuracy rate reflects the accuracy of the model in classifying the types of surface defects of automotive plastic parts; the intersection over union measures the overlap degree between the defect regions predicted by the model and the real defect regions. When the FID is lower than 15 and the IoU is higher than 0.85, the model snapshot is triggered to be saved, and the parameters of the current model are saved for subsequent use or further optimization.

[0058] After the training is completed, in order to improve the efficiency of the model in practical applications and the convenience of deployment, the generative adversarial network is compressed into a lightweight recognition model through knowledge distillation technology. During this process, more than 97% of the accuracy of the original model is retained as much as possible, so that the lightweight recognition model has a smaller model size and a faster inference speed while maintaining a high recognition accuracy rate.

[0059] Step S500: Real-time collect the surface image of the automotive plastic part, and input it into the surface defect recognition model of the automotive plastic part to obtain the surface defect recognition result of the automotive plastic part.

[0060] In this embodiment, the surface image of the automotive plastic part is collected in real time, and input into the surface defect recognition model of the automotive plastic part to obtain the surface defect recognition result of the automotive plastic part. First, image collection is performed by a high-resolution line array CCD camera deployed on the production line. The camera collects images with an original resolution of 2048×2048 at a rate of 1 frame per second. Such a collection rate and resolution can ensure that clear and complete surface image information of the automotive plastic part is obtained. The collected images are preprocessed in real time. The preprocessing process is the same as the preprocessing flow when establishing the dataset before, that is, the cv2.bilateralFilter function is used for bilateral filtering and noise reduction, and then histogram equalization processing is performed according to whether the image is grayscale or color. Finally, the cv2.resize function is used to scale the image to a resolution of 512×512 pixels.

[0061] The preprocessed image is input into the surface defect recognition model of the automotive plastic part. The generator network in the model processes the input image and outputs a 512×512 feature map with a defect heat map. The defect heat map can intuitively show the areas where defects may exist in the image. The darker the color, the greater the possibility of defects in that area. The discriminator network performs multi-scale analysis on the input image and generates multi-scale confidence scores, which reflect the possibility of defects in the image at different scales. The classification module combines the two types of outputs of the generator network and the discriminator network to generate structured data containing defect types, position coordinates, and confidence levels. Specifically, the classification module analyzes the generated feature map and confidence scores according to the feature patterns learned during previous training, determines the specific types of defects in the image (such as scratches, bubbles, material shortages, or weld lines), calculates the position coordinates of the defects in the image, and gives the confidence level of this judgment.

[0062] To accurately determine whether there are actually defects in the image, a dynamic threshold determination mechanism is adopted. The dynamic decision threshold is dynamically adjusted according to the performance of the model in actual applications and the data distribution. When the comprehensive confidence level exceeds the dynamic decision threshold, it is considered that there are defects in the image. At this time, an audible and visual alarm is triggered to attract the attention of the staff. At the same time, the system records the defect image and coordinate information for subsequent analysis and processing.

[0063] Visualize the decision-making process through the model interpretation module. The Class Activation Mapping (CAM) technology is adopted, which can highlight the areas in the image that play a key role in the model's decision-making, namely the defect areas. In addition, the system will also generate an interactive detection report containing timestamps, production line workstations, and defect statistical analysis. The timestamp records the specific time of the detection, the production line workstation information can help locate the specific position where the problem occurs, and the defect statistical analysis statistically analyzes information such as the types and quantities of defects detected within a period of time, providing strong support for the optimization of the production process and quality control. Staff can view, filter, and analyze the detection report through the interactive interface to take corresponding measures in a timely manner.

[0064] Embodiment 2: The present invention provides a method for identifying surface defects of automotive plastic parts based on vision. The method further includes constructing a self-evolving model update mechanism to achieve continuous learning and dynamic optimization.

[0065] In this embodiment, the surface images of automotive plastic parts in actual production are captured in real time by cameras deployed on the automotive plastic part production line to construct an incremental dataset update channel. The acquisition terminal is built with an edge computing unit, and a multi-level cache queue is used to manage newly acquired images. First, online quality screening is performed. The pre-trained lightweight ResNet-18 network is used to score the image clarity and lighting uniformity, and blurred, overexposed, or underexposed images are excluded. Only qualified images with a quality score higher than 0.85 are retained and put into the queue to be processed. After the qualified images are processed through the same preprocessing process as in Embodiment 1, the automatic annotation process is triggered, and a semi-supervised annotation strategy is adopted: for defect areas with a model prediction confidence higher than 0.95, the model prediction results are directly used as pseudo-labels; for prediction results with a confidence of 0.7 - 0.95, uncertain samples are screened through an improved active learning algorithm and pushed to the manual annotation interface for review and annotation; samples with a confidence lower than 0.7 trigger the expert review mechanism for manual secondary confirmation. The annotated incremental data dynamically updates the training set in a sliding window mechanism, retaining the effective data of the most recent 30 days and randomly retaining 20% of the historical representative samples to prevent the performance degradation of the model caused by data distribution drift.

[0066] In the model update phase, a dual-channel verification strategy is adopted to construct an A / B test environment with physical isolation. When the incremental dataset accumulates to reach a preset threshold (usually 15% of the scale of the original training set), the model fine-tuning process is initiated: First, freeze the encoder part of the generator network, and only update the parameters of the decoder and classification modules. Adopt a progressive learning rate strategy, set the initial learning rate to 0.0001, and adjust it according to the cosine annealing algorithm after each round of training. During the fine-tuning process, introduce the Elastic Weight Consolidation (EWC) regularization term, calculate the Fisher information matrix of the key parameters on the old task, and constrain the parameter update direction to maintain the original knowledge. Every 5 rounds of fine-tuning, cross-validation is initiated. The current model (Model B) and the online production model (Model A) are tested in parallel on an independent validation set. The validation set contains 2000 balanced samples covering various defects, and the evaluation metrics comprehensively include three dimensions: the FID value, mAP, and false detection rate. When the mAP of Model B relatively increases by more than 1.2%, the FID decreases by more than 0.8%, and the increase rate of the false detection rate does not exceed 0.5% in three consecutive validations, the model hot-swap protocol is triggered, and Model B is gradually replaced by Model A through the weight smooth migration technology. The sliding average algorithm is used in the migration process to complete the progressive update of the model parameters within 10 batches to ensure the uninterrupted online recognition service. If the performance of Model B does not meet the upgrade standard, a rollback mechanism is initiated, Model A is retained, and the performance bottleneck is analyzed. The samples of the corresponding failure mode are added to the enhanced training set, and the update process is triggered again after the next round of incremental data accumulation is completed.

[0067] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.

Claims

1. A method for identifying surface defects of automobile plastic parts based on vision, characterized in that: The method comprises: Step S100: acquiring historical automobile plastic part surface images, preprocessing the images, and establishing an automobile plastic part surface dataset; Step S200: annotating the automobile plastic parts surface dataset and dividing it into a training set and a validation set; Step S300: Building a deep learning model based on a generative adversarial network; Step S400: inputting the training set into a deep learning model based on a generative adversarial network for training, and using a validation set for validation during the training process to obtain a surface defect recognition model for automobile plastic parts; Step S500: collecting the surface image of the automobile plastic part in real time, and inputting the automobile plastic part surface defect recognition model to obtain the automobile plastic part surface defect recognition result.

2. The method for identifying surface defects of automobile plastic parts according to claim 1, characterized in that: The historical automobile plastic parts surface images are collected by a high-resolution linear array CCD camera, and include plastic parts surface images taken under different lighting conditions and at multiple angles, and cover sample data within a production cycle of at least 12 months.

3. The method for identifying surface defects of automobile plastic parts according to claim 2, characterized in that: The pre-processing comprises: Use bilateral filtering algorithm to reduce image noise; Enhance image contrast through histogram equalization; Scale the image to 512×512 pixel resolution.

4. The method for identifying surface defects of automobile plastic parts according to claim 1, characterized in that: LabelImg is used to annotate the preprocessed automotive plastic parts surface dataset, and the annotation types include four types of defects: scratches, bubbles, missing materials, and weld lines.

5. The method for identifying surface defects of automobile plastic parts according to claim 1, characterized in that: The deep learning model based on the generative adversarial network includes: The generator network uses a U-Net architecture and contains an 8-layer downsampling-upsampling structure to gradually extract and restore the surface image features of automotive plastic parts. Each layer integrates a residual block and a self-attention mechanism module. The 3rd to 5th layers are embedded with Transformer encoder units to capture long-distance dependencies and learn the feature associations of cross-region defects. The output uses a spectral normalization convolution layer to generate a 512×512 high-resolution feature map. The discriminator network builds a multi-scale pyramid architecture, which includes three parallel convolutional streams, processing original resolution, 1 / 2 downsampled and 1 / 4 downsampled images respectively, to capture defects of different scales on the surface of automotive plastic parts. Each branch uses deep separable convolution and adaptive instance normalization technology to adapt to different image feature distributions. The end integrates multi-scale image features through a gated fusion module; The classification module consists of a parallel attention mechanism layer and a fully connected layer. The attention mechanism layer adopts a dynamic multi-head self-attention mechanism, sets 12 attention heads and integrates a channel attention module to analyze the surface defect characteristics of automotive plastic parts from different angles. The fully connected layer adopts a hybrid expert system architecture, which includes 32 expert networks and a learnable gated controller to determine the defect type.

6. The method for identifying surface defects of automobile plastic parts according to claim 5, characterized in that: The generator network specifically includes: The input layer uses a 4-channel tensor with position encoding to mark the defect position features; The downsampling stage contains 4 residual blocks, each of which consists of 2 3×3 spectral normalization convolutional layers, with a windowed self-attention module based on Swin Transformer inserted in between; The bottleneck layer integrates three cascaded Transformer encoding units, each of which contains 12 attention heads and a 2048-dimensional feedforward network; Sub-pixel convolution technology is used in the upsampling stage, combined with adaptive instance normalization to perform multi-level image feature fusion; The output layer uses the tanh activation function and combines the channel attention gating mechanism to generate a synthetic image with enhanced local details.

7. The method for identifying surface defects of automobile plastic parts according to claim 6, characterized in that: The discriminator network specifically includes: The multi-scale input branch uses adaptive resampling technology to dynamically adjust the resolution of the image of the surface of automobile plastic parts input by each branch; Each convolutional stream consists of five stages, each of which uses a combination of depthwise separable convolution and adaptive instance normalization, and uses the LeakyReLU activation function to extract multi-scale features of defects of different sizes and shapes on the surface of automotive plastic parts; The feature fusion module uses a topological structure generated by differentiable neural architecture search, integrates a cross-scale feature interaction mechanism, and fuses multi-scale features of defects of different sizes and shapes on the surface of automotive plastic parts; The discriminant head includes a region focusing mechanism, which applies a 3x weight to key regions through a saliency map generated in real time; The training process adopts the Wasserstein loss function, which integrates gradient penalty and spectral regularization constraints to improve the training stability and convergence speed of the discriminator network.

8. The method for identifying surface defects of automobile plastic parts according to claim 7, characterized in that: The classification module specifically includes: The attention mechanism layer adopts a spatiotemporal joint attention architecture, which includes 12 dynamically configured attention heads, of which 4 heads are dedicated to spatial dimension interactions, 6 heads process inter-channel relationships, and 2 heads analyze cross-layer feature associations to mine the associations between different channel features and defect types; The channel attention module integrates a dynamic channel compression mechanism, and the compression rate is automatically adjusted according to the input features to highlight the features related to defects; The hybrid expert system of the fully connected layer contains 32 heterogeneous expert networks, of which 16 experts use a 3-layer MLP structure to process local features of defects, 8 experts use a graph convolutional network to classify defect features, and 8 experts use a time series analysis module to analyze defect change trends; The output layer uses dynamic weight fusion technology to integrate the outputs of the attention branch and the fully connected branch through a learnable gating parameter α; Label smoothing regularization and contrastive learning auxiliary loss are introduced during training, and the classification decision threshold is dynamically adjusted according to the performance of the validation set.

9. The method for identifying surface defects of automobile plastic parts according to claim 1, characterized in that: The step S400 specifically includes: Set the training parameters, including the Adam optimizer with an initial learning rate of 0.002, a batch size of 32, a maximum training round of 200 rounds, and configure the early stopping mechanism to terminate the training when the validation set loss does not decrease for 10 consecutive rounds; An alternating training strategy is adopted. In the first round, the parameters of the generator network are frozen and the discriminator network is trained using the Wasserstein loss function. In the second round, the parameters of the discriminator network are frozen and the generator network is optimized by combining adversarial loss, L1 reconstruction loss and perceptual loss. After every 5 rounds of training, the validation set evaluation is started, and the FID index, defect classification accuracy, and intersection over union ratio of the generated image and the real image are calculated synchronously. When the FID is lower than 15 and the IoU is higher than 0.85, the model snapshot is saved; After training is completed, the generative adversarial network is compressed into a lightweight recognition model through knowledge distillation technology, retaining more than 97% of the original model accuracy.

10. The method for identifying surface defects of automobile plastic parts according to claim 8, characterized in that: The step S500 specifically includes: The high-resolution linear array CCD camera deployed on the production line captures 2048×2048 original resolution images at a frame rate of 1 per second and performs pre-processing operations in real time. The preprocessed image is input into the automotive plastic surface defect recognition model. The generator network outputs a 512×512 feature map with a defect heat map. The discriminator network generates multi-scale confidence scores. The classification module combines the two types of outputs to generate structured data including defect type, location coordinates and confidence. Adopting dynamic threshold judgment mechanism, when the comprehensive confidence exceeds the dynamic decision threshold, the sound and light alarm is triggered and the defect image and defect coordinate information are recorded; The model interpretation module visualizes the decision basis, uses class activation mapping technology to highlight defective areas, and generates interactive inspection reports with timestamps, production line positions, and defect statistics.

Citation Information

Patent Citations

  • Product surface defect detection method and device based on deep learning and machine vision

    CN111445471A

Cited By

  • Visual inspection system and method for tiny flaws of industrial products

    CN120612327A

  • Surface flaw detection method and system for PE carpet

    CN120635056A

  • Ultra-clean barrel production quality inspection system and method based on machine vision and AI

    CN121032994A

  • Model automatic generation method based on bottle body label quality detection

    CN121169925A

  • Robot-assisted automobile injection molding part defect sorting method and system

    CN121424608A