Automatic Detection Method and System for Surface Defects of Substrate Glass Based on Machine Vision
By using multimodal data processing and graph neural network technology in the detection of substrate glass surface defects and combining cloud models for defect identification, the problems of inconsistent detection results and poor accuracy in the existing technology are solved, and efficient and accurate defect detection is achieved.
Patent Information
- Application Number
- CN202411490084.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-10-24
AI Technical Summary
The prior art has problems such as inconsistent detection of substrate glass surface defects and poor accuracy in the detection of substrate glass, which is difficult to meet the efficient and accurate detection needs of modern production.
Using a machine vision-based method, by obtaining multimodal data (image, acoustic, thermal imaging) on the substrate glass surface, pre-processing and feature extraction, combining graph neural networks and multi-task learning networks, topological relationships between pixels are established, structural and context information of defects are extracted, and precise identification is used using cloud models.
It realizes efficient and accurate detection of defects on the surface of the substrate glass, improves the accuracy and robustness of the detection, and can more accurately detect and locate complex defect forms.
Smart Images

Figure CN119006469B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to defect detection technology, and particularly to an automatic detection method and system for surface defects of substrate glass based on machine vision. Background Art
[0002] Substrate glass is an important component of panel display devices, and its surface quality directly affects the display effect. During the production process, various defects such as scratches, particles, and stains are likely to appear on the surface of substrate glass. These defects will seriously affect the display quality and even lead to the scrapping of the entire display device. The traditional manual visual inspection method is time-consuming and laborious, and it is difficult to ensure the consistency and accuracy of the inspection results. Therefore, there is an urgent need for an efficient and accurate automatic detection method to meet the requirements of modern production.
[0003] The development of machine vision technology provides strong support for automatic defect detection. Existing defect detection methods mainly include traditional methods based on image processing and intelligent methods based on deep learning. Traditional methods usually use manually designed image processing algorithms such as image filtering and edge detection to locate and classify defect areas. However, these methods are sensitive to noise and complex backgrounds and are difficult to adapt to variable defect morphologies.
[0004] In recent years, deep learning technology has made breakthrough progress in the fields of image processing and pattern recognition, bringing new opportunities to the defect detection task. Intelligent methods based on deep learning can automatically learn feature representations from a large amount of data and have strong adaptability to complex defect morphologies. However, existing deep learning methods still have certain deficiencies in defect location, multi-scale feature fusion, and model optimization, and are difficult to meet the high-precision requirements of substrate glass defect detection. Therefore, an innovative automatic defect detection method and system are needed to effectively solve the above problems and achieve efficient and accurate detection of surface defects of substrate glass. Summary of the Invention
[0005] Embodiments of the present invention provide an automatic detection method and system for surface defects of substrate glass based on machine vision, which can solve the problems in the prior art.
[0006] In the first aspect of the embodiments of the present invention,
[0007] An automatic detection method for surface defects of substrate glass based on machine vision is provided, including:
[0008] Obtain multimodal data on the surface of the substrate glass and perform preprocessing. The multimodal data includes image data, acoustic data, and thermal imaging data. Feature extraction is performed on the preprocessed image data, preprocessed acoustic data, and preprocessed thermal imaging data respectively. Cross-modal feature fusion is performed on the extracted image features, acoustic features, and thermal imaging features to obtain a fused feature map;
[0009] Input the fused feature map into a pre-constructed graph neural network. Through node embedding and graph convolution operations, establish the topological relationship between pixels in the fused feature map, extract the structure and context information of the defect, and obtain a topological feature map representing the defect structure. Concatenate the topological feature map with the fused feature map and input it into a multi-task learning network. Use the multi-task learning network to simultaneously complete the segmentation of the defect area and the preliminary classification of the defect type in the concatenated feature map, and obtain the segmented suspected defect area and the corresponding defect type;
[0010] For the segmented suspected defect area, use a pre-trained object detection model to perform defect localization, determine the exact position of the defect through bounding box regression, and obtain an image of the target defect area. Upload the image of the target defect area to the cloud, and use the defect detection model deployed on the cloud to accurately identify the defect type of the target defect area to obtain the final defect detection result.
[0011] In an alternative embodiment,
[0012] Input the fused feature map into a pre-constructed graph neural network model. Through node embedding and graph convolution operations, establish the topological relationship between pixels in the fused feature map, and extract the structure and context information of the defect. Obtaining a topological feature map representing the defect structure includes:
[0013] Divide the fused feature map into multiple local regions. Each local region corresponds to a node in the graph. By performing average pooling on the pixel features within each local region, obtain node features;
[0014] Construct an adjacency matrix of the graph according to the similarity between node features. The elements of the adjacency matrix represent the topological connection relationship between nodes;
[0015] Construct positive and negative sample pairs. Among them, the positive sample pairs are composed of different pixel features in the same local region, and the negative sample pairs are composed of graph representations of different local regions. By minimizing the distance between positive sample pairs and maximizing the distance between negative sample pairs, design a graph contrast loss function. Based on the obtained graph contrast loss function and graph classification loss function, construct a joint loss function. Based on the joint loss function, update the parameters of the graph neural network model through a gradient optimization algorithm, and finally obtain a trained graph neural network model;
[0016] Update and propagate the node features using the trained graph neural network model, and learn the high-level feature representation of the nodes by aggregating the neighborhood information of the nodes;
[0017] Adopt graph pooling operation to aggregate the high-level features of the nodes into graph-level feature representations, extract the structural and contextual information of the defects, and obtain the topological feature map representing the defect structure.
[0018] In an alternative embodiment,
[0019] Based on the obtained graph contrast loss function and graph classification loss function, the calculation formula for constructing the joint loss function is as follows:
[0020] ;
[0021] where L represents the joint loss function, B represents the number of graph samples, C represents the number of categories, y i,k represents the true label of g i belonging to category k, p(·) represents the predicted probability distribution, g i represents the node features, λ represents the balance factor, s(·) represents the similarity, z i represents the i-th node feature in the positive sample pair, represents another node feature in the same local region as z i , τ represents the temperature parameter, represents the j-th node feature in the negative sample pair.
[0022] In an alternative embodiment,
[0023] For the segmented suspected defect regions, use the pre-trained object detection model to locate the defects, and determine the exact positions of the defects through bounding box regression, and obtain the target defect region images including:
[0024] Obtain the suspected defect region images, adjust the suspected defect region images to a fixed size, and obtain the adjusted suspected defect region images;
[0025] Input the adjusted suspected defect region images into the pre-trained object detection model, extract the multi-scale feature maps of the defect region images through the backbone network, and use the feature pyramid structure to fuse the multi-scale feature maps to obtain the fused feature maps;
[0026] Introduce the spatial attention weight map and perform element-wise multiplication with the fused feature maps to obtain the spatially attention-enhanced feature maps. At the same time, introduce the channel attention weight vector and perform channel-wise multiplication with the spatially attention-enhanced feature maps to obtain the enhanced target feature maps;
[0027] The enhanced target feature map is input into the detection head for defective target detection. The detection head generates preliminary detection results by predicting the coordinates of the target center point, the width and height of the bounding box, and the target confidence level;
[0028] Perform non-maximum suppression processing on the preliminary detection results, filter redundant detection boxes, obtain the filtered detection results, perform bounding box regression optimization on the filtered detection results, determine the precise location of the defect, and obtain the final target defect area image.
[0029] In an alternative embodiment,
[0030] Before defect localization using the pre-trained target detection model for the segmented suspected defect area, the steps include:
[0031] Obtain labeled training samples containing the class information and location information of the defect area, and perform data augmentation processing to obtain enhanced training samples;
[0032] Input the enhanced training samples into the target detection model. The target detection model includes a feature extraction network, a region proposal network, and a detection head. The feature extraction network uses a deep convolutional neural network to extract multi-scale feature maps of the defect area. The region proposal network generates candidate regions through a region proposal algorithm. The detection head uses a fully connected layer to classify and regress the candidate regions;
[0033] According to the prediction results and the ground truth labels of the target detection model, calculate the difference between the minimum enclosing region and the union of the predicted box and the ground truth box. At the same time, introduce the weight factor and modulation factor of the sample difficulty level to dynamically adjust the weights of positive and negative samples, and construct the target detection loss function;
[0034] Use the stochastic gradient descent algorithm to iteratively optimize the parameters of the target detection model by minimizing the target detection loss function, enabling the model to learn the feature representations of bounding box regression and difficult-to-classify samples. Repeat the iteration until the preset termination condition is met to obtain the trained target detection model.
[0035] In an alternative embodiment,
[0036] Upload the target defect area image to the cloud, and use the defect detection model deployed in the cloud to accurately identify the defect type of the target defect area. The final defect detection results include:
[0037] Upload the target defect area image to the cloud server;
[0038] The cloud server uses the deployed feature extraction model to extract features from the received defect area image to obtain the multi-scale high-level semantic feature vector of the defect image;
[0039] Input the extracted multi-scale high-level semantic feature vectors into the first model based on the attention mechanism. Through the attention mechanism, adaptively fuse the multi-scale features to obtain the fused defect feature vectors, and at the same time output the soft labels of the defect types;
[0040] Using the output soft labels of the defect types as the supervision signals, train the second model ensemble by minimizing the difference between the prediction results of the second model ensemble and the soft labels of the defect types;
[0041] Input the fused defect feature vectors into the trained second model ensemble, and combine the defect type prediction results output by each second model to obtain the final defect type recognition result.
[0042] In an alternative embodiment,
[0043] Input the extracted multi-scale high-level semantic feature vectors into the first model based on the attention mechanism. Through the attention mechanism, adaptively fuse the multi-scale features to obtain the fused defect feature vectors, and at the same time output the soft labels of the defect types, including:
[0044] Introduce a normalization factor, and the calculation formula for the soft labels of the defect types is as follows:
[0045] ;
[0046] where, T represents the soft label of the defect type, N represents the total number of defect types, t a represents the probability of the a-th defect type, and t d represents the probability of the d-th defect type;
[0047] Using the output soft labels of the defect types as the supervision signals, training the second model ensemble by minimizing the difference between the prediction results of the second model ensemble and the soft labels of the defect types includes:
[0048] For each feature channel, calculate the difference between the features extracted by the teacher model and the student model on the feature channel, and at the same time introduce the KL divergence between the prediction probabilities of the first model and the second model to construct the objective function, and the calculation formula is as follows:
[0049] ;
[0050] where, J represents the objective function, α 1 represents the weight of the KL divergence term, Q represents the dimension of the prediction probability distribution, KL(·) represents the KL divergence, w s represents the prediction probability of the first model for the s-th defect category, q s represents the prediction probability of the second model for the s-th defect category, α2 The weight coefficient representing the alignment of the feature maps, R represents the number of feature channels, H m represents the height of the feature map, W m represents the width of the feature map, x l represents the l th sample in the sample set, represents the feature extracted by the first model for the sample x l on the m-th channel, represents the feature extracted by the second model for the sample x l on the m-th channel, ||·|| represents the norm.
[0051] In the second aspect of the embodiments of the present invention,
[0052] A substrate glass surface defect automatic detection system based on machine vision is provided, including:
[0053] A first unit for acquiring and preprocessing multi-modal data on the surface of the substrate glass, the multi-modal data including image data, acoustic data, and thermal imaging data, respectively performing feature extraction on the preprocessed image data, preprocessed acoustic data, and preprocessed thermal imaging data, and performing cross-modal feature fusion on the extracted image features, acoustic features, and thermal imaging features to obtain a fused feature map;
[0054] A second unit for inputting the fused feature map into a pre-constructed graph neural network, establishing the topological relationship between pixels in the fused feature map through node embedding and graph convolution operations, extracting the structure and context information of the defect to obtain a topological feature map representing the defect structure; cascading the topological feature map with the fused feature map and inputting it into a multi-task learning network, and using the multi-task learning network to simultaneously complete the segmentation of the defect region and the preliminary classification of the defect type in the cascaded feature map to obtain the segmented suspected defect region and the corresponding defect type;
[0055] A third unit for, for the segmented suspected defect region, using a pre-trained object detection model to perform defect localization, determining the precise position of the defect through bounding box regression to obtain a target defect region image; uploading the target defect region image to the cloud, and using a defect detection model deployed on the cloud to accurately identify the defect type of the target defect region to obtain the final defect detection result.
[0056] In the third aspect of the embodiments of the present invention,
[0057] An electronic device is provided, including:
[0058] A processor;
[0059] A memory for storing processor-executable instructions;
[0060] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0061] In the fourth aspect of the embodiments of the present invention,
[0062] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.
[0063] In this embodiment, by fusing data of three modalities, namely images, acoustics, and thermal imaging, the defect information on the surface of the substrate glass can be characterized more comprehensively, improving the accuracy and robustness of defect detection. Different modality data can complement different side features of defects, which is beneficial to improving the performance of defect detection. Using a graph neural network to establish the topological relationship between pixels can effectively extract the structural and contextual information of defects, capturing the overall features of defects, rather than being limited to local pixel information. This helps to more accurately detect and locate complex defect morphologies. Dividing the defect detection task into two stages, this end-cloud collaborative processing mode can balance computing resources and detection accuracy, complete preliminary processing on edge devices, reduce the computing pressure on the cloud, and at the same time use the powerful computing power of the cloud for more refined defect recognition. By simultaneously completing defect region segmentation and preliminary classification through a multi-task learning network, the parameters of the feature extractor can be shared, promoting each other, and improving the generalization ability and detection efficiency of the model. Using an object detection model to accurately locate the defect region and then combining it with the defect detection model on the cloud for precise identification can obtain more accurate defect type and location information, providing a reliable basis for subsequent defect processing. Description of the Drawings
[0064] Figure 1 It is a schematic flowchart of the method for automatically detecting defects on the surface of substrate glass based on machine vision according to the embodiments of the present invention;
[0065] Figure 2 It is a schematic structural diagram of the system for automatically detecting defects on the surface of substrate glass based on machine vision according to the embodiments of the present invention. Detailed Embodiments
[0066] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0067] The technical solution of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0068] Figure 1 FIG. is a schematic flow chart of an automatic substrate glass surface defect detection method based on machine vision according to an embodiment of the present invention. As Figure 1 shown, the method includes:
[0069] S101. Obtain multi-modal data on the surface of the substrate glass and perform preprocessing. The multi-modal data includes image data, acoustic data, and thermal imaging data. Feature extraction is respectively performed on the preprocessed image data, preprocessed acoustic data, and preprocessed thermal imaging data, and cross-modal feature fusion is performed on the extracted image features, acoustic features, and thermal imaging features to obtain a fused feature map.
[0070] Exemplarily, first, obtain multi-modal data on the surface of the substrate glass, including image data, acoustic data, and thermal imaging data. Among them, the image data is collected by a high-resolution industrial camera, the acoustic data is collected by an ultrasonic sensor, and the thermal imaging data is collected by an infrared thermal imager. For the image data, the preprocessing steps include: image denoising, image enhancement, image correction, etc. Among them, the image denoising adopts a method combining wavelet transform and adaptive median filtering, which can effectively suppress Gaussian noise and salt-and-pepper noise; the image enhancement adopts an adaptive histogram equalization algorithm to improve the contrast and clarity of the image; the image correction adopts a perspective transformation method based on corner detection to eliminate the geometric distortion of the image. For the acoustic data, the preprocessing steps include: signal filtering, signal enhancement, feature extraction, etc. Among them, the signal filtering adopts a method of wavelet packet decomposition and reconstruction, which can effectively remove high-frequency noise in the acoustic signal; the signal enhancement adopts an envelope extraction method based on Hilbert transform to highlight the amplitude characteristics of the acoustic signal; the feature extraction adopts a short-time Fourier transform (STFT) method to obtain the time-frequency characteristics of the acoustic signal. For the thermal imaging data, the preprocessing steps include: thermal map denoising, thermal map enhancement, thermal map segmentation, etc. Among them, the thermal map denoising adopts an adaptive guided filtering algorithm, which can effectively suppress high-frequency noise in the thermal imaging data; the thermal map enhancement adopts an adaptive plateaus histogram equalization method to improve the contrast and clarity of the thermal map; the thermal map segmentation adopts a binary method based on Otsu threshold to extract the high-temperature region in the thermal map.
[0071] Feature extraction is performed on the preprocessed image data, acoustic data, and thermal imaging data respectively. Among them, for image feature extraction, a multi-scale residual convolutional neural network is used. By designing receptive fields and residual connections of different scales, multi-scale texture and structural features of the image are extracted; for acoustic feature extraction, a long short-term memory network based on the attention mechanism is used. By introducing an attention mechanism between the encoder and decoder of the LSTM, the temporal features of the acoustic signal are adaptively extracted; for thermal imaging feature extraction, a heatmap pyramid pooling convolutional neural network is used. Through spatial pyramid pooling operations, multi-scale regional features of the heatmap are extracted.
[0072] Cross-modal feature fusion is performed on the extracted image features, acoustic features, and thermal imaging features. First, the image features, acoustic features, and thermal imaging features are aligned to the same scale and dimension through feature mapping; then, a self-attention fusion module is used to perform adaptive weighted fusion on the aligned features. By calculating the correlation between features of different modalities, a fusion weight matrix is generated to achieve cross-modal interaction and fusion of features; finally, the fused features are compressed and refined through convolutional layers and pooling layers to obtain a compact fused feature map.
[0073] In this embodiment, for data of three different modalities, namely images, acoustics, and thermal imaging, suitable preprocessing methods are respectively adopted, such as denoising, enhancement, correction, etc., to improve the quality and feature separability of each modality data. The self-attention fusion module is used to adaptively generate a fusion weight matrix according to the correlation between features of different modalities, realizing the interaction and fusion of features, and making full use of the complementary advantages of multi-modal data. Through technical means such as multi-modal data preprocessing, deep learning feature extraction, cross-modal feature fusion, and edge-cloud collaborative processing, the rich information of image, acoustic, and thermal imaging data can be fully explored and utilized, improving the accuracy, robustness, and efficiency of defect detection, and achieving high-precision detection and classification of defects on the surface of substrate glass.
[0074] S102. Input the fused feature map into a pre-constructed graph neural network. Through node embedding and graph convolution operations, establish the topological relationship between pixels in the fused feature map, extract the structural and contextual information of the defect, and obtain a topological feature map representing the defect structure; cascade the topological feature map with the fused feature map and input it into a multi-task learning network. Use the multi-task learning network to simultaneously complete the segmentation of the defect region and the preliminary classification of the defect type in the cascaded feature map, and obtain the segmented suspected defect region and the corresponding defect type.
[0075] Among them, the multi-task learning network adopts an encoder-decoder architecture. The encoder part consists of several convolutional layers and pooling layers, which are used to extract high-level semantic features of the cascaded feature maps. The convolutional layers use 3×3 convolutional kernels. By performing convolutional operations on the feature maps, local features are extracted and the receptive field is enlarged. The pooling layers use 2×2 max pooling. By downsampling the feature maps, the spatial resolution of the feature maps is reduced, and at the same time, the receptive field is increased. After multiple convolutions and poolings, the spatial resolution of the feature maps continuously decreases, but the feature dimension continuously increases, obtaining high-level semantic features with strong discriminative power. The decoder part consists of several upsampling layers and transposed convolutional layers, which are used to restore the spatial resolution of the feature maps. The upsampling layers use the nearest neighbor interpolation method to double the spatial resolution of the feature maps. The transposed convolutional layers use 3×3 convolutional kernels. By performing convolutional operations on the upsampled feature maps, local features are extracted and the feature representation is refined. After multiple upsamplings and transposed convolutions, the spatial resolution of the feature maps gradually returns to the same as that of the input image, obtaining a segmentation result with the same size as the input image.
[0076] To achieve multi-task learning, two parallel fully connected layers are introduced at the last layer of the encoder, corresponding to the defect region segmentation task and the defect type classification task respectively. The fully connected layer for the segmentation task maps the high-level semantic features to a pixel-level segmentation result, predicting the probability that each pixel belongs to the defect region. The fully connected layer for the classification task maps the high-level semantic features to a class-level classification result, predicting the probability that the input image belongs to each defect type.
[0077] The strategy of gradient averaging is adopted for multi-task learning. By jointly optimizing the segmentation task and the classification task, end-to-end defect segmentation and classification are achieved. During the training process, the cascaded feature maps and the corresponding segmentation labels and class labels are input into the multi-task learning network, the loss functions of the segmentation task and the classification task are calculated, and the network parameters are updated by gradient averaging. After multiple rounds of iterative training, a multi-task learning network with the ability to segment defect regions and classify defect types simultaneously is obtained.
[0078] During the inference process, the cascaded feature map of the substrate glass surface image to be measured is input into the trained multi-task learning network, and the segmentation result and classification result are obtained through forward propagation. The segmentation result is a probability map with the same size as the input image, indicating the probability that each pixel belongs to the defect area. The probability map is binarized to obtain the suspected defect area after segmentation. The classification result is a multi-dimensional probability vector, indicating the probability that the input image belongs to each defect type. The category with the highest probability is taken as the final defect type. Through the inference of the multi-task learning network, the suspected defect area on the substrate glass surface and the corresponding defect type can be obtained simultaneously, realizing the automatic segmentation and recognition of the defect area. Compared with training the segmentation network and classification network separately, the multi-task learning network can make full use of the correlation between different tasks, promote each other between tasks through joint optimization, and improve the accuracy and efficiency of defect segmentation and classification.
[0079] In an alternative embodiment,
[0080] The fused feature map is input into a pre-constructed graph neural network model. Through node embedding and graph convolution operations, the topological relationship between pixels in the fused feature map is established, and the structural and contextual information of the defect is extracted. The topological feature map representing the defect structure includes:
[0081] The fused feature map is divided into multiple local regions, and each local region corresponds to a node in the graph. By performing average pooling on the pixel features within each local region, the node features are obtained;
[0082] An adjacency matrix of the graph is constructed according to the similarity between the node features. The elements of the adjacency matrix represent the topological connection relationship between nodes;
[0083] Positive and negative sample pairs are constructed. Among them, the positive sample pairs are composed of different pixel features in the same local region, and the negative sample pairs are composed of graph representations in different local regions. By minimizing the distance between positive sample pairs and maximizing the distance between negative sample pairs, a graph contrast loss function is designed. Based on the obtained graph contrast loss function and graph classification loss function, a joint loss function is constructed. Based on the joint loss function, the parameters of the graph neural network model are updated through the gradient optimization algorithm, and finally the trained graph neural network model is obtained;
[0084] The trained graph neural network model is used to update and propagate the node features, and the high-level feature representation of the nodes is learned by aggregating the neighborhood information of the nodes;
[0085] Graph pooling operations are adopted to aggregate the high-level features of the nodes into graph-level feature representations, extract the structural and contextual information of the defect, and obtain the topological feature map representing the defect structure.
[0086] Exemplarily, the fused feature map is divided into multiple local regions, and each local region corresponds to a node in the graph. For each local region, a node feature vector is obtained by performing average pooling operation on the pixel features therein. Based on the similarity between node features, the adjacency matrix of the graph is constructed. Similarity metrics (such as Euclidean distance, cosine similarity, etc.) can be used to calculate the similarity between node features, and the similarity is transformed into the elements of the adjacency matrix, representing the topological connection relationship between nodes.
[0087] Positive and negative sample pairs are constructed. The positive sample pairs are composed of different graph representations of the same defect region, and the negative sample pairs are composed of graph representations of different defect regions. Specifically, for a given defect region, multiple different graph representations are generated through data augmentation techniques (such as rotation, flipping, cropping, etc.) as positive sample pairs. For different defect regions, their graph representations are randomly selected as negative sample pairs.
[0088] Based on the obtained graph contrast loss function and graph classification loss function, a joint loss function is constructed. Based on the joint loss function, the parameters of the graph neural network model are updated using a gradient optimization algorithm (such as stochastic gradient descent) to minimize the joint loss function. Through the iterative training process, the model parameters are continuously adjusted to improve the performance of the model in graph classification tasks and graph contrast tasks. The trained graph neural network model is used to update and propagate node features. By aggregating the neighborhood information of nodes, high-level feature representations of nodes are learned. Graph neural network structures such as graph convolutional networks and graph attention networks can be used to achieve the update and propagation of node features. Graph pooling operations are adopted to aggregate the high-level features of nodes into graph-level feature representations. Graph pooling operations can use methods such as graph convolution and graph attention to aggregate node features into a graph-level feature vector. In this way, the structural and contextual information of the defect can be extracted to obtain a topological feature map representing the defect structure.
[0089] In an alternative embodiment,
[0090] The calculation formula for constructing the joint loss function based on the obtained graph contrast loss function and graph classification loss function is as follows:
[0091] ;
[0092] where L represents the joint loss function, B represents the number of graph samples, C represents the number of categories, y i,k represents the true label that g i belongs to category k, p(·) represents the predicted probability distribution, g i represents the node feature, λ represents the balance factor, s(·) represents the similarity, z i represents the i-th node feature in the positive sample pair, represents the one corresponding to zi Another node feature of the same local area, where τ represents the temperature parameter, represents the j-th node feature in the negative sample pair.
[0093] In this embodiment, based on the graph structure representation method, it can effectively capture the spatial proximity relationship and structural association between pixels in the image, exceeding the limitation of traditional convolutional neural networks that only consider local neighborhoods. The graph structure-based feature learning method helps to extract the structural and contextual information of defects, capture the topological features of defects in the overall image, rather than being limited to local pixel-level features. The topological feature map not only contains the local texture and edge information of defects, but also contains the structural features and contextual information of defects in the overall image, which helps to more accurately detect and classify complex defect morphologies. Based on the contrast learning optimization strategy, the model can learn more discriminative feature representations, improve the accuracy and performance of defect detection, and thus achieve efficient detection and classification of defects on the surface of substrate glass.
[0094] S103. For the segmented suspected defect regions, use the pre-trained object detection model to locate the defects, determine the exact positions of the defects through bounding box regression, and obtain the target defect region images; upload the target defect region images to the cloud, and use the defect detection model deployed on the cloud to accurately identify the defect types of the target defect regions to obtain the final defect detection results.
[0095] Among them, the object detection model adopts a two-stage detection framework, such as Faster R-CNN or Mask R-CNN, generates candidate regions through the Region Proposal Network (RPN), and classifies and performs bounding box regression on the candidate regions. It mainly consists of three parts: the backbone network, the neck network, and the detection head. The backbone network extracts multi-scale features through residual connections and CSP (Cross Stage Partial) modules; the neck network adopts the PANet structure and realizes the fusion of different scale features through feature pyramids and adaptive feature fusion; the detection head adopts an anchor-free design and directly predicts the center point and bounding box of the target.
[0096] In an alternative embodiment,
[0097] For the segmented suspected defect regions, using the pre-trained object detection model to locate the defects and determining the exact positions of the defects through bounding box regression to obtain the target defect region images includes:
[0098] Obtain the suspected defect region image, adjust the suspected defect region image to a fixed size to obtain the adjusted suspected defect region image;
[0099] Input the adjusted image of the suspected defect area into a pre-trained object detection model. Extract multi-scale feature maps of the defect area image through the backbone network, and use the feature pyramid structure to fuse the multi-scale feature maps to obtain the fused feature map;
[0100] Introduce a spatial attention weight map and perform element-wise multiplication with the fused feature map to obtain a spatially attention-enhanced feature map. At the same time, introduce a channel attention weight vector and perform channel-wise multiplication with the spatially attention-enhanced feature map to obtain an enhanced target feature map;
[0101] Input the enhanced target feature map into the detection head for defect object detection. The detection head generates a preliminary detection result by predicting the coordinates of the target center point, the width and height of the bounding box, and the target confidence;
[0102] Perform non-maximum suppression on the preliminary detection result to filter out redundant detection boxes, obtain the filtered detection result, perform bounding box regression optimization on the filtered detection result, determine the precise location of the defect, and obtain the final target defect area image.
[0103] Exemplarily, first, obtain the image of the suspected defect area and adjust it to a fixed size to meet the input requirements of the object detection model. Usually, operations such as scaling and cropping are used to obtain the adjusted image of the suspected defect area.
[0104] Input the adjusted image of the suspected defect area into a pre-trained object detection model. The object detection model usually consists of a backbone network and a detection head. The backbone network uses a deep convolutional neural network, such as ResNet, VGG, etc., to extract multi-scale feature maps of the defect area image. To fuse feature information of different scales, a feature pyramid structure (FPN) is introduced. Through the top-down path and lateral connections, high-level semantic features are fused with low-level detail features to obtain the fused feature map.
[0105] To further enhance the representation ability of the defect region, an attention mechanism is introduced. First, by generating a spatial attention weight map and performing an element-wise product with the fused feature map, a spatially attention-enhanced feature map is obtained. The spatial attention weight map can be obtained through a convolutional layer and a Sigmoid activation function, which is used to highlight the spatial location information of the defect region. At the same time, a channel attention weight vector is introduced, and a channel-wise product is performed with the spatially attention-enhanced feature map to obtain an enhanced target feature map. The channel attention weight vector can be obtained through global average pooling, a fully connected layer, and a Sigmoid activation function, which is used to adaptively adjust the importance of different channels. The enhanced target feature map is input into the detection head for defect target detection. The detection head generates a preliminary detection result by predicting the target center point coordinates, bounding box width and height, and target confidence. Commonly used detection heads include SSD, YOLO, Faster R-CNN, etc., and the corresponding detection head is selected according to the specific model. Non-maximum suppression processing is performed on the preliminary detection result, and a non-maximum suppression algorithm (such as IOU threshold, confidence threshold) can be used to filter redundant detection boxes to obtain a filtered detection result. Boundary box regression optimization is performed on the filtered detection result, and methods such as linear regression and smooth L1 loss can be used. According to the predicted bounding box offset and scaling factor of the model, the position of the bounding box of the preliminary detection result is adjusted to obtain a more accurate defect position.
[0106] According to the optimized bounding box position, the defect region image is extracted from the original image. An image processing library (such as OpenCV) can be used to perform image cropping operations to obtain the final target defect region image.
[0107] In this embodiment, by fusing multi-scale features through a feature pyramid structure, different scale features of the defect region can be effectively captured. By predicting the target center point coordinates, bounding box width and height, and target confidence, a preliminary detection result is generated, and non-maximum suppression processing and bounding box regression optimization are performed, which can accurately locate the position and range of the defect to obtain the final target defect region image. Through the attention mechanism, the attention to the defect region and feature channels can be adaptively enhanced, improving the sensitivity of the model to defect features, thereby improving the accuracy of defect localization. Deploying the defect localization task on the edge side and using a pre-trained target detection model for defect localization can complete preliminary processing on edge devices, reducing the computational pressure on the cloud. Only after the target defect region is determined, the target defect region image is uploaded to the cloud for further accurate identification. This edge-cloud collaborative processing mode can improve the efficiency and performance of the overall system. By extracting multi-scale feature maps through the backbone network and using a feature pyramid structure to fuse multi-scale feature maps, different scale defect features can be effectively captured and fused, improving the model's detection ability for defects.
[0108] In an alternative embodiment,
[0109] Before using the pre-trained object detection model to locate the defects in the segmented suspected defect regions, the steps include:
[0110] Obtain the labeled training samples containing the class information and location information of the defect regions, and perform data augmentation processing to obtain the augmented training samples;
[0111] Input the augmented training samples into the object detection model, where the object detection model includes a feature extraction network, a region proposal network, and a detection head. The feature extraction network uses a deep convolutional neural network to extract multi-scale feature maps of the defect regions. The region proposal network generates candidate regions through a region proposal algorithm. The detection head uses a fully connected layer to classify and regress the candidate regions;
[0112] According to the prediction results and ground truth labels of the object detection model, calculate the difference between the minimum enclosing region and the union of the predicted bounding box and the ground truth bounding box. At the same time, introduce the weight factor and modulation factor of the sample difficulty level to dynamically adjust the weights of positive and negative samples, and construct the object detection loss function;
[0113] Use the stochastic gradient descent algorithm to iteratively optimize the parameters of the object detection model by minimizing the object detection loss function, so that the model learns the feature representations of bounding box regression and difficult-to-classify samples. Repeat the iteration until the preset termination condition is met to obtain the trained object detection model.
[0114] Exemplarily, first obtain the labeled training samples containing the class information and location information of the defect regions, and perform data augmentation processing on the training samples, including operations such as random flipping, random cropping, and random scaling, to obtain the augmented training samples. Input the augmented training samples into the object detection model for training. The object detection model includes a feature extraction network, a region proposal network, and a detection head. The feature extraction network uses a deep convolutional neural network to extract multi-scale feature maps of the defect regions. The region proposal network generates candidate regions through a region proposal algorithm. The detection head uses a fully connected layer to classify and regress the candidate regions. According to the prediction results and ground truth labels of the object detection model, calculate the difference between the minimum enclosing region and the union of the predicted bounding box and the ground truth bounding box, and construct the object detection loss function. At the same time, introduce the weight factor and modulation factor of the sample difficulty level to dynamically adjust the weights of positive and negative samples to strengthen the learning of difficult-to-classify samples.
[0115] Use the stochastic gradient descent algorithm to iteratively optimize the parameters of the object detection model by minimizing the object detection loss function, so that the model learns the feature representations of bounding box regression and difficult-to-classify samples. Repeat the iteration until the preset termination condition is met to obtain the trained object detection model.
[0116] In this embodiment, by performing data augmentation processing on the labeled training samples, such as flipping, rotating, scaling and other operations, the training data set can be effectively expanded, and the diversity and richness of the data can be improved. The multi-scale feature extraction strategy helps the model detect and locate defects of different sizes and shapes, and improves the generalization ability of the model. By introducing the Region Proposal Network and generating candidate regions through the region proposal algorithm, the number of regions to be processed can be effectively reduced, and the computing efficiency can be improved. The dynamic sample weight adjustment strategy enables the model to pay more attention to difficult samples, improves the learning ability of bounding box regression and difficult classification samples, and thus improves the accuracy of defect detection. Through technical means such as data augmentation, multi-scale feature extraction, region proposal algorithm, dynamic sample weight adjustment and edge-cloud collaborative processing, a high-performance object detection model can be trained to achieve accurate positioning of the defects on the substrate glass surface.
[0117] In an alternative embodiment,
[0118] Upload the target defect area image to the cloud, and use the defect detection model deployed in the cloud to accurately identify the defect type of the target defect area, and the final defect detection result includes:
[0119] Upload the target defect area image to the cloud server;
[0120] The cloud server uses the deployed feature extraction model to extract features from the received defect area image, and obtains the multi-scale high-level semantic feature vector of the defect image;
[0121] Input the extracted multi-scale high-level semantic feature vector into the first model based on the attention mechanism, adaptively fuse the multi-scale features through the attention mechanism to obtain the fused defect feature vector, and at the same time output the defect type soft label;
[0122] Using the output defect type soft label as the supervision signal, train the second model set by minimizing the difference between the prediction results of the second model set and the defect type soft label;
[0123] Input the fused defect feature vector into the trained second model set, and combine the defect type prediction results output by each second model to obtain the final defect type recognition result.
[0124] Exemplarily, an image of an object to be detected is collected by a terminal device, and the collected image of the object to be detected is transmitted to an edge server; the edge server uses a deployed multi-task learning network to segment the defect area of the received image of the object to be detected, obtaining a segmented defect area image, and uses a deployed object detection model to accurately locate the segmented defect area image, obtaining a located defect area image; the located defect area image obtained in the edge server is uploaded to a cloud server.
[0125] The cloud server uses a deployed feature extraction model to extract features from the received defect area image. The feature extraction model usually adopts a pre-trained deep convolutional neural network, such as EfficientNet, ResNet, etc., to extract multi-scale high-level semantic feature vectors of the defect image. By outputting feature maps at different convolutional layers, different-scale feature representations are obtained to capture the multi-scale semantic information of the defect area.
[0126] The extracted multi-scale high-level semantic feature vectors are input into a first model based on an attention mechanism. The first model adaptively fuses the multi-scale features through the attention mechanism, obtaining a fused defect feature vector. Specifically, by generating an attention weight vector and performing weighted summation with the multi-scale feature vectors, the importance of different-scale features is highlighted, obtaining a fused defect feature vector. At the same time, the first model also outputs a soft label of the defect type as a supervision signal for subsequent model training.
[0127] According to the soft label of the defect type output by the first model, a second model set is trained. The second model set consists of multiple independent defect type recognition models, and each model corresponds to a defect type. By minimizing the difference between the prediction results of the second model set and the soft label of the defect type, using a cross-entropy loss function or a KL divergence loss function, the second model set is trained to learn the discriminative features of the defect type.
[0128] The fused defect feature vector is input into the trained second model set for defect type recognition. Each second model independently classifies the defect feature vector and outputs the prediction probability of the corresponding defect type. By combining the prediction results of each second model, methods such as a voting mechanism or weighted averaging are used to obtain the final defect type recognition result.
[0129] In an optional embodiment,
[0130] The extracted multi-scale high-level semantic feature vectors are input into a first model based on an attention mechanism. Through the attention mechanism, the multi-scale features are adaptively fused to obtain a fused defect feature vector. At the same time, the output soft label of the defect type includes:
[0131] A normalization factor is introduced, and the calculation formula for the soft label of the defect type is as follows:
[0132] ;
[0133] where T represents the soft label of the defect type, N represents the total number of defect types, t a represents the probability of the a-th defect type, and t d represents the probability of the d-th defect type;
[0134] Using the output soft label of the defect type as the supervision signal, training the second model set by minimizing the difference between the prediction result of the second model set and the soft label of the defect type includes:
[0135] For each feature channel, calculate the difference between the features extracted by the teacher model and the student model on the feature channel, and at the same time introduce the KL divergence between the prediction probability of the first model and the prediction probability of the second model to construct the objective function. The calculation formula is as follows:
[0136] ;
[0137] where J represents the objective function, α 1 represents the weight of the KL divergence term, Q represents the dimension of the prediction probability distribution, KL(·) represents the KL divergence, w s represents the prediction probability of the first model for the s-th defect category, q s represents the prediction probability of the second model for the s-th defect category, α 2 represents the weight coefficient of feature map alignment, R represents the number of feature channels, H m represents the height of the feature map, W m represents the width of the feature map, x l represents the l -th sample in the sample set, represents the feature extracted by the first model for the sample x l on the m-th channel, represents the feature extracted by the second model for the sample x l on the m-th channel, and ||·|| represents the norm.
[0138] In this embodiment, by uploading the target defect area image to the cloud server and leveraging the powerful computing capabilities and storage resources of the cloud, more complex and high-performance defect detection models can be deployed to achieve accurate identification of defect types. The cloud server can extract multi-scale high-level semantic feature vectors from the defect area image using the deployed feature extraction model, capturing the texture, edges, and structural information of the defect at different scales. The first model based on the attention mechanism can adaptively focus on the important parts of different-scale features, improving the sensitivity to defect features and thus enhancing the accuracy of defect type identification. At the same time, the first model can also output soft labels of defect types, providing effective supervision signals for subsequent model training. By inputting the fused defect feature vectors into the trained second model ensemble and combining the output results of each model, model integration can be achieved, improving the robustness and generalization ability of defect type identification. A processing mode of edge-cloud collaboration is adopted, where preliminary defect localization is completed on the edge device, and the target defect area image is uploaded to the cloud for fine defect type identification. By leveraging techniques such as the powerful computing capabilities of the cloud, multi-scale feature extraction and fusion, attention mechanism to enhance feature representation, model integration and knowledge distillation, and edge-cloud collaboration processing, high-precision type identification of substrate glass surface defects can be achieved, providing crucial support for subsequent defect repair and quality control.
[0139] Figure 2 FIG. is a schematic structural diagram of an automatic detection system for substrate glass surface defects based on machine vision according to an embodiment of the present invention, as Figure 2 shown, the system includes:
[0140] A first unit for acquiring and preprocessing multi-modal data on the surface of the substrate glass, where the multi-modal data includes image data, acoustic data, and thermal imaging data, respectively extracting features from the preprocessed image data, preprocessed acoustic data, and preprocessed thermal imaging data, and performing cross-modal feature fusion on the extracted image features, acoustic features, and thermal imaging features to obtain a fused feature map;
[0141] A second unit for inputting the fused feature map into a pre-constructed graph neural network, establishing the topological relationship between pixels in the fused feature map through node embedding and graph convolution operations, extracting the structural and contextual information of the defect to obtain a topological feature map representing the defect structure; cascading the topological feature map with the fused feature map and inputting it into a multi-task learning network, and using the multi-task learning network to simultaneously complete the segmentation of the defect area and the preliminary classification of the defect type in the cascaded feature map to obtain the segmented suspected defect area and the corresponding defect type;
[0142] The third unit is used to locate the defects in the segmented suspected defect areas by using a pre-trained object detection model, determine the exact positions of the defects through bounding box regression, and obtain the target defect area images; upload the target defect area images to the cloud, and use the defect detection model deployed on the cloud to accurately identify the defect types of the target defect areas, so as to obtain the final defect detection results.
[0143] In the third aspect of the embodiments of the present invention,
[0144] a kind of electronic device is provided, including:
[0145] a processor;
[0146] a memory for storing instructions executable by the processor;
[0147] wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0148] In the fourth aspect of the embodiments of the present invention,
[0149] a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0150] The present invention can be a method, a device, a system and / or a computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are uploaded.
[0151] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for automatically detecting defects on the surface of a substrate glass based on machine vision, characterized in that: include: Acquiring and preprocessing multimodal data of the surface of the substrate glass, the multimodal data including image data, acoustic data and thermal imaging data, respectively extracting features from the preprocessed image data, the preprocessed acoustic data and the preprocessed thermal imaging data, and performing cross-modal feature fusion on the extracted image features, acoustic features and thermal imaging features to obtain a fused feature map; The fused feature map is input into a pre-built graph neural network, and the topological relationship between pixels in the fused feature map is established through node embedding and graph convolution operations, and the structure and context information of the defect are extracted to obtain a topological feature map representing the defect structure; the topological feature map is cascaded with the fused feature map, and input into a multi-task learning network, and the multi-task learning network is used to simultaneously complete the segmentation of the defect area and the preliminary classification of the defect type in the cascaded feature map, and obtain the segmented suspected defect area and the corresponding defect type; For the suspected defect area after segmentation, the pre-trained target detection model is used to locate the defect, and the precise position of the defect is determined by bounding box regression to obtain the target defect area image; Upload the target defect area image to the cloud, use the defect detection model deployed on the cloud to accurately identify the defect type of the target defect area, and obtain the final defect detection result; For the suspected defect area after segmentation, the pre-trained target detection model is used to locate the defect, and the precise position of the defect is determined by bounding box regression. The target defect area image is obtained, including: Acquire a suspected defect area image, adjust the suspected defect area image to a fixed size, and obtain an adjusted suspected defect area image; The adjusted suspected defect area image is input into the pre-trained target detection model, the multi-scale feature map of the defect area image is extracted through the backbone network, and the multi-scale feature map is fused using the feature pyramid structure to obtain the fused feature map; The spatial attention weight map is introduced and element-wise multiplied with the fused feature map to obtain the feature map enhanced with spatial attention. At the same time, the channel attention weight vector is introduced and channel-wise multiplied with the feature map enhanced with spatial attention to obtain the enhanced target feature map. The enhanced target feature map is input into the detection head for defect target detection. The detection head generates preliminary detection results by predicting the target center point coordinates, bounding box width and height, and target confidence. Perform non-maximum suppression processing on the preliminary detection results, filter redundant detection frames, obtain filtered detection results, perform bounding box regression optimization on the filtered detection results, determine the precise location of the defect, and obtain the final target defect area image; The target defect area image is uploaded to the cloud, and the defect detection model deployed on the cloud is used to accurately identify the defect type of the target defect area. The final defect detection results include: Upload the target defect area image to the cloud server; The cloud server uses the deployed feature extraction model to extract features from the received defect area image to obtain a multi-scale high-level semantic feature vector of the defect image; The extracted multi-scale high-level semantic feature vector is input into the first model based on the attention mechanism, and the multi-scale features are adaptively fused through the attention mechanism to obtain the fused defect feature vector, and the defect type soft label is output at the same time; The second model set is trained by minimizing the difference between the prediction result of the second model set and the soft label of the defect type according to the output soft label of the defect type as a supervisory signal; The fused defect feature vector is input into the trained second model set, and the defect type prediction results output by each second model are combined to obtain the final defect type recognition result.
2. The method according to claim 1, characterized in that The fused feature map is input into the pre-built graph neural network model. Through node embedding and graph convolution operations, the topological relationship between pixels in the fused feature map is established, the structure and context information of the defect are extracted, and the topological feature map representing the defect structure is obtained, including: The fused feature map is divided into multiple local regions, each of which corresponds to a node in the map. The node features are obtained by averaging the pixel features in each local region. The adjacency matrix of the graph is constructed based on the similarity between node features. The elements of the adjacency matrix represent the topological connection relationship between nodes. Construct positive and negative sample pairs, where the positive sample pairs are composed of different pixel features in the same local area, and the negative sample pairs are composed of graph representations of different local areas. By minimizing the distance between positive sample pairs and maximizing the distance between negative sample pairs, design a graph contrast loss function, and construct a joint loss function based on the obtained graph contrast loss function and graph classification loss function. Based on the joint loss function, update the parameters of the graph neural network model through a gradient optimization algorithm, and finally obtain a trained graph neural network model; Use the trained graph neural network model to update and propagate node features, and learn high-level feature representations of nodes by aggregating node neighborhood information; Graph pooling operation is used to aggregate high-level features of nodes into graph-level feature representation, extract the structure and context information of the defect, and obtain a topological feature map that represents the defect structure.
3. The method according to claim 2, characterized in that The calculation formula for constructing the joint loss function based on the obtained image contrast loss function and image classification loss function is as follows: Among them, L represents the joint loss function, B represents the number of graph samples, C represents the number of categories, and y i,k Indicates g i The true label of category k, p(·) represents the predicted probability distribution, g i represents node features, λ represents the balance factor, s(·) represents the similarity, z i represents the i-th node feature in the positive sample pair, Indicates that z i Another node feature of the same local area, τ, represents the temperature parameter, Represents the jth node feature in the negative sample pair.
4. The method according to claim 1, characterized in that: For the segmented suspected defect area, the steps before using the pre-trained target detection model to locate the defect include: Obtain labeled training samples containing category information and location information of defective areas, and perform data enhancement processing to obtain enhanced training samples; Inputting the enhanced training samples into the target detection model, the target detection model includes a feature extraction network, a region proposal network and a detection head, the feature extraction network uses a deep convolutional neural network to extract multi-scale feature maps of defect areas, the region proposal network generates candidate areas through a region proposal algorithm, and the detection head uses a fully connected layer to classify and regress candidate areas; According to the prediction results and true labels of the target detection model, the difference between the minimum closure area and the union of the predicted box and the true box is calculated. At the same time, the weight factor and modulation factor of the sample difficulty are introduced to dynamically adjust the weights of positive and negative samples and construct the target detection loss function. The stochastic gradient descent algorithm is used to iteratively optimize the parameters of the target detection model by minimizing the target detection loss function, so that the model can learn the feature representation of bounding box regression and difficult-to-classify samples. The iteration is repeated until the preset termination condition is met to obtain a trained target detection model.
5. The method according to claim 1, characterized in that The extracted multi-scale high-level semantic feature vector is input into the first model based on the attention mechanism. The multi-scale features are adaptively fused through the attention mechanism to obtain the fused defect feature vector. At the same time, the defect type soft labels are output including: The normalization factor is introduced, and the calculation formula for the defect type soft label is as follows: Where T represents the soft label of the defect type, N represents the total number of defect types, and t a represents the probability of the a-th defect type, t d represents the probability of the d-th defect type; According to the output defect type soft label as the supervisory signal, the second model set is trained by minimizing the difference between the prediction result of the second model set and the defect type soft label, including: For each feature channel, the difference between the features extracted by the teacher model and the student model on the feature channel is calculated, and the KL divergence between the predicted probability of the first model and the predicted probability of the second model is introduced to construct the objective function. The calculation formula is as follows: Where J represents the objective function, α1 represents the weight of the KL divergence term, Q represents the dimension of the predicted probability distribution, KL(·) represents the KL divergence, and w s represents the predicted probability of the first model for the s-th defect category, q s represents the predicted probability of the second model for the s-th defect category, α2 represents the weight coefficient of feature map alignment, R represents the number of feature channels, and H m Represents the height of the feature map, W m Indicates the feature map width, x l represents the lth sample in the sample set, φ m (x l ) represents the first model for sample x l The features extracted on the mth channel, Represents the second model for sample x l The features extracted on the mth channel, ||·|| represents the norm.
6. A system for automatically detecting defects on the surface of a substrate glass based on machine vision, used to implement the method described in any one of claims 1 to 5, characterized in that: include: The first unit is used to obtain and preprocess multimodal data of the surface of the substrate glass, wherein the multimodal data includes image data, acoustic data and thermal imaging data, and to extract features from the preprocessed image data, the preprocessed acoustic data and the preprocessed thermal imaging data, respectively, and to perform cross-modal feature fusion on the extracted image features, acoustic features and thermal imaging features to obtain a fused feature map; The second unit is used to input the fused feature map into a pre-built graph neural network, establish the topological relationship between pixels in the fused feature map through node embedding and graph convolution operations, extract the structure and context information of the defect, and obtain a topological feature map that characterizes the defect structure; cascade the topological feature map with the fused feature map, input them into a multi-task learning network, and use the multi-task learning network to simultaneously complete the segmentation of the defect area in the cascaded feature map and the preliminary classification of the defect type, and obtain the segmented suspected defect area and the corresponding defect type; The third unit is used to locate the defect in the segmented suspected defect area using a pre-trained target detection model, determine the exact position of the defect through bounding box regression, and obtain the target defect area image; The target defect area image is uploaded to the cloud, and the defect detection model deployed on the cloud is used to accurately identify the defect type of the target defect area to obtain the final defect detection result.
7. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 5.
8. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Surface defect detection method based on multi-scale information fusion
CN113610822A
Glass panel surface defect detection method based on small sample learning
CN114092389A
Cited By
Food defect real-time detection method and system based on image processing
CN120783334A
Real-time Food Defect Detection Method and System Based on Image Processing
CN120783334B