Aerogel felt surface defect identification method and device
Through the combination of deep convolutional neural network and SVM classifier, the problem of surface defect detection of aerogel wool felt is solved, efficient and accurate defect recognition is achieved, and it is suitable for quality control of aerogel wool felt.
Patent Information
- Application Number
- CN202510740003.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-04
AI Technical Summary
The prior art is difficult to effectively detect and identify defects on the surface of aerogel felt, which affects its performance and quality, especially in application scenarios such as high temperature insulation or aerospace, resulting in the inability to meet specific requirements.
The deep convolutional neural network model is used to combine HOG feature vectors and SVM classifiers, and through multi-scale feature extraction and fusion, combined with anchor-free bounding box prediction method, the defects on the surface of aerogel wool felt are accurately identified.
It improves the efficiency and accuracy of surface defect detection of aerogel felt, ensures product quality and production efficiency, adapts to complex production environments, and provides detailed defect location and category information.
Smart Images

Figure CN120259637A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of aerogel preparation, and particularly to a method and device for identifying surface defects of aerogel felts. Background Art
[0002] As a high-performance nanomaterial, aerogel is widely used in fields such as aerospace, building insulation, and electronic devices due to its unique structural properties, such as low density, high porosity, and low thermal conductivity. However, defects on the surface of aerogel felts may have a significant impact on their performance, such as a decrease in heat insulation performance and a weakening of durability. In application scenarios with extremely high requirements for material performance, such as high-temperature insulation or aerospace, the surface quality of aerogel felts is crucial, and any defect may cause them to fail to meet specific requirements. To fully utilize the performance advantages of aerogel felts, surface defects must be strictly controlled during the production process, and effective measures must be taken to solve problems such as powder shedding. Therefore, the detection of surface defects of aerogel felts is of great significance. Through visual detection methods, real-time defect detection of aerogel felts during the conveying process can be carried out, thereby helping to improve product quality and production efficiency. Summary of the Invention
[0003] In view of this, this application provides a method and device for identifying surface defects of aerogel felts, which can improve the detection efficiency and accuracy of surface defects of aerogel felts.
[0004] Specifically, this application is implemented through the following technical solutions:
[0005] The first aspect of this application provides a method for identifying surface defects of aerogel felts, and the method includes:
[0006] Obtain the image data after preprocessing the surface of the aerogel felt, and extract the HOG feature vector of the image data;
[0007] Label the preprocessed image data, build a deep convolutional neural network model for identifying aerogel felt defects, and train the deep convolutional neural network model with the labeled data;
[0008] Among them, the deep convolutional neural network model includes a CSPDarknet53, a spatial pyramid pooling module, and a path aggregation network module connected in sequence;
[0009] The CSPDarknet53 takes image data as input and outputs multi-scale feature maps of the image data; the spatial pyramid pooling module performs multi-scale pooling operations on the feature maps output by the CSPDarknet53 and outputs the pooling results to the path aggregation network module; the path aggregation network module receives the multi-scale feature maps output by the spatial pyramid pooling module and the CSPDarknet53, performs feature fusion and enhancement processing, and outputs the fused feature maps; the detection head of the deep convolutional neural network model adopts an anchor-free bounding box prediction method, is connected to the path aggregation network module, takes the fused feature maps as input, and outputs the predicted defect positions and category information;
[0010] Input the HOG feature vector into a pre-trained SVM classifier, and use the SVM classifier to screen out the first image, where the first image is an image with defects on the surface of the aerogel felt;
[0011] Input the first image into the deep convolutional neural network model to output defect positions and category information.
[0012] A second aspect of this application provides a device for identifying defects on the surface of an aerogel felt. The device includes an extraction module, a training module, a screening module, and an identification module.
[0013] Among them, the extraction module is used to obtain the preprocessed image data on the surface of the aerogel felt and extract the HOG feature vector of the image data;
[0014] The training module is used to label the preprocessed image data, build a deep convolutional neural network model for identifying aerogel felt defects, and train the deep convolutional neural network model with the labeled data;
[0015] Among them, the deep convolutional neural network model includes a CSPDarknet53, a spatial pyramid pooling module, and a path aggregation network module connected in sequence;
[0016] The CSPDarknet53 takes image data as input and outputs multi-scale feature maps of the image data; the spatial pyramid pooling module performs multi-scale pooling operations on the feature maps output by the CSPDarknet53 and outputs the pooling results to the path aggregation network module; the path aggregation network module receives the multi-scale feature maps output by the spatial pyramid pooling module and the CSPDarknet53, performs feature fusion and enhancement processing, and outputs the fused feature maps; the detection head of the deep convolutional neural network model adopts an anchor-free bounding box prediction method, is connected to the path aggregation network module, takes the fused feature maps as input, and outputs the predicted defect positions and category information;
[0017] The screening module is configured to input the HOG feature vector into a pre-trained SVM classifier, and use the SVM classifier to screen out the first image, where the first image is an image with defects on the surface of the aerogel felt;
[0018] The recognition module is configured to input the first image into the deep convolutional neural network model and output defect location and category information.
[0019] The method and device for identifying surface defects of aerogel felt provided by this application improve the feature extraction link, and fuse HOG features and multi-scale features. Among them, HOG feature vectors are selected to accurately capture the local shape and texture features of defects such as pits, stains, floating glue, and indentations on the surface of aerogel felt, that is, use HOG features to capture changing, local features; use CSPDarknet53 in the deep convolutional neural network model to extract multi-scale features, which can capture features of different scales on the surface of aerogel felt, and can effectively extract both tiny texture changes and larger defect contours, that is, extract the information content of the image from different scales. The spatial pyramid pooling module further fuses multi-scale features to enhance the adaptability to defects of different sizes. Defects of different sizes present different feature scales in the image. The spatial pyramid pooling module divides the feature map into sub-regions of different scales and performs max pooling operations, which can extract the most significant feature information at different scales. After these feature information of different scales are fused, the model can simultaneously focus on the local details and macroscopic features in the image. The path aggregation network module builds a communication bridge between features at different levels through upsampling and downsampling operations, realizes the bidirectional fusion and enhancement of semantic information and detail features, and further improves the model's recognition ability for complex defects. Among them, semantic information mainly refers to the higher-level and more abstract features in the image, which contains the overall understanding and cognition of the image content. In the identification of surface defects of aerogel felt, semantic information can be understood as an abstract description of the defect category. For example, through the analysis of image features, it is judged that the defect is a pit, a stain or other types. Detail features refer to the specific and subtle information in the image, such as the texture and color changes at the edge of the defect. The method provided by the present invention obtains feature information of multiple scales and different granularities by fusing HOG local strong change features and CSPDarknet53 multi-scale features during the feature extraction process, which improves the accuracy of feature extraction.
[0020] Furthermore, the detection head provided by the present invention adopts an anchor-free bounding box prediction method, which abandons the cumbersome and disadvantages of traditional anchor settings, directly locates the defect location and category accurately from the feature map, and significantly improves the detection efficiency and accuracy.
[0021] Finally, on the basis of accurate feature extraction, the method provided by the present invention first performs rough classification on the feature information, that is, introducing an SVM classifier. In the face of a large number of aerogel felt images, it can quickly screen out the images that may have defects, greatly reducing the processing burden of the deep convolutional neural network model and improving the operating efficiency of the entire system. After rough classification, accurate detection is performed on the suspected defective images, improving the efficiency of defect recognition. Description of the Drawings
[0022] Figure 1 It is a flowchart of the first embodiment of the method for identifying surface defects of aerogel felt provided by this application;
[0023] Figure 2 It is a schematic structural diagram of the surface defect detection system of aerogel felt shown in this application;
[0024] Figure 3 It is a schematic diagram of the surface image defect categories of aerogel felt shown in this application;
[0025] Figure 4 It is a data transmission architecture diagram of the surface defect detection system of aerogel felt shown in this application;
[0026] Figure 5 It is a schematic structural diagram of the second embodiment of the surface defect device of aerogel felt provided by this application. Detailed Embodiments
[0027] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application.
[0028] The terms used in this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The singular forms "a", "the", and "said" used in this application are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0029] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" used herein may be interpreted as "when" or "while" or "in response to determining".
[0030] Specific embodiments are given below to introduce the technical solutions of the present application in detail.
[0031] Embodiment 1:
[0032] Figure 1 It is a flowchart of Embodiment 1 of the method for identifying surface defects of the aerogel felt provided by the present application. Please refer to Figure 1 , the method provided in this embodiment may include:
[0033] S101. Obtain the image data after preprocessing the surface of the aerogel felt, and extract the HOG feature vector of the image data.
[0034] It should be noted that the image data of the surface of the aerogel felt can be obtained by building a surface defect detection system for the aerogel felt. Specifically, before obtaining the image data after preprocessing the surface of the aerogel felt, it includes: calculating the arrangement positions of each camera among multiple cameras according to the conveyor belt position, and the three-dimensional space formed by the maximum shooting of each camera covers all angles of the aerogel product conveyed on the conveyor belt; setting the conveyor belt speed of the aerogel product according to the acquisition time of the multiple cameras and the image line interval error requirement; based on the multiple cameras, shooting images of each angle of the aerogel product; and adopting photometric stereo vision technology to fuse multiple images with multi-angle exposure into a complete defect image.
[0035] As an optional embodiment, fusing the multiple images collected into a complete defect image specifically includes: establishing a spatial coordinate system with the starting point of the conveyor belt as the origin; determining the first spatial coordinate of the aerogel product and the acquisition angle of each camera at the moment of image acquisition; determining the imaging area of each camera at the first spatial coordinate according to the light propagation path when each camera is rotated at the acquisition angle, and the imaging area is the area where the camera captures the aerogel product; determining the overlapping shooting area of the aerogel product according to the imaging areas of each camera; identifying the iconic image features in the overlapping shooting area; sorting the multiple images based on the iconic image features and the coordinate positions of each camera; and fusing the sorted images according to the boundary features of the overlapping shooting area. Specifically, the overlapping shooting area refers to the area on the aerogel product, which will be repeatedly photographed by multiple cameras. At this time, this area on the aerogel product is first located, and then the landmark features of this area are identified, and then the adjacent relationship between multiple images is quickly found, that is, images with the same landmark features are adjacent, and the order relationship between adjacent images is determined based on the position of the landmark features in the image. Furthermore, the order of the images taken by the camera is determined according to the location of the camera, and then the order between multiple images taken in a camera is determined according to the landmark features. Among the multiple images taken by each camera, the images at the head and tail of the order are determined according to the landmark features, and the internal order between multiple images taken by a camera is determined according to the rotation angle and shooting time of the camera. After the sorting is completed, the image fusion between the cameras quickly locates the fusion area of the two images through the border features of the overlapping shooting area, that is, the tail image of the previous group of images and the head image of the next group of images, and each group is multiple images taken by one camera.
[0036] According to the given image fusion method, the overlapping shooting area is accurately located and the iconic features are identified, which can ensure the complete splicing of the aerogel felt images taken from multiple angles. In actual production, the surface defects of aerogel felt are diverse, and a single image is difficult to show the whole picture. For example, a tiny crack may only be shown partially in one image, but the fused image can be fully presented, providing comprehensive information for subsequent defect detection and avoiding misjudgment or omission of defects due to missing information. The aerogel felt moves continuously on the conveyor belt. If the image is sorted incorrectly, it will cause the position of the defect to deviate, affecting the subsequent analysis. By accurately sorting the fused images, the surface condition of the aerogel felt can be truly reflected, providing a reliable basis for defect location and analysis, and improving the accuracy of defect identification. In addition, the production conditions of aerogel felt are complex and changeable, and factors such as lighting and camera position may differ. This fusion method can effectively integrate images taken under different conditions by establishing a coordinate system, determining the imaging area and the overlapping shooting area, and has strong adaptability. Whether the lighting is uneven or the camera position is slightly deviated, the image can be accurately fused to ensure the stability and reliability of defect detection.
[0037] Figure 2 The following is a schematic structural diagram of the aerogel felt surface defect detection system shown in this application. Please refer to Figure 2 , a high-resolution ordinary industrial lens is used above the conveyor belt, and a multi-angle light source combination (such as at 45°, 90°, and 135° obliquely) is configured to achieve multiple exposures, avoiding the loss of defect area or incomplete defect information caused by single-angle exposure. Among them, a line-scan camera (single-field of view 550 mm) is selected for the camera layout, and 3 cameras are arranged along the width of the felt to cover a full width of 1500 mm. For defects with directionality (such as scratches and pressing points), the multi-angle light source combination method can ensure that defects in different directions appear in at least one image. Therefore, photometric stereo vision technology is used to fuse multiple images with multi-angle exposures into a complete defect image for subsequent detection. After the camera is deployed on the aerogel product conveyor belt, the speed range of the aerogel product conveyor belt, the camera acquisition interval time, and the image line interval error requirements are set, etc., to complete the deployment of the aerogel felt defect detection system. Among them, the conveyor belt speed range should be 300 - 2000 mm / s; the maximum acquisition interval time of the camera is 50 μs; the image line interval error should be within 0.1 mm (no omission should be guaranteed at the highest operating speed of the felt), and the data collected for each frame of the image includes felt texture, boundary, lighting characteristics, and potential defects.
[0038] The aerogel felt defect detection system completed by deployment can obtain the surface image data of the aerogel felt. In order to further improve the image quality and highlight the image features for subsequent processing, the obtained image data is first preprocessed. Specifically, the preprocessing can include grayscale processing, Gaussian blur, edge enhancement, and adaptive histogram equalization; the process of the preprocessing includes: based on the collected surface image data of the aerogel felt, using the weighted average method to convert the red, green, and blue components into grayscale values to obtain a grayscale image; applying Gaussian blur technology to smooth the grayscale image to obtain a denoised image; using the Sobel operator to perform edge detection on the denoised image, obtaining a gradient amplitude image by calculating the brightness gradient of the denoised image, and using the gradient amplitude image as the original image channel for edge enhancement to obtain an edge-enhanced image; through the CLAHE algorithm, dividing the edge-enhanced image into multiple non-overlapping sub-blocks, and independently performing histogram equalization on each sub-block according to the pixel values of the original image, the total number of pixels of the non-overlapping sub-blocks, the number of gray levels, and the cumulative distribution function of the histogram.
[0039] Among them, the grayscale processing includes: converting the red, green, and blue components into grayscale values based on the acquired image data using the weighted average method. It should be noted that the purpose of grayscale processing is to simplify the image data, eliminate color information, and only retain the luminance information, so that the subsequent algorithms can focus more on the luminance changes and structural features of the image, thereby reducing the complexity of subsequent processing. Specifically, the pixel value range of grayscale is from 0 to 255. The grayscale processing is achieved by the weighted average method, and the specific formula is as follows:
[0040] ;
[0041] where R, G, and B respectively represent the red, green, and blue components of the pixel, and Igray represents the grayscale value. This weighting method simulates the human eye's sensitivity to different colors, being most sensitive to green and least sensitive to blue.
[0042] The Gaussian blur includes: applying Gaussian blur technology to smooth the grayscale image. Among them, each pixel value of the grayscale image is replaced by the weighted average of the pixels within a specified area, and the weight is determined by the Gaussian kernel. It should be noted that Gaussian blur is an effective noise suppression technology. By reducing image noise and details, it helps to reduce the risk of false detection caused by noise in the subsequent edge detection process. Through Gaussian blur, the edges of the image become smoother, and at the same time, the high-frequency noise components in the image are reduced, providing a clearer image basis for accurate edge detection. Specifically, in this processing process, the weight is determined by the Gaussian kernel, and the weight of the central pixel is the largest. As the distance from the central pixel increases, the weight gradually decreases. The new value of each pixel is the weighted average of the surrounding pixel values, thereby achieving the smoothing processing of the image. The specific formula is as follows:
[0043] ;
[0044] where is the pixel value of the original image at the position. is the Gaussian kernel. k is the radius of the kernel, representing the size of the kernel, which determines the intensity of the blur effect. A larger kernel will result in a stronger blur effect. It is usually an odd number, such as 3x3, 5x5, etc., so as to ensure there is a center point. represents the new value of the pixel at the position after Gaussian blur processing. This value is the weighted average of the pixel values in its neighborhood, and the weight is determined by the Gaussian kernel.
[0045] The Gaussian kernel construction formula is as follows:
[0046] ;
[0047] where x and y are coordinates, is the standard deviation, which determines the smoothness of the Gaussian kernel, The larger it is, the stronger the blurring effect. When specifically implemented, the OpenCV library can be used to automatically calculate the appropriate value.
[0048] The edge enhancement includes: performing edge detection on the denoised image after Gaussian blurring using the Sobel operator, highlighting the edge information by calculating the image brightness gradient, and using the gradient magnitude image as a channel of the original image for edge enhancement. It should be noted that the Sobel operator is a discrete differential operator used to calculate the approximate gradient of the image grayscale function to highlight the edge information in the image. Since Gaussian blurring has reduced the noise, the Sobel operator can extract the edge information of the image on a relatively clean basis, which is directly related to the accurate extraction of subsequent image features. Specifically, the Sobel operator is used to calculate the gradient of the image brightness to highlight the edge information, that is, to determine the boundary by comparing the color changes around the pixel points. The gradient calculation formula of the Sobel operator is as follows:
[0049] ;
[0050] ;
[0051] Among them, , are the filter kernels in the horizontal and vertical directions. is the pixel value of the image at the position. , refers to the gradient of the image in the horizontal and vertical directions. On this basis, the gradient magnitude is calculated, and the formula is as follows:
[0052] ;
[0053] Among them, G is the gradient magnitude representing the edge strength, and Gx and Gy are the gradients in the horizontal and vertical directions respectively. The gradient magnitude image is a grayscale image, where the grayscale value of each pixel represents the edge strength at that point. In this image, the higher the grayscale value, the stronger the edge, and the lower the grayscale value in the non-edge area. In order to integrate the edge information and the information of the denoised image obtained by Gaussian blurring, the gradient magnitude image is used as a channel (usually the brightness channel) of the original image for edge information enhancement, for example, directly superimposing the gradient magnitude on the original image to enhance the visual effect of the edge.
[0054] The adaptive histogram equalization includes: applying the CLAHE algorithm to the edge-enhanced image, dividing the image into multiple non-overlapping sub-blocks, and independently performing histogram equalization on each sub-block according to the pixel values of the original image, the total number of pixels in the sub-block, the number of gray levels, and the cumulative distribution function of the histogram. It should be noted that CLAHE is a local contrast enhancement technique, especially suitable for enhancing the local contrast of images, especially in the dark regions of images. CLAHE makes edges and textures clearer by enhancing details, and also helps to further reduce the impact of uneven illumination in the image. Specifically, CLAHE divides the image into multiple small non-overlapping sub-blocks, independently performs histogram equalization on each sub-block, and limits the degree of contrast enhancement to avoid noise amplification. The formula for histogram equalization is:
[0055] ;
[0056] where, is the enhanced pixel value; I(x, y) is the pixel value of the original image; is the total number of pixels in the sub-block; L is the number of gray levels (usually 256); CDF is the cumulative distribution function of the histogram.
[0057] Specifically, subtract the minimum value of the CDF of the whole image from the CDF function of the input pixel point, that is, the brightness difference between the current pixel and the darkest point of the whole image. After normalizing the result, map it to the gray level range of [0, L-1] to achieve the equalization adjustment of the current pixel point. In addition, CLAHE clips the histogram and limits the frequency of each gray level, thereby equalizing and enhancing the image contrast.
[0058] Combined with the above description, the obtaining of the image data after preprocessing the surface of the aerogel felt and performing feature extraction to obtain the HOG feature vector includes: dividing the preprocessed image into cells, calculating the gradient magnitude of each cell, and calculating the gradient direction based on the gradient ratio in the vertical and horizontal directions of the gradient magnitude; converting the gradient direction into multiple direction intervals, constructing a first feature vector for each direction interval; weighted accumulating the gradient magnitude of each cell to the first feature vector in the corresponding direction interval to obtain a second feature vector; normalizing multiple cells, and fusing each of the second feature vectors to form a third feature vector; connecting all the third feature vectors into a long vector as the final HOG feature vector.
[0059] In specific implementation, since defects (such as stains and indentations) often exhibit abnormal gradient distributions, which are significantly different from the characteristics of blank areas, the HOG algorithm is selected to extract the local texture features of the image, which can minimize the computational complexity. The HOG algorithm divides the image into small cells (such as 8×8 pixels), calculates the gradient magnitude in each cell, and calculates the gradient direction based on this, that is, the direction of the edge or texture at each pixel point:
[0060] ;
[0061] where Gx and Gy are the gradients in the horizontal and vertical directions respectively. Since the gradient direction is continuous (such as from 0° to 180°), direct processing will increase the computational complexity. Therefore, it is divided into several direction intervals (such as 9 direction intervals, each interval covering 20°). For each cell, a feature vector is constructed (such as dividing into 9 direction intervals, then the feature vector is 9-dimensional). The gradient magnitude (G) of the pixel points in the unit cell is weighted and accumulated into the corresponding direction interval to obtain the corresponding feature vector. Multiple cells are normalized and fused into a single feature vector (the number of dimensions remains unchanged), and then all feature vectors are concatenated into a long vector as the final HOG feature vector (the number of dimensions changes).
[0062] S102. Label the preprocessed image data, build a deep convolutional neural network model for aerogel felt defect recognition, and train the deep convolutional neural network model with the labeled data.
[0063] Among them, the deep convolutional neural network model includes a CSPDarknet53, a spatial pyramid pooling module, and a path aggregation network module connected in sequence; the CSPDarknet53 takes the image data as input and outputs a multi-scale feature map of the image data; the spatial pyramid pooling module performs multi-scale pooling operations on the feature map output by the CSPDarknet53 and outputs the pooling results to the path aggregation network module; the path aggregation network module receives the multi-scale feature maps output by the spatial pyramid pooling module and the CSPDarknet53, performs feature fusion and enhancement processing, and outputs the fused feature map; the detection head of the deep convolutional neural network model uses an anchor-free bounding box prediction method, is connected to the path aggregation network module, takes the fused feature map as input, and outputs the predicted defect location and category information.
[0064] It should be noted that the HOG feature vectors and the corresponding labels are used to form the training set and the test set. The labels are divided into positive samples and negative samples. The positive samples are marked as 1, indicating that there are defects in the surface image of the aerogel felt. For example, in the image, there are cracks, holes, breakages, pits, stains, floating glue, indentations, etc. on the aerogel felt. When training and testing the model, these positive samples serve as examples for the model to identify "existing defects", enabling the model to learn the characteristics of defective aerogel felt images. The negative samples are marked as 0, representing that there are no defects in the surface image of the aerogel felt, that is, the surface of the aerogel felt in the image is smooth and has no visible problems. The negative samples are used to inform the model that such images are normal and defect-free samples, enabling the model to distinguish between normal images and defective images.
[0065] In addition, since the result of SVM is to judge whether there are defects in the image, and the output result of the trained neural network model is the defect location and type, when constructing the training set and the test set, in addition to identifying whether there are defects in the samples, for the samples with defects (positive samples), specific defect information should also be identified, such as whether the defect is a crack, a pit or other types, so that the neural network model can learn the characteristic patterns corresponding to different defect types during the training process, and then in practical applications, it can not only accurately judge whether a new image has defects, but also accurately determine the location and category of the defects.
[0066] Figure 3 For the schematic diagram of the defect categories of the surface image of the aerogel felt shown in this application, please refer to Figure 3 , the defect categories shown in this application include pits, stains, floating glue, indentations, etc.
[0067] After annotating the preprocessed image data, it also includes operations such as scaling and normalizing the images to be trained to make them suitable for network input. Process the annotated data and convert it into a format that the model can understand. At the same time, apply data augmentation techniques such as rotation, flipping, and cropping to improve the generalization ability of the model.
[0068] Combined with the above description, it should be noted that the operation steps of CSPDarknet53 include:
[0069] (1) Use the preprocessed feature map of the surface image of the aerogel felt as the input to obtain multi-scale feature maps.
[0070] It should be noted that multi-scale feature maps can be obtained by using deep convolutional layers and lightweight convolutional layers. Among them, the deep convolutional layers mine the abstract features in the image through multiple convolutional operations, while the lightweight convolutional layers reduce the computational cost while ensuring the feature extraction ability.
[0071] (2) Incorporate the local texture and edge feature information contained in the HOG feature vector into the multi-scale feature map according to a preset fusion rule to form a first fusion feature map.
[0072] It should be noted that the fusion rule can be determined according to different feature importance and position information. For example, for the texture information at a specific position, add it to the feature map at the corresponding position and assign different weights according to the significance of the feature to enhance the representation ability of the feature map for local details.
[0073] (3) Perform a splitting operation on the first fusion feature map to obtain multiple sub-feature sets.
[0074] Specifically, the splitting can be performed according to factors such as the scale, semantic level, or spatial position of the feature.
[0075] As an optional embodiment, identify the object boundaries in the surface image of the aerogel felt, where the object boundaries are the boundaries of all objects in the image; determine the first splitting method of the feature map according to the object boundaries, and split the feature map according to the first splitting method to obtain a first feature set. The feature map within one boundary is split into a complete sub-block, and the number of complete sub-blocks is equal to the number of object boundaries; determine the minimum scale for splitting each first sub-feature according to the semantic meaning of each first sub-feature; take the mean of the minimum scales between the first sub-features with semantic similarity greater than the threshold as the target splitting scale for the first sub-features with semantic similarity greater than the threshold, and determine the splitting gap of each first sub-feature according to the target splitting scale, where the splitting gap is the gap from the current scale of the first sub-feature to the target splitting scale; split the first sub-features according to the semantic level and the splitting gap of the first sub-features, and each first sub-feature is split to obtain a set of second sub-features. The set composed of each group of second sub-features is used as the result of splitting the feature map to obtain multiple sub-feature sets.
[0076] (4) Recombine the split sub-feature sets to obtain the recombined multi-scale feature information and obtain a recombined multi-scale feature map.
[0077] Specifically, recombine the split sub-feature sets, and recombine the sub-feature sets through a parallel processing architecture to obtain the recombined multi-scale feature information. Among them, the parallel processing method can accelerate feature fusion and information integration, and improve the richness and diversity of feature representation.
[0078] The operation steps of the spatial pyramid pooling module in the identification of surface defects of aerogel felt include:
[0079] (1)Receive the output of CSPDarknet53 to reconstruct the multi-scale feature map, divide the reconstructed multi-scale feature map into sub-regions of multiple different scales, and perform max pooling operations on the sub-regions of each scale respectively.
[0080] The spatial pyramid pooling module is connected to CSPDarknet53 and can receive the feature map output from CSPDarknet53.
[0081] It should be noted that these sub-regions cover different sizes and resolutions, covering both local details and more macroscopic information. Max pooling operations are performed on the sub-regions of each scale respectively to extract the most significant feature information at different scales, while retaining the key information in the local features and avoiding information loss.
[0082] (2)Combine the results of pooling the sub-regions of different scales according to the preset splicing rules.
[0083] It should be noted that the splicing rules can determine the order and weights according to the importance of different scales and task requirements, forming a comprehensive feature vector. This feature vector integrates the feature information at multiple scales and provides a more comprehensive information basis for subsequent feature processing.
[0084] The operation steps of the path aggregation network module in the identification of surface defects of aerogel felt include:
[0085] (1)Receive the feature map output from the spatial pyramid pooling module and the reconstructed multi-scale feature map of CSPDarknet53.
[0086] (2)Enlarge the first range layer feature map through upsampling operation, and fuse the abstract semantic feature information of the first range layer with the second range layer feature map according to the predetermined fusion rules to obtain the first fusion result; the first range layer is at a deeper position in the convolutional neural network than the second range layer.
[0087] It should be noted that the first range layer feature map can be a high-level feature map, and the second range layer feature map can be a low-level feature map. Specifically, the high-level feature map is enlarged through upsampling operation to match its size with the low-level feature map, and the abstract semantic feature information of the high level is fused with the low-level feature map according to the predetermined fusion rules. The fusion process can adopt the method of weighted summation, and different weights are assigned according to the importance of the feature map to allow the low-level feature map to incorporate more semantic information.
[0088] (3)Shrink the second range layer feature map through downsampling operation, and add the detailed features of the second range layer to the first range layer feature map according to the predetermined fusion rules to obtain the second fusion result.
[0089] The low-level feature map is reduced through downsampling operations to match the size of the high-level feature map. The detailed features of the low level are added to the high-level feature map according to a predetermined fusion rule to enhance the high-level feature map's perception ability of local details. The fusion method can be element-wise addition or concatenation, and the weights are adjusted according to task requirements.
[0090] (4) Through multiple interaction and fusion operations on the first fusion result and the second fusion result, feature enhancement of the final fused feature map is achieved.
[0091] Through multiple top-down and bottom-up information interaction and fusion operations, effective fusion and supplementation of features are achieved. Convolution operations are used to enhance the features of the fused feature map, highlighting important features and suppressing irrelevant information, and finally the fused feature map is output.
[0092] It should also be noted that the detection head adopts an anchor-free bounding box prediction method and is connected to the path aggregation network module, taking the fused feature map output by the path aggregation network module as input. Through this method, the object detection task is directly completed based on the extracted features, avoiding the complexity of traditional anchor settings and improving detection efficiency and accuracy.
[0093] Specifically, the detection head is connected to the path aggregation network module. After the path aggregation network module performs feature fusion and enhancement processing, the output fused feature map is passed to the detection head. The detection head receives these fused feature maps and performs specific operations of object detection on them.
[0094] Among them, the detection head is composed of multiple convolutional layers. Specifically, the detection head divides the feature map into multiple "grid points" and uses convolutional layers to predict information such as "bounding box center point", "width", "height", "confidence", and "category" at each position.
[0095] In specific implementation, a convolutional layer is used to predict the offsets (Δx, Δy) of the center point coordinates relative to the current grid point at each grid point position. To ensure that the values of the offsets are between 0 and 1, the Sigmoid activation function is used to normalize them. Then, the normalized offsets are added to the grid coordinates to obtain the bounding box center point coordinates. The calculation formula is: x = grid_x + Δx, y = grid_y + Δy, where grid_x and grid_y are the coordinates of the grid point, and x and y are the bounding box center point coordinates. Since a bounding box center point coordinate can be calculated for each grid point, a convolutional layer is needed to predict the confidence for each grid point. Finally, through NMS (Non-Maximum Suppression) and "threshold screening", the coordinates with high confidence are selected as the final output.
[0096] Use another set of convolutional layers to predict the logarithmic values log(w) and log(h) of the bounding box width and height at each grid point. To obtain the actual width and height, apply the exponential function to these logarithmic values for inverse normalization, i.e., w = e log(w) , h = e log(h) . Similarly, relying on the confidence prediction, select the width and height with high confidence as the output.
[0097] Furthermore, the probability distribution of each bounding box belonging to each category is output through the convolutional layer. To convert the output into probability values, apply the Softmax activation function to the category prediction results. Finally, relying on the confidence prediction, select the category with high confidence as the output.
[0098] Finally, compare the data such as the center point, width, height, confidence, and category of the predicted bounding box with the real data, and calculate the loss function to optimize the training effect of the detection head. Specifically, the following 3 types of loss functions are used:
[0099] Bounding box regression loss: Calculate the difference between the predicted bounding box and the real bounding box, usually using IoU loss or GIoU loss.
[0100] Category classification loss: Calculate the cross-entropy loss between the predicted category probability and the real category label.
[0101] Confidence loss: Calculate the binary cross-entropy loss between the predicted confidence and the real target / background label.
[0102] The calculated values of the above 3 types of loss functions represent the quality of the prediction results. To minimize the loss function, that is, to make the prediction results reach the best, use the SGD (Stochastic Gradient Descent) optimizer for parameter optimization. SGD calculates the gradient by randomly selecting a small batch of data and then updates the model parameters, which has the characteristics of fast convergence speed and high computational efficiency. It should be noted that after the above model improvement based on the YOLOv9 architecture, it also includes the training and evaluation of the model. Specifically, the training and evaluation adjustment of the deep convolutional neural network model based on the identification of aerogel felt defects includes:
[0103] (1) Set the hyperparameters required for training, select a composite loss function that includes classification loss, localization loss, and confidence loss, and an optimizer for optimizing the neural network model parameters.
[0104] It should be noted that hyperparameters required for training are set, including learning rate, batch size, number of training epochs, etc. The learning rate can be dynamically adjusted according to the training stage and the convergence of the model. The batch size is reasonably selected according to the hardware resources and data volume, and the number of training epochs is determined according to the convergence of training and performance metrics. For example, the initial learning rate is set to 0.01; the learning rate scheduler uses CosineDecay to gradually decay the learning rate; the batch size is 32; the number of training epochs is set to 300; and the weight decay is 0.0005.
[0105] Select an appropriate loss function. For example, a composite loss function including classification loss, localization loss, and confidence loss can be used to comprehensively consider the accuracy of defect class prediction and location prediction. Specifically, the classification loss function selects Binary Cross-Entropy Loss, the localization loss is CIoU Loss, and the confidence loss is Focal Loss; the optimizer selects Stochastic Gradient Descent (SGD) or Adam to optimize the parameters of the neural network. By calculating gradients and updating parameters, the model gradually converges. The metrics select mAP, Recall, and Precision.
[0106] (2) Divide the labeled image data into a training set, a validation set, and a test set. Input the training set data into the constructed deep convolutional neural network model. During training, calculate the loss function and update the parameters of the deep convolutional neural network model according to the optimizer's strategy; use the validation set to evaluate the trained deep convolutional neural network model, and adjust the hyperparameters and model structure according to the evaluation results.
[0107] For example, the training set accounts for 80% of the total data, the validation set accounts for 10%, and the test set accounts for 10%. Use the validation set to evaluate the trained model, and adjust the hyperparameters and model structure according to the evaluation results. Evaluation metrics such as accuracy, recall, and F1 score can be used to evaluate the model's detection and classification capabilities for different types of defects. If the evaluation results do not meet the expectations, adjust the hyperparameters (such as adjusting the learning rate, changing the batch size), or fine-tune the model structure, such as adjusting the convolutional layer parameters of CSPDarknet53, the pooling scales of the Spatial Pyramid Pooling module, and the fusion rules of the Path Aggregation Network module. Repeat the model training and validation evaluation process until the model's performance on the validation set reaches satisfactory metrics, ensuring that the model has good generalization ability.
[0108] (3) When the deep convolutional neural network model reaches the specified performance on the validation set, use the test set to perform a final performance evaluation on the deep convolutional neural network model that has reached the specified performance.
[0109] After the model achieves satisfactory performance on the validation set, the trained model is evaluated for its final performance using the test set. The test set data is input into the model, and various performance metrics of the model on the test set are calculated, such as accuracy, recall, F1-score, and mean average precision (mAP), etc., to comprehensively evaluate the performance of the model on unseen data and verify whether the model truly has the ability to accurately detect and classify surface defects of aerogel felt.
[0110] S103. Input the HOG feature vector into a pre-trained SVM classifier, and use the SVM classifier to screen out the first image, where the first image is an image with defects on the surface of the aerogel felt.
[0111] It should be noted that the HOG feature vector can well describe the local shape and texture information of objects in an image and has strong expressive ability for features such as edges and contours on the surface of the aerogel felt. By calculating the gradient information of the image, HOG can highlight the detailed changes on the surface of the aerogel felt, and these changes are often closely related to the presence of defects. SVM is a classic machine learning classification algorithm and performs well in dealing with binary classification problems. It can find an optimal hyperplane to separate data points of different classes as accurately as possible. For the problem of identifying surface defects of aerogel felt, the images with defects and the defect-free images are regarded as two different classes. The SVM classifier can learn the boundary for distinguishing these two types of images based on the differences in HOG feature vectors, thereby accurately judging whether an image has defects. In addition, SVM has good generalization ability and can effectively learn the internal laws of the data when the training data is limited, and also has a good classification effect on unknown test data. This enables the SVM classifier trained based on the HOG feature vector to maintain a certain degree of accuracy and stability when facing aerogel felt images collected in different batches and different environments, and has strong adaptability.
[0112] In the entire process of identifying surface defects of aerogel felt, using the SVM classifier to screen based on the HOG feature vector can quickly perform preliminary processing on a large number of images. The images that are obviously defect-free are screened out, reducing the number of images that the subsequent deep convolutional neural network model needs to process, improving the operating efficiency of the entire defect recognition system, and the results of the SVM classifier can be used as auxiliary information to provide reference for the subsequent deep convolutional neural network model.
[0113] Specifically, the process of pre-training the SVM classifier includes:
[0114] (1) Combine the HOG feature vector with the corresponding label to construct a training set and a test set.
[0115] It should be noted that the HOG feature vectors are extracted from the surface images of aerogel felts, which contain the local shape and texture information of these images. To enable the SVM classifier to learn how to distinguish defective and non-defective images, we need to assign labels to these feature vectors. Positive samples (defective images) are labeled as 1, and negative samples (non-defective images) are labeled as 0. These labeled HOG feature vectors are divided into a training set and a test set, usually in a certain proportion, such as 70% of the data as the training set and 30% of the data as the test set. The training set is used to train the SVM classifier to learn how to judge whether an image is defective based on the HOG features; the test set is used to evaluate the performance of the trained classifier on unseen data.
[0116] (2) Select a kernel function and a regularization parameter to train the SVM classifier.
[0117] It should be noted that the role of the kernel function is to map the data in the low-dimensional space to the high-dimensional space, so that the data that is linearly inseparable in the low-dimensional space becomes linearly separable in the high-dimensional space. Common kernel functions include linear kernel, polynomial kernel, radial basis function (RBF) kernel, etc. Selecting different kernel functions will affect the performance of the SVM classifier. The regularization parameter C is used to balance the complexity of the model and the degree of fitting to the training data. During the training process, an appropriate C value needs to be selected according to the characteristics of the data and the experimental results to avoid overfitting or underfitting.
[0118] (3) Calculate the classification accuracy rates of the training set and the test set, and adjust the parameters of the HOG features and the SVM parameters based on the training results.
[0119] During the training process, the data in the training set is input into the SVM classifier for training, and then the data in the test set is used for testing. The classification accuracy rate (which is the ratio of the number of correctly classified samples to the total number of samples) is calculated. By calculating the classification accuracy rate, the performance of the model on the training set and the test set can be evaluated, and it can be understood whether the model is overfitting or underfitting to the data.
[0120] The parameters of the HOG features include window size, block size, cell size, the number of intervals of the gradient direction, etc. According to the results of the classification accuracy rate, if the model performs poorly, these parameters can be tried to be adjusted. For example, changing the window size may affect the extraction range of the features, thus changing the extracted feature information and further affecting the performance of the classifier. In addition, in addition to the regularization parameter C, the parameters of the kernel function can also be adjusted to optimize the performance of the SVM classifier.
[0121] (4)Use the SVM classifier for classification prediction. If it is determined that the current image has no defects, repeat the operation of using the HOG algorithm to extract features and inputting them into the SVM classifier; if it is determined that there are defects currently, input the image into the trained deep convolutional neural network model.
[0122] Use the trained SVM classifier to perform classification prediction on new images. For each input image, first use the HOG algorithm to extract its feature vector, and then input the feature vector into the SVM classifier. If the SVM classifier determines that the image has no defects (the predicted label is 0), it is considered that the image does not require further in-depth analysis for the time being. To ensure the accuracy of the results, the operation of using the HOG algorithm to extract features and inputting them into the SVM classifier will be repeated, and multiple judgments are made to ensure that it is not a misjudgment caused by accidental errors in feature extraction.
[0123] If the SVM classifier determines that the image has defects (the predicted label is 1), input the image into the trained deep convolutional neural network model, because the deep convolutional neural network model can perform more accurate localization and classification of the position and category of the defects, and can provide more detailed information, such as determining the position and category of the defects (such as pits, stains, floating glue, indentations, etc.), rather than just determining whether there are defects.
[0124] S104. Input the first image into the deep convolutional neural network model and output the defect position and category information.
[0125] It should be noted that the operation of inputting the screened images with defects into the deep convolutional neural network model and outputting the defect position and category information includes:
[0126] (1) Input the image data of the aerogel felt with defects screened by the SVM classifier into the trained deep convolutional neural network model optimized for aerogel felt defects.
[0127] After screening by the SVM classifier, we have selected the image data of the aerogel felt that may have defects. These image data contain various information on the surface of the aerogel felt, but the specific position and category information of the defects are not yet clear. Inputting them into the deep convolutional neural network model is to further utilize the powerful feature extraction and classification capabilities of the deep convolutional neural network to accurately find the position and category of the defects.
[0128] (2) Based on the CSPDarknet53, spatial pyramid pooling module, and path aggregation network module included in the deep convolutional neural network model, perform multi-scale feature extraction, multi-scale feature fusion, and fusion and enhancement operations between different-level features on the input image in sequence.
[0129] Specifically, for the introduction of each module, please refer to the above description and will not be elaborated here.
[0130] (3) Using the anchor-free bounding box prediction method adopted by the detection head, based on the analyzed feature information, locate the specific position of the defect on the surface of the aerogel felt in the image, and present the area range where the defect is located in the form of a bounding box.
[0131] It should be noted that the anchor-free bounding box prediction method is a method used by the detection head of the deep convolutional neural network model. It directly predicts the position of the defect bounding box based on the comprehensive feature map output by the path aggregation network module. Traditional anchor-based methods require pre-setting anchors of different sizes and shapes, while the anchor-free bounding box prediction method avoids this complex process and directly learns the position information of the defect from the feature map. For example, it can determine the position of the defect on the surface of the aerogel felt according to the feature intensity and distribution in certain regions of the feature map, and mark the area range where the defect is located with a bounding box (usually a rectangular area represented by the upper left corner coordinates and the lower right corner coordinates), visually showing the position of the defect in the image.
[0132] (4) Based on the different defect category feature patterns learned by the deep convolutional neural network model, determine the category to which the defect in the input image belongs.
[0133] During the training process, the deep convolutional neural network model has learned the feature patterns of different defect categories (such as pits, stains, floating glue, indentations, etc.). When an image with defects is input, according to the finally obtained feature map, the model can compare the defects in the image with the learned different category feature patterns to determine which category the defect belongs to. This is completed by the classifier part in the model. The classifier can be a fully connected layer or other forms. It will output the probabilities of the defect belonging to each category according to the information of the feature map, and finally determine the most likely category to which the defect belongs.
[0134] (5) Integrate and output the determined defect position information and the determined defect category information.
[0135] Integrate the defect position information (bounding box coordinates) obtained by the anchor-free bounding box prediction method and the defect category information determined by the classifier. In this way, the information of an aerogel felt image with defects can be presented in a clear and definite manner. For example, the output result can be: there is a pit in the area from (x1, y1) to (x2, y2) in the image, and there is a stain in the area from (x3, y3) to (x4, y4), etc.
[0136] Figure 4 For the data transmission architecture diagram of the aerogel felt surface defect detection system shown in this application, please refer toFigure 4 The images captured by the camera will be transmitted to the edge computing device, the industrial intelligent machine SX20. This device integrates a PLC control terminal and a Linux system platform. The conveyor belt and the rewinder are controlled by the PLC layer of this device. The image processing, HOG algorithm, classifier, and defect annotation model are all deployed on the Linux system platform of this device. The line scan camera transmits image information to the Linux layer based on the RTSP protocol. The control parameters of the conveyor belt and the rewinder are transmitted from the PLC layer to the Linux layer in the form of the TCP protocol.
[0137] The image processing algorithm performs preliminary processing on the data transmitted using the RTSP protocol and sends it to the HOG algorithm and classifier for screening. After detecting defects, it is transmitted to the annotation model for defect annotation and recognition. The model results combine the running speeds of the conveyor belt and the rewinder and the image acquisition time to calculate the location of the defect in the current batch of products starting time.
[0138] It should be noted that for the automatically recognized and classified images with defects, the defect location and category information are output, including: inputting the data of the suspected defect images screened out into the deep convolutional neural network model; calculating the location of the defect according to the on-site conveyor belt speed and the camera image acquisition time when the surface defect detection system is deployed; based on the defect category already annotated by the defect annotation model and the calculated location of the defect, outputting the location and category of the defect, and storing the defect location information and the defect category together. Specifically, the calculation method is as follows:
[0139] Defect location (L) = Conveyor belt running speed (V) × (Defect detection time (T) - Model calculation time (t_model));
[0140] For example, conveyor belt running speed (V): 2 m / s, model calculation time (t_model): 0.2 seconds, start time of the current batch of products (T_start): 10:00:00, defect detection time (T_detect): 10:00:05, then the defect detection time difference (T): T = T_detect - T_start = 5 seconds; corrected time difference (T_corrected): T_corrected = T - t_model = 5 - 0.2 = 4.8 seconds; defect location (L): L = V × T_corrected = 2 × 4.8 = 9.6 meters, that is, the location of the felt defect is at 9.6 meters. At this time, further, the defect location information and the defect category are jointly stored in the Mysql database.
[0141] In the method provided in this embodiment, in the image data acquisition and preprocessing stage, the well-deployed surface defect detection system of aerogel felt gives full play to the advantages of multi-camera multi-angle exposure and photometric stereo vision technology, successfully obtains comprehensive and accurate image data, and effectively avoids the problem of missing defect information caused by single-angle exposure. A series of preprocessing operations such as grayscale processing, Gaussian blur, edge enhancement, and adaptive histogram equalization significantly improve the image quality and highlight the defect features. The deep convolutional neural network model evolves continuously during the strict training and evaluation adjustment process. The reasonably set hyperparameters and the carefully selected composite loss function including classification loss, localization loss, and confidence loss guide the model to steadily move in the direction of accurately identifying defects. During the training process, the hyperparameters and model structure are flexibly adjusted according to the evaluation results of the validation set to ensure that the model has excellent generalization ability and can stably play a role in the complex and changeable actual production environment and accurately identify various defects. In addition, the HOG feature vector provides highly discriminative information for the SVM classifier with its strong expression ability for local texture and edge features of the image. The SVM classifier quickly screens out the defective images with its excellent performance in binary classification problems and good generalization ability, greatly shortening the detection time and improving the operation efficiency of the entire system. After the screened defective images are input into the deep convolutional neural network model, with the close cooperation of its internal modules and advanced detection head technology, the defect position and category information can be accurately output, providing detailed and accurate basis for the quality control of aerogel felt.
[0142] Embodiment 2:
[0143] Corresponding to the foregoing embodiment of the method for surface defects of aerogel felt, the present application also provides an embodiment of a device for surface defects of aerogel felt.
[0144] Figure 5 It is a schematic structural diagram of Embodiment 2 of the device for surface defects of aerogel felt provided by the present application. Please refer to Figure 5 The device provided in this embodiment includes an extraction module 510, a training module 520, a screening module 530, and an identification module 540.
[0145] Among them, the extraction module 510 is used to obtain the preprocessed image data on the surface of the aerogel felt and extract the HOG feature vector of the image data.
[0146] The training module 520 is used to label the preprocessed image data, build a deep convolutional neural network model for identifying aerogel felt defects, and train the deep convolutional neural network model with the labeled data.
[0147] Among them, the deep convolutional neural network model includes a CSPDarknet53, a spatial pyramid pooling module, and a path aggregation network module connected in sequence;
[0148] The CSPDarknet53 takes image data as input and outputs a multi-scale feature map of the image data; the spatial pyramid pooling module performs multi-scale pooling operations on the feature map output by the CSPDarknet53 and outputs the pooling results to the path aggregation network module; the path aggregation network module receives the multi-scale feature maps output by the spatial pyramid pooling module and the CSPDarknet53, performs feature fusion and enhancement processing, and outputs the fused feature map; the detection head of the deep convolutional neural network model uses an anchor-free bounding box prediction method, is connected to the path aggregation network module, takes the fused feature map as input, and outputs the predicted defect location and category information;
[0149] The screening module 530 is configured to input the HOG feature vector into a pre-trained SVM classifier and use the SVM classifier to screen out the first image, where the first image is an image with defects on the surface of the aerogel felt;
[0150] The recognition module 540 is configured to input the first image into the deep convolutional neural network model and output defect location and category information.
[0151] The device of this embodiment can be used to execute Figure 1 the steps of the method embodiment shown. The specific implementation principle and process are similar and will not be elaborated here.
[0152] For the implementation processes of the functions and roles of each unit in the above device, refer to the implementation processes of the corresponding steps in the above method for details, which will not be elaborated here.
[0153] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this application. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0154] The above are only the preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this application shall be included within the scope of protection of this application.
Claims
1. A method for identifying surface defects of aerogel felt, characterized in that, The method includes: Obtain the image data after surface preprocessing of the aerogel felt, and extract the HOG feature vectors of the image data; Annotate the preprocessed image data, build a deep convolutional neural network model for defect recognition of the aerogel felt, and train the deep convolutional neural network model with the annotated data; Among them, the deep convolutional neural network model includes a CSPDarknet53, a spatial pyramid pooling module, and a path aggregation network module connected in sequence; The CSPDarknet53 takes the image data as input and outputs multi-scale feature maps of the image data; the spatial pyramid pooling module performs multi-scale pooling operations on the feature maps output by the CSPDarknet53 and outputs the pooling results to the path aggregation network module; the path aggregation network module receives the multi-scale feature maps output by the spatial pyramid pooling module and the CSPDarknet53, performs feature fusion and enhancement processing, and outputs the fused feature maps; the detection head of the deep convolutional neural network model adopts an anchor-free bounding box prediction method, is connected to the path aggregation network module, takes the fused feature maps as input, and outputs the predicted defect positions and category information; Input the HOG feature vectors into a pre-trained SVM classifier, and use the SVM classifier to screen out the first images, where the first images are the images with defects on the surface of the aerogel felt; Input the first images into the deep convolutional neural network model, and output the defect positions and category information.
2. The method according to claim 1, characterized in that, Based on the training and evaluation adjustment of the deep convolutional neural network model for defect recognition of the aerogel felt, it includes: Set the hyperparameters required for training, select a composite loss function including classification loss, localization loss, and confidence loss, and an optimizer for optimizing the neural network model parameters; Divide the annotated image data into a training set, a validation set, and a test set. Input the training set data into the built deep convolutional neural network model. During the training process, calculate the loss function and update the parameters of the deep convolutional neural network model according to the optimizer's strategy; use the validation set to evaluate the trained deep convolutional neural network model, and adjust the hyperparameters and model structure according to the evaluation results; When the deep convolutional neural network model reaches the specified performance on the validation set, use the test set to perform the final performance evaluation on the deep convolutional neural network model that has reached the specified performance.
3. The method according to claim 1, characterized in that The operation steps of the CSPDarknet53 include: Take the preprocessed surface image feature map of the aerogel felt as input and obtain multi-scale feature maps; Integrate the local texture and edge feature information contained in the HOG feature vectors into the multi-scale feature maps according to the preset fusion rules to form the first fused feature maps; Perform a splitting operation on the first fused feature maps to obtain multiple sub-feature sets; Recombine the split sub-feature sets to obtain the recombined multi-scale feature information and obtain the recombined multi-scale feature maps.
4. The method according to claim 1, wherein The operation steps of the spatial pyramid pooling module in the defect recognition of the aerogel felt surface include: Receive the output of CSPDarknet53 to reorganize the multi-scale feature map, divide the reorganized multi-scale feature map into sub-regions of multiple different scales, and perform max pooling operations on the sub-regions of each scale respectively; Combine the results after pooling the sub-regions of different scales according to the preset splicing rules.
5. The method according to claim 1, wherein The operation steps of the path aggregation network module in the identification of surface defects of aerogel felt include: Receive the feature map output from the spatial pyramid pooling module and the reorganized multi-scale feature map of CSPDarknet53; Enlarge the feature map of the first range layer through upsampling operation, and fuse the abstract semantic feature information of the first range layer with the feature map of the second range layer according to the preset fusion rules to obtain the first fusion result; the first range layer is in a deeper position in the convolutional neural network than the second range layer; Reduce the feature map of the second range layer through downsampling operation, and add the detailed features of the second range layer to the feature map of the first range layer according to the preset fusion rules to obtain the second fusion result; Realize the feature enhancement of the finally fused feature map according to the multiple interaction and fusion operations of the first fusion result and the second fusion result.
6. The method according to claim 1, wherein The process of pre-training the SVM classifier includes: Combine the HOG feature vectors with the corresponding labels to construct a training set and a test set; Select a kernel function and a regularization parameter to train the SVM classifier; Calculate the classification accuracy rates of the training set and the test set, and adjust the parameters of the HOG features and the SVM parameters based on the training results; Use the SVM classifier for classification prediction. If it is determined that the current image has no defects, repeat the operation of using the HOG algorithm to extract features and input them into the SVM classifier; if it is determined that there are current defects, input the image into the trained deep convolutional neural network model.
7. The method according to claim 1, wherein The preprocessing includes grayscale processing, Gaussian blur, edge enhancement, and adaptive histogram equalization; The process of the preprocessing includes: Based on the collected image data of the surface of the aerogel felt, use the weighted average method to convert the red, green, and blue components into grayscale values to obtain a grayscale image; Apply Gaussian blur technology to smooth the grayscale image to obtain a denoised image; Use the Sobel operator to perform edge detection on the denoised image, obtain the gradient amplitude image by calculating the brightness gradient of the denoised image, and use the gradient amplitude image as the original image channel for edge enhancement to obtain an edge-enhanced image; Through the CLAHE algorithm, divide the edge-enhanced image into multiple non-overlapping sub-blocks, and perform histogram equalization on each sub-block independently according to the pixel values of the original image, the total number of pixels of the non-overlapping sub-blocks, the number of gray levels, and the cumulative distribution function of the histogram.
8. The method according to claim 1, wherein Before obtaining the preprocessed image data of the surface of the aerogel felt, it includes: Calculate the arrangement positions of each camera among multiple cameras according to the conveyor belt position, and the stereo space covered by the maximum shooting of each camera covers all angles of the aerogel product conveyed on the conveyor belt; Set the conveyor belt speed of the aerogel product according to the acquisition time of the multiple cameras and the requirement of the image line interval error; Based on the images of various angles of the aerogel product captured by the multiple cameras; Using photometric stereo vision technology, multiple images with multi-angle exposures are fused into a complete defect image.
9. The method according to claim 1, characterized in that, The obtaining of the image data after preprocessing the surface of the aerogel felt and the extraction of HOG feature vectors includes: Dividing the preprocessed image into cells, calculating the gradient magnitude of each cell, and calculating the gradient direction based on the gradient ratio in the vertical and horizontal directions in the gradient magnitude; Converting the gradient direction into multiple direction intervals, and constructing a first feature vector for each direction interval; Weighted accumulation of the gradient magnitude of each cell to the first feature vector in the corresponding direction interval to obtain a second feature vector; Performing normalization processing on multiple cells, and fusing each of the second feature vectors to form a third feature vector; Connecting all the third feature vectors into a long vector as the final HOG feature vector.
10. An apparatus for identifying surface defects of aerogel felt, characterized in that, The device includes an extraction module, a training module, a screening module, and an identification module, wherein the extraction module is used to obtain the image data after preprocessing the surface of the aerogel felt and extract the HOG feature vectors of the image data; The training module is used to label the preprocessed image data, build a deep convolutional neural network model for defect recognition of the aerogel felt, and train the deep convolutional neural network model with the labeled data; wherein the deep convolutional neural network model includes a CSPDarknet53, a spatial pyramid pooling module, and a path aggregation network module connected in sequence; The CSPDarknet53 takes the image data as input and outputs a multi-scale feature map of the image data; the spatial pyramid pooling module performs multi-scale pooling operations on the feature map output by the CSPDarknet53 and outputs the pooling result to the path aggregation network module; the path aggregation network module receives the multi-scale feature maps output by the spatial pyramid pooling module and the CSPDarknet53, performs feature fusion and enhancement processing, and outputs the fused feature map; the detection head of the deep convolutional neural network model adopts an anchor-free bounding box prediction method, is connected to the path aggregation network module, takes the fused feature map as input, and outputs the predicted defect position and category information; The screening module is used to input the HOG feature vectors into a pre-trained SVM classifier, and use the SVM classifier to screen out the first image, and the first image is an image with defects on the surface of the aerogel felt; The identification module is used to input the first image into the deep convolutional neural network model and output the defect position and category information.
Citation Information
Patent Citations
Visual-based appearance defect detection method for an earphone silica gel gasket
CN113989196A
YOLOv4-based battery piece defect detection method and system
CN114943830A
PCB defect detection system and method based on improved YOLO algorithm
CN119540152A
Method for fine defects Inspection of Leather using Deep Learning Model
KR102336110B1