Industrial product surface defect detection method based on AI large model

Through the industrial product surface defect detection method based on AI large model, the existing technology has solved the problem of insufficient robustness and insufficient extraction of diversified defect features under complex lighting conditions, and efficient real-time detection is achieved, which is suitable for complex industrial environments.

CN120013927AActive Publication Date: 2025-05-16BEIJING MIAOXIANG SCIENCE & TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510465030.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-16
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

The existing industrial product surface defect detection technology is insufficiently robust under complex lighting conditions, insufficient extraction of diversified defect features, limited generalization capabilities of model, and low real-time detection efficiency of edge equipment.

Method used

Using the surface defect detection method of industrial products based on AI large models, light invariant features and multi-scale features are generated through data preprocessing, deep feature extraction is used using the Transformer architecture, and multi-level defect feature maps are generated through multi-scale feature fusion. Through knowledge distillation and model quantization techniques, the model is optimized, and the lightweight model is generated and deployed to the edge device, and the detection threshold is dynamically adjusted to optimize the confidence judgment during the detection process.

Benefits of technology

It improves detection robustness under complex lighting conditions, enhances the ability to extract diverse defect features, improves the generalization ability of the model, and realizes real-time detection efficiency on edge devices with limited resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013927A_ABST
    Figure CN120013927A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of industrial detection, and discloses an industrial product surface defect detection method based on an AI large model, and the method comprises the following steps: S1, carrying out the data preprocessing of an industrial product surface image, and generating an illumination invariant feature and a multi-scale feature; s2, performing depth feature extraction on the preprocessed image by using an AI large model, and generating a multi-level defect feature map through multi-scale feature fusion; s3, optimizing the AI large model through a knowledge distillation technology to generate a lightweight model; s4, deploying the lightweight model to edge equipment, and performing surface defect detection; and S5, optimizing confidence judgment in the detection process by dynamically adjusting the detection threshold. Through the technical scheme of combining illumination invariant feature extraction and multi-scale feature decomposition, the effect of extracting surface detail features under a complex illumination condition is achieved, the problem of unstable detection caused by illumination change is solved, and the robustness of detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of industrial detection technology, and specifically to an industrial product surface defect detection method based on an AI big model. Background Art

[0002] Surface defect detection of industrial products is one of the important links in ensuring product quality and production efficiency in modern manufacturing. Existing technologies mainly include manual visual inspection and automatic inspection methods based on machine vision. Manual visual inspection relies on the experience and sensitivity of workers. Although it is still used in some small-scale production scenarios, it has low efficiency and consistency when facing large-scale assembly line operations. Automatic inspection technology based on machine vision uses regularized image processing algorithms to extract and analyze features of industrial surface images, thereby improving the degree of automation while reducing the shortcomings of manual intervention. These technologies show good applicability in scenarios with stable lighting conditions, relatively simple surface textures and defect features, and improve detection efficiency to a certain extent. However, with the increasing complexity and diversity of industrial production environments, existing technologies are gradually facing bottlenecks in adaptability.

[0003] There are still some shortcomings in the practical application of existing technologies. Traditional rule-based machine vision methods often have difficulty in processing complex textures and diverse defect features, especially in industrial environments with variable lighting conditions and diverse materials, and their accuracy and robustness are significantly limited. Although deep learning-based defect detection technology has made significant progress in automation and detection accuracy, the generalization ability of existing methods is relatively limited. The models are usually optimized for specific defect types or specific scenarios. When applied to scenarios with changes in materials, lighting or defect types, the detection performance is prone to decline. In addition, deep learning methods have a high demand for large-scale annotated data, while industrial site defect data is often unevenly distributed and has high collection and annotation costs, which makes model optimization more difficult. At the same time, deep learning models usually require high computing resource support, and their large-scale deployment in resource-limited edge devices is limited, making it difficult to meet the real-time detection needs of industrial sites. The above problems restrict the widespread application of existing technologies in complex industrial environments. Summary of the invention

[0004] In response to the shortcomings of the existing technology, the present invention provides an industrial product surface defect detection method based on an AI large model, which solves the problems of insufficient robustness of existing industrial product surface defect detection technology under complex lighting conditions, insufficient extraction of diversified defect features, limited model generalization capability, and low real-time detection efficiency of edge devices.

[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions: an industrial product surface defect detection method based on an AI large model comprises the following steps: S1. Perform data preprocessing on the surface images of industrial products to generate illumination invariant features and multi-scale features; S2. Use the AI ​​big model to extract deep features from the preprocessed image, and generate multi-level defect feature maps through multi-scale feature fusion; S3. Optimize the AI ​​large model through knowledge distillation technology to generate a lightweight model; S4, deploy the lightweight model to the edge device for surface defect detection; S5. Optimize the confidence determination during the detection process by dynamically adjusting the detection threshold.

[0006] Preferably, in step S1, the illumination invariant feature extraction of data preprocessing includes: The pixel gradient of the input image is calculated to obtain the gradient direction and amplitude of the pixel value, and the pixel gradient is normalized to eliminate the influence of light intensity and generate an illumination invariant feature map.

[0007] Preferably, in step S1, the multi-scale feature decomposition of data preprocessing includes: Through multi-scale decomposition, high-frequency features and low-frequency features are extracted from the input image, and the high-frequency features and low-frequency features are weighted fused according to the set weight parameters to generate a fused feature map containing multi-scale information.

[0008] Preferably, in the step S2, deep feature extraction is performed by a multi-head self-attention mechanism based on the Transformer architecture, which specifically includes the following steps: a. Generate query matrix, key matrix and value matrix; b. Generate attention weights through calculation between query matrix and key matrix; c. Use the attention weights to perform weighted summation on the value matrix to generate deep features; d. Extract global and local features through multi-head attention mechanism.

[0009] Preferably, in the step S2, multi-scale feature fusion is achieved through a feature pyramid network, and feature maps of different scales are synthesized into a unified multi-scale feature after weight optimization.

[0010] Preferably, in the step S3, the knowledge distillation technology includes the following steps: a. Use the teacher model to generate deep features for input data; b. Use the student model to learn the characteristic output of the teacher model; c. The student model is optimized by minimizing the cross entropy loss between the true labels and the student model output, and the difference loss between the teacher model and the student model output.

[0011] Preferably, in the step S3, the lightweight model is optimized by model quantization, and the model quantization discretizes the weights of the model, converts the weights from floating-point representation to quantized representation, and uses the quantized weights to perform matrix operations.

[0012] Preferably, in the step S4, the lightweight model in the edge device is deployed through a deep learning inference acceleration tool, and the inference acceleration tool is used to optimize the model inference speed and computing efficiency.

[0013] Preferably, in step S5, dynamically adjusting the detection threshold comprises the following steps: a. Statistical defect detection results confidence distribution, calculate the mean and standard deviation of the confidence; b. According to the set adjustment parameters, the detection threshold is dynamically adjusted using the mean and standard deviation.

[0014] Preferably, in the step S2, the weight of the multi-scale feature fusion is set by optimizing the response degree of the high-frequency features and the low-frequency features, and is used for weighted synthesis of local detail defects and global features.

[0015] The present invention provides an industrial product surface defect detection method based on an AI large model. It has the following beneficial effects: 1. The present invention adopts a technical solution that combines illumination-invariant feature extraction and multi-scale feature decomposition to achieve the effect of extracting surface detail features under complex illumination conditions. Compared with the detection method in the prior art that relies on fixed illumination conditions or global features, it solves the problem of unstable detection caused by illumination changes and improves the robustness of detection.

[0016] 2. The present invention achieves the feature expression capability of taking into account both global information and local details by performing deep feature extraction based on a large Transformer model and combining it with a multi-scale feature fusion network. Compared with the single-scale or single-network architecture solutions in the prior art, it effectively solves the problem of insufficient feature extraction for diversified defects and adapts to the complex needs of industrial surface defect detection.

[0017] 3. This invention uses knowledge distillation and model quantization technology to optimize large models into lightweight models, achieving the effect of efficient operation on edge devices. Compared with the existing solutions that rely on high-performance hardware support, it solves the problems of limited resources and high computing costs in industrial sites and realizes the deployment and application of lightweight models.

[0018] 4. The present invention achieves the effect of real-time optimization of missed detection rate and false detection rate through the technical solution of dynamically adjusting the detection threshold. Compared with the solution of detecting defects by using a fixed threshold method in the prior art, it effectively solves the problem of unstable detection accuracy under various defect types and production scenarios, and improves the detection reliability and adaptability in industrial environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 The figure is a flow chart of the method steps of the present invention. DETAILED DESCRIPTION

[0020] The following will be combined with the drawings in the specification of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0021] Please refer to the attached Figure 1 , an embodiment of the present invention provides an industrial product surface defect detection method based on an AI large model, comprising the following steps: S1. Perform data preprocessing on the surface images of industrial products to generate illumination invariant features and multi-scale features; S2. Use the AI ​​big model to extract deep features from the preprocessed image, and generate multi-level defect feature maps through multi-scale feature fusion; S3. Optimize the AI ​​large model through knowledge distillation technology to generate a lightweight model; S4, deploy the lightweight model to the edge device for surface defect detection; S5. Optimize the confidence determination during the detection process by dynamically adjusting the detection threshold.

[0022] Specifically, by combining AI big models, data preprocessing, model optimization and dynamic adjustment mechanisms, efficient detection of surface defects of industrial products can be achieved. First, the system performs data preprocessing on the input surface image of industrial products to generate illumination invariant features and multi-scale features. This step eliminates the influence of illumination intensity and separates local details from global background features, providing more robust input features for subsequent models. Then, after deep feature extraction of the AI ​​large model, the system uses a multi-head self-attention mechanism to comprehensively model the global and local information of the image. At the same time, through a multi-scale feature fusion network, it further integrates information at different levels and scales to generate a multi-level feature map suitable for defect detection. After that, in order to improve the operating efficiency in industrial scenarios, the system uses knowledge distillation technology to transfer the key knowledge of the large model to a lightweight model, and combines model quantization technology to further reduce storage and computing requirements, so that the model can adapt to edge devices. Finally, in the actual detection process, the system adopts a dynamic adjustment mechanism to adaptively optimize the detection threshold according to the real-time confidence distribution to balance the missed detection rate and the false detection rate, ensuring stable detection capabilities in complex industrial environments. Each step is closely linked, from feature input to optimized deployment to real-time detection, forming a complete and efficient defect detection process.

[0023] In step S1, the illumination invariant feature extraction of data preprocessing includes: Calculate the pixel gradient of the input image, obtain the gradient direction and amplitude of the pixel value, and normalize the pixel gradient to eliminate the influence of light intensity and generate an illumination invariant feature map; In step S1, the multi-scale feature decomposition of data preprocessing includes: Through multi-scale decomposition, high-frequency features and low-frequency features are extracted from the input image, and the high-frequency features and low-frequency features are weighted fused according to the set weight parameters to generate a fused feature map containing multi-scale information.

[0024] Specifically, step S1 mainly extracts illumination-invariant features and multi-scale features by preprocessing the surface images of industrial products, providing more robust input for subsequent large-model deep feature extraction. In an industrial environment, the surface of a product may have large differences in brightness or shadow due to different lighting conditions, and traditional image processing methods are difficult to eliminate the impact of illumination changes. In addition, the types, shapes, and sizes of surface defects vary greatly, and a single-scale feature extraction method cannot fully express these diversities. Through illumination-invariant feature extraction and multi-scale feature decomposition, the present invention can effectively reduce the interference of illumination changes, enhance the expressiveness of features, and provide high-quality input features for deep learning models.

[0025] Generally, illumination invariant feature extraction processes the gradient changes of pixels in an image to eliminate the influence of illumination intensity. As an alternative, multi-scale feature decomposition captures local details and global background information by separating high-frequency features from low-frequency features of an image. In some embodiments, the present invention can also optimize the weights of high-frequency and low-frequency features according to defect types and image characteristics, thereby improving the ability to express different types of defects. In this embodiment, for the input industrial product surface image, illumination invariant feature extraction is first performed. Specifically: In one possible implementation, the input image For each pixel point, calculate its gradient value to reflect the intensity change in the neighborhood of the pixel point. The calculation formula of the gradient value is: ; in: The coordinates in the image are The pixel value of Represents the intensity change of the pixel in the horizontal direction, that is, along Partial derivatives of direction; Represents the intensity change of the pixel in the vertical direction, that is, along Partial derivatives in direction.

[0026] As an option, in order to eliminate the influence of light intensity changes on the gradient, the calculated gradient is normalized. The specific formula for normalization is: ; in: Represents the normalized gradient direction; Represents the modulus of the gradient, and its calculation formula is: ; The purpose of normalization is to make the illumination invariant feature mainly depend on the direction of the gradient, and not be affected by changes in illumination intensity. In some embodiments, the illumination invariant feature map can be further transformed by function mapping. For example, the gradient direction can be mapped to pixel values ​​using an angle function to generate an illumination invariant feature map: ; Through the above processing, the illumination invariant feature map is obtained It can effectively capture surface detail features and eliminate the interference of illumination on the detection results. In another implementation, the original gradient information and the normalized gradient information can be directly combined to further improve the performance of the feature map.

[0027] In this embodiment, in order to solve the problem of diversity of defect types, a multi-scale feature decomposition technology is used to process the input image. Specifically: In one possible implementation, the illumination invariant feature map is decomposed at multiple scales by discrete wavelet transform. The input image is decomposed into two parts: high-frequency features and low-frequency features. High-frequency features are used to describe surface details, such as cracks and scratches, while low-frequency features retain the global background information of the image, such as large-scale undulations or depressions on the surface.

[0028] In general, the calculation formula of discrete wavelet transform is: ; in: Indicates that the image is at scale ,Location The wavelet component below; Represents the illumination invariant feature map; represents the wavelet basis function, is its conjugate complex number; represents the scale parameter, controlling the frequency range of the decomposition; Represents the translation parameter, which controls the position of the decomposition.

[0029] In one possible implementation, the high frequency features and low frequency characteristics They are calculated by different wavelet components. Specifically, the high-frequency part retains the detail information, and the low-frequency part retains the smooth background.

[0030] In some other embodiments, the high-frequency features and the low-frequency features can be further processed by weighted fusion. The fusion formula is: ; in: is the fused multi-scale feature map; is the weight parameter of high-frequency features, and its value range is 0≤ ≤1, the specific value can be set according to actual needs.

[0031] As an option, in some possible implementations, the weight parameter It can be dynamically adjusted according to the response degree of different types of defects. For example, for small scratch detection, the weight of high-frequency features can be appropriately increased; while for large-area dents or deformation detection, the weight of low-frequency features can be increased.

[0032] Through the above multi-scale decomposition process, the multi-scale feature map generated It can effectively express the global and local information of surface defects and provide high-quality input for subsequent deep feature extraction of large models.

[0033] In step S2, deep feature extraction is completed through a multi-head self-attention mechanism based on the Transformer architecture, which specifically includes the following steps: a. Generate query matrix, key matrix and value matrix; b. Generate attention weights through calculation between query matrix and key matrix; c. Use the attention weights to perform weighted summation on the value matrix to generate deep features; d. Extract global and local features through multi-head attention mechanism; In step S2, multi-scale feature fusion is achieved through the feature pyramid network, and feature maps of different scales are synthesized into unified multi-scale features after weight optimization; In step S2, the weights of multi-scale feature fusion are set by optimizing the response levels of high-frequency features and low-frequency features, and are used to perform weighted synthesis of local detail defects and global features.

[0034] Specifically, step S2 uses the AI ​​big model to perform deep feature extraction and multi-scale feature fusion on these preprocessed features based on the illumination invariant feature map and multi-scale feature map generated by data preprocessing. The design focus of step S2 is to extract the deep semantic features of the input data through the big model of the Transformer architecture, and combine the features of different levels and scales through the multi-scale feature fusion network to generate a multi-level defect feature map containing global information and local details.

[0035] In general, the types of defects on the surface of industrial products include scratches, cracks, dents, etc. These defects are diverse in morphology, and defects of different sizes and scales require different feature extraction methods to accurately express. Therefore, using AI large models for deep feature extraction can better capture the complex patterns of surface defects. At the same time, combined with multi-scale feature fusion, it is possible to make full use of features of different scales and improve the overall feature expression ability of the model. As an option, this step uses a Transformer-based multi-head self-attention mechanism to achieve deep feature extraction, and realizes multi-scale feature fusion through a feature pyramid network (FPN).

[0036] In this embodiment, for the preprocessed feature map generated in step S1, a large model based on the Transformer architecture is used to perform deep feature extraction.

[0037] Specifically, in one possible implementation, the input feature map First, generate the query matrix through linear transformation , key matrix Sum Matrix The calculation formula is: ; in: , , They are query matrix, key matrix and value matrix respectively; is the input feature map; , , are weight parameters for the query, key, and value matrices.

[0038] In general, the query matrix and key matrix The dot product of can measure the correlation between different positions in the input feature map. By normalizing the dot product result, the attention weight matrix can be obtained The calculation formula of attention weight is: ; in: is the attention weight matrix; is the dimension of the key matrix, used to normalize the dot product result; is the normalization function.

[0039] As an alternative, using the attention weight matrix Pair Matrix Perform a weighted sum to generate the output feature matrix: ; in, Represents the extracted deep features.

[0040] In some embodiments, in order to further enhance the feature extraction capability, a multi-head self-attention mechanism is used to model the input feature map from multiple angles. Specifically, the query matrix, key matrix, and value matrix are divided into multiple subspaces, the attention weights and weighted sum results are calculated in each subspace, and the results of all subspaces are concatenated to form a multi-head attention output: ; in: and Respectively represent The attention weights and value matrices of the subspaces; Indicates the number of long positions; is the linear transformation weight after concatenation.

[0041] Through the multi-head self-attention mechanism, this embodiment can simultaneously capture the global information and local detail information in the input feature map.

[0042] In this embodiment, a feature pyramid network (FPN) is used to implement multi-scale feature fusion for the multi-scale features generated in step S1 and the features obtained by deep feature extraction.

[0043] In one possible implementation, multi-scale feature maps are extracted from different layers of the Transformer model. , ,…, , where each feature map corresponds to a different semantic level and scale. To achieve multi-scale fusion, these feature maps are upsampled and weighted fused layer by layer. The fusion formula is: ; in: is the fused multi-scale feature map; For the Feature map of the layer; For the Weight parameters of the layer feature map.

[0044] In general, the weight parameter The setting of can be adjusted according to the importance of features at different levels. In some embodiments, the weight parameters can be automatically determined by back propagation optimization during the training process.

[0045] As an option, in order to further improve the effect of multi-scale feature fusion, different weight parameters can be set for high-frequency features and low-frequency features. Specifically: ; in: Represents high-frequency features; Represents low-frequency features; is the weight parameter of high-frequency features, with a value range of 0≤ ≤1.

[0046] Through the above-mentioned multi-scale feature fusion method, this embodiment can effectively combine the features of different levels and scales, so that the generated multi-level defect feature map Has stronger representation ability.

[0047] In step S3, the knowledge distillation technique includes the following steps: a. Use the teacher model to generate deep features for input data; b. Use the student model to learn the characteristic output of the teacher model; c. Optimize the student model by minimizing the cross entropy loss between the true label and the student model output, and the difference loss between the teacher model and the student model output; In step S3, the lightweight model is optimized through model quantization. Model quantization discretizes the weights of the model, converts the weights from floating-point representation to quantized representation, and uses the quantized weights for matrix operations.

[0048] Specifically, the core goal of the S3 step is to optimize the generated large model and transfer the knowledge of the large model to the lightweight model using knowledge distillation technology, thereby significantly reducing the computational complexity and model size while maintaining high detection accuracy. The lightweight optimized model can run more efficiently on resource-constrained edge devices, laying the foundation for real-time surface defect detection in subsequent steps.

[0049] In general, the high performance of deep learning models usually relies on complex structures and a large number of parameters, but such models have high requirements for computing resources and are difficult to deploy directly on edge devices in industrial scenarios. As an option, through knowledge distillation technology, key features and knowledge can be extracted from complex large models (i.e., teacher models) and migrated to simple small models (i.e., student models), thereby achieving lightweight models while maintaining performance as much as possible. In some embodiments, in order to further reduce the complexity of the model, the present invention also combines model quantization technology to reduce storage requirements and computing overhead by converting the floating-point weights of the model into fixed-point weights.

[0050] In this embodiment, the optimization from a large model to a lightweight model is achieved through knowledge distillation technology. Specifically, the teacher model is used to generate deep features or predict outputs, and the student model is trained by imitating the behavior of the teacher model, thereby achieving the purpose of streamlining the model.

[0051] In one possible implementation, the knowledge distillation process includes the following steps: First, the teacher model is used to generate deep features or prediction outputs of the input data as soft labels. Generally, the prediction output of the teacher model not only contains the classification results, but also retains the correlation information between different categories. This correlation information is critical to the optimization of the model.

[0052] Specifically, the predicted probability distribution of the teacher model is It is usually calculated by the following formula: ; in: Indicates The predicted probability of the class; Represents the output of the teacher model The raw score of the class; Represents the temperature parameter used to smooth the prediction distribution.

[0053] As an alternative, in order to enable the student model to learn the probability distribution of the teacher model output, by minimizing the teacher model output distribution and the student model output distribution The KL divergence between is optimized. Its loss function is expressed as: ; in: Represents the student model Class prediction probability; Represents the difference loss between the teacher model and the student model.

[0054] In addition, the training of the student model also incorporates the cross entropy loss between the true labels and the model predictions The final loss function is defined as: ; in: is the weighting coefficient of loss; represents the total loss of knowledge distillation.

[0055] In some embodiments, in order to further improve the learning ability of the student model, real labels and soft labels can be input simultaneously during the training process so that the student model can learn the knowledge of the teacher model more comprehensively.

[0056] In this embodiment, in order to further reduce the computational complexity and storage requirements of the lightweight model, model quantization technology is applied to the student model.

[0057] Generally, the weight parameters of deep learning models are represented by floating-point numbers, and floating-point operations have a large resource overhead. As an option, by quantizing floating-point weights into low-precision fixed-point numbers, the computational complexity and storage requirements can be effectively reduced.

[0058] In one possible implementation, model quantization is achieved through the following formula: ; in: Represents the quantized weight value; represents the raw floating point weights; Represents the quantization step size, and its calculation formula is: ; in, and are the maximum and minimum values ​​of the weight, respectively. is the quantization bit width.

[0059] As a possible implementation, the quantized weights and activation values ​​are stored and calculated using integers, thereby reducing the overhead of floating-point calculations.

[0060] In some embodiments, a post-quantization training method may also be combined to fine-tune the model after quantization to reduce the accuracy loss introduced by quantization.

[0061] In step S4, the lightweight model in the edge device is deployed through the deep learning inference acceleration tool, which is used to optimize the model inference speed and computing efficiency.

[0062] Specifically, step S4 is to deploy the optimized lightweight model to the industrial edge device to achieve real-time detection of surface defects of industrial products. This step mainly solves the problem of limited computing resources and high real-time requirements in industrial environments. By combining deep learning reasoning acceleration tools (such as TensorRT), the present invention can greatly improve the reasoning speed of the model while ensuring detection performance, thereby meeting the requirements for real-time processing in a pipeline environment.

[0063] Generally, when a deep learning model runs on an edge device, it needs to be adapted and optimized for the hardware architecture (such as GPU, TPU or FPGA). As an option, the present invention reduces the storage requirements and computational complexity of the lightweight model to an acceptable range for the edge device through model quantization and structure optimization technology, and further optimizes the reasoning process of the model through reasoning acceleration tools. In some embodiments, the calculation graph structure can also be adjusted for a specific hardware platform to further improve the operating efficiency of the model.

[0064] In this embodiment, the lightweight model generated in step S3 is deployed to the edge device to meet the application requirements of the industrial site.

[0065] Specifically, in one possible implementation, the lightweight model is first converted and optimized through a deep learning inference acceleration tool such as TensorRT. TensorRT is a high-performance optimization tool for deep learning inference that can reorganize the model's computational graph and apply a variety of optimization strategies, including layer fusion, precision calibration, and dynamic memory management.

[0066] Generally speaking, the first step of model optimization is to convert the lightweight model into an intermediate representation (IR). In TensorRT, this process can be achieved in the following ways: ; in: represents the original lightweight model; Represents the intermediate representation after conversion.

[0067] As an option, based on the intermediate representation, the model's computational graph is further optimized by layer fusion. The main purpose of layer fusion is to merge multiple consecutive operations (such as convolution, activation function, and batch normalization) into one operation to reduce the number of memory accesses and computational overhead.

[0068] Specifically, suppose the original computation graph contains two consecutive operations and , then after layer fusion optimization, they can be combined into a single operation: ; In addition, in some embodiments, in order to further improve the reasoning speed, the calculation accuracy of the model can be calibrated (PrecisionCalibration). Generally, deep learning models use 32-bit floating point numbers for calculations, but in practical applications, the calculation accuracy can be reduced to 16-bit floating point numbers or 8-bit integers. Through precision calibration, not only can the calculation overhead be significantly reduced, but also high detection performance can be maintained in most scenarios.

[0069] In this embodiment, the optimized lightweight model performs real-time surface defect detection tasks on edge devices.

[0070] Specifically, in a possible implementation, the real-time detection process includes the following steps: First, the surface images collected in the industrial assembly line are input into the model in batch form for inference. Assume that the input image is a batch size of Tensor ,in: Indicates the batch size of images; Indicates the number of channels of an image, such as an RGB image and Represents the height and width of the image respectively.

[0071] The output of the model is the defect detection result ,in Represents the number of categories or defect types predicted by the model.

[0072] As an option, in order to improve real-time performance, non-essential steps in the inference process (such as storage of intermediate feature maps) can be pruned to reduce memory usage and latency. In some embodiments, the batch size can also be dynamically adjusted based on the pipeline speed. , to balance inference speed and throughput.

[0073] In addition, in order to ensure the accuracy of the detection results, this embodiment adds a confidence calculation module in the reasoning process. Generally, the confidence calculation formula is: ; in: Indicates The highest confidence level of samples; Indicates The samples belong to The probability of the class.

[0074] Generally speaking, the hardware architecture of edge devices (such as NVIDIA Jetson, Google TPU or other embedded platforms) has limited computing resources. Therefore, during the model deployment process, it is necessary to adapt and optimize according to the hardware characteristics.

[0075] As an option, when running on edge devices targeting GPU architecture, the convolution operation can be accelerated using TensorRT's CUDA kernel. The optimization formula for the convolution operation is: ; in: is the output of the convolution operation; is the input feature map; is the convolution kernel weight; For bias.

[0076] In some embodiments, the parallel computing capability of FPGA can also be utilized to improve the throughput of the model by adjusting the structure of the data flow graph.

[0077] In step S5, dynamically adjusting the detection threshold includes the following steps: a. Statistical defect detection results confidence distribution, calculate the mean and standard deviation of the confidence; b. According to the set adjustment parameters, the detection threshold is dynamically adjusted using the mean and standard deviation.

[0078] Specifically, step S5 aims to optimize the surface defect detection process in step S4, and improve the adaptability and stability of the model in industrial application scenarios by dynamically adjusting the detection threshold. Surface defect detection in industrial environments usually requires a balance between missed detection rate and false detection rate, and a single fixed threshold may not be able to cope with the complex requirements under different industrial conditions. The present invention uses a dynamic threshold adjustment mechanism to adaptively adjust the threshold according to the confidence distribution of the detection result to ensure the reliability of the detection result.

[0079] In general, dynamic threshold adjustment is optimized based on the confidence distribution of the model output. As an option, the present invention dynamically calculates the detection threshold by statistically analyzing the mean and standard deviation of the confidence and combining adjustable parameters, so that the system can adapt to the defect characteristics in different scenarios. In some embodiments, the threshold adjustment mechanism can be further optimized in combination with the global characteristics of batch detection data to enhance the robustness of the system.

[0080] In this embodiment, the process of dynamically adjusting the detection threshold includes two main steps: confidence statistics and threshold calculation.

[0081] In one possible implementation, the output of the model on a batch of detection samples is first obtained, and the model outputs the probability distribution of the category to which each sample belongs. For each sample, the system extracts the highest confidence value from the output probability distribution as the detection confidence of the sample.

[0082] In order to calculate the distribution characteristics of the overall test results, the mean and fluctuation range of the confidence of all samples in the current batch can be calculated. For example, by calculating the average value of the sample confidence, the overall credibility of the model in the current batch can be reflected, and by calculating the dispersion index of the confidence, the fluctuation degree of the model output results can be quantified.

[0083] Specifically, the adjustment of the dynamic threshold is based on the distribution characteristics of the confidence. The dynamically adjusted detection threshold can be calculated by combining the central value and fluctuation range of the confidence. As an option, an adjustable parameter can be introduced to control the offset of the detection threshold relative to the confidence distribution, so as to adapt to the detection needs in different scenarios.

[0084] In some embodiments, in order to further improve the stability of threshold adjustment, the detection threshold can be smoothed by combining the statistical information of the current batch and the historical batches. For example, the statistical results of the current batch are combined with the historical data in a weighted manner to reduce the impact of a single batch of abnormal data on the detection threshold.

[0085] In this embodiment, in order to adapt to the detection requirements in different industrial scenarios, the dynamic adjustment mechanism can be optimized.

[0086] In a possible implementation, when the industrial assembly line speed is fast and the real-time requirement is high, a larger batch size can be selected to ensure the stability of the statistical results. Conversely, when there are many categories of defect detection and the confidence distribution is complex, a smaller batch size can be selected to reduce the impact of differences between batches on threshold calculation.

[0087] Specifically, the choice of batch size is closely related to the production scenario. For example, in high-speed pipeline detection, hundreds of images may need to be processed per second. In this case, a larger sample batch size can be selected to ensure that the dynamic adjustment of the threshold can fully reflect the global characteristics.

[0088] In some embodiments, the dynamic adjustment mechanism of the present invention can be used in a variety of industrial scenarios. For example, for mass-produced assembly line products, when the confidence of the detection results is concentrated in the high range, the threshold can be appropriately increased to reduce false detections; and when the confidence of the detection results is more dispersed, the threshold can be lowered to reduce missed detections.

[0089] Specifically, when the dynamically adjusted threshold is higher than the basic threshold of the system, the number of detection results may decrease, but the reliability of the detection results is higher; when the dynamically adjusted threshold is lower than the basic threshold, the system can capture more potential defects, but false detection results may need to be further processed.

[0090] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. The method for detecting surface defects of industrial products based on AI large model is characterized by: The following steps are involved: S1. Perform data preprocessing on the surface images of industrial products to generate illumination invariant features and multi-scale features; S2. Use the AI ​​big model to extract deep features from the preprocessed image, and generate multi-level defect feature maps through multi-scale feature fusion; S3. Optimize the AI ​​large model through knowledge distillation technology to generate a lightweight model; S4, deploy the lightweight model to the edge device for surface defect detection; S5. Optimize the confidence determination during the detection process by dynamically adjusting the detection threshold.

2. The method for detecting surface defects of industrial products based on AI large model according to claim 1 is characterized in that: In the step S1, the illumination invariant feature extraction of data preprocessing includes: The pixel gradient of the input image is calculated to obtain the gradient direction and amplitude of the pixel value, and the pixel gradient is normalized to eliminate the influence of light intensity and generate an illumination invariant feature map.

3. The industrial product surface defect detection method based on AI large model according to claim 1 is characterized in that: In the step S1, the multi-scale feature decomposition of data preprocessing includes: Through multi-scale decomposition, high-frequency features and low-frequency features are extracted from the input image, and the high-frequency features and low-frequency features are weighted fused according to the set weight parameters to generate a fused feature map containing multi-scale information.

4. The method for detecting surface defects of industrial products based on AI large model according to claim 1 is characterized in that: In the S2 step, deep feature extraction is completed through a multi-head self-attention mechanism based on the Transformer architecture, which specifically includes the following steps: a. Generate query matrix, key matrix and value matrix; b. Generate attention weights through calculation between query matrix and key matrix; c. Use the attention weights to perform weighted summation on the value matrix to generate deep features; d. Extract global and local features through multi-head attention mechanism.

5. The method for detecting surface defects of industrial products based on AI large model according to claim 1 is characterized in that: In the step S2, multi-scale feature fusion is achieved through a feature pyramid network, and feature maps of different scales are synthesized into a unified multi-scale feature after weight optimization.

6. The method for detecting surface defects of industrial products based on AI large model according to claim 1 is characterized in that: In the S3 step, the knowledge distillation technology includes the following steps: a. Use the teacher model to generate deep features for input data; b. Use the student model to learn the characteristic output of the teacher model; c. The student model is optimized by minimizing the cross entropy loss between the true labels and the student model output, and the difference loss between the teacher model and the student model output.

7. The method for detecting surface defects of industrial products based on AI large model according to claim 1 is characterized in that: In the step S3, the lightweight model is optimized by model quantization, and the model quantization discretizes the weights of the model, converts the weights from floating point representation to quantized representation, and uses the quantized weights to perform matrix operations.

8. The method for detecting surface defects of industrial products based on AI large model according to claim 1 is characterized in that: In the S4 step, the lightweight model in the edge device is deployed through a deep learning inference acceleration tool, and the inference acceleration tool is used to optimize the model inference speed and computing efficiency.

9. The method for detecting surface defects of industrial products based on AI large model according to claim 1 is characterized in that: In the step S5, dynamically adjusting the detection threshold comprises the following steps: a. Statistical defect detection results confidence distribution, calculate the mean and standard deviation of the confidence; b. According to the set adjustment parameters, the detection threshold is dynamically adjusted using the mean and standard deviation.

10. The method for detecting surface defects of industrial products based on AI large model according to claim 1, characterized in that: In the step S2, the weight of the multi-scale feature fusion is set by optimizing the response degree of the high-frequency features and the low-frequency features, and is used to perform weighted synthesis of local detail defects and global features.

Citation Information

Patent Citations

  • Strip steel surface area type defect identification and classification method

    CN104866862A

  • Defect detection method for irregular metal machining surface based on depth learning

    CN109636772A

  • Steel plate surface defect detection system and method based on machine vision

    CN110873718A

  • Steel tiny defect detection method suitable for edge equipment deployment

    CN118396947A

  • Image detection method for strip steel surface defects

    CN118446981A

Cited By

  • Lightweight 3D medical image real-time reasoning method and system based on edge calculation

    CN120339267A

  • Lightweight 3D medical image real-time reasoning method and system based on edge computing

    CN120339267B

  • Multi-scale interpretable deep learning micron-sized surface defect detection method

    CN120746971A

  • Silica gel product surface defect detection method based on AI vision

    CN122243985A

  • An AI vision-based silicone product surface defect detection method

    CN122243985B