Surface defect detection method and system

By unifying the global and local feature fusion method of teacher-student learning network, the problem of insufficient information in surface defect detection is solved, and high-precision and low-cost detection effects are achieved.

CN120278971APending Publication Date: 2025-07-08FUDAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510355820.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art is difficult to take into account fine-grained texture and complex background context information in surface defect detection, resulting in missed or missed detection, and the model parameters and calculation amount are too large, increasing deployment cost.

Method used

A unified teacher-student learning network is adopted, and the teacher model is used to extract global context features, combined with the depth separation convolution and multi-attention mechanism in the student model, capture fine-grained local information, and achieve efficient fusion of global and local features.

Benefits of technology

It improves the detection accuracy and robustness of small targets and diversified defects, reduces model parameters and calculation amount, and is suitable for surface defect detection in complex industrial environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278971A_ABST
    Figure CN120278971A_ABST
Patent Text Reader

Abstract

The invention discloses a surface defect detection method and system, and particularly relates to the technical field of defect detection. The method comprises the following steps: inputting an industrial surface image into a feature mapping module for slicing and vectoring to generate embedded features in a unified form; the embedded features are input into a teacher model for general feature extraction, and a global context coding result is obtained; a coding result is input into a student model, fine-grained local features are captured by integrating an efficient multi-attention module, and feature fusion of global background information and local details is completed through a global local attention module; and finally, the types and positions of the surface defects are output through the classification head module, so that the detection effect on small targets and diversified defects in a complex industrial environment is improved, the detection accuracy and robustness are improved, excellent performance is shown in complex background and small target detection, and the method can be widely applied to industrial quality detection scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of defect detection, and more specifically, to a surface defect detection method and system. Background Art

[0002] With the increasingly wide application of surface defect detection in industrial production. Traditional small target detection or defect recognition methods usually rely only on single local or global features, and it is difficult to take into account fine-grained textures and complex background context information, resulting in easy omission or misdetection when discriminating multiple types of defects (such as linear, punctate or complex-shaped defects). When the prior art extracts and fuses defect region features: it tends to only learn defect features while ignoring general representations, and it is difficult to handle multi-morphological defect scenarios; too simple feature splicing or dual-branch fusion will result in insufficient features or missing context when facing defects with different sizes and texture differences, which is not conducive to improving the robustness of the model; multi-branch or external fusion modules are likely to bring additional model parameters and computational complexity, causing cost pressure for deployment and application.

[0003] To solve the above problems, a technical solution is provided now. Summary of the Invention

[0004] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a surface defect detection method and system to solve the problems raised in the above background art.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] A surface defect detection method includes the following steps:

[0007] Input an industrial surface image into a feature mapping module for slicing and vectorization processing to generate embedded features in a unified form;

[0008] Input the embedded features into a teacher model for general feature extraction to obtain a global context encoding result;

[0009] Input the global context encoding result into a student model, capture fine-grained local features through an integrated efficient multi-attention module, and complete the feature fusion of global background information and local details through a global-local attention module;

[0010] Output the type and location of the surface defect through a classification head module.

[0011] In a preferred embodiment, inputting an industrial surface image into a feature mapping module for slicing and vectorization processing to generate embedded features in a unified form is specifically:

[0012] Slice the entire industrial surface image into multiple non-overlapping image patches of the same size;

[0013] Flatten the image patch to convert the two-dimensional image patch into a one-dimensional feature vector.

[0014] In a preferred embodiment, the embedded features are input into the teacher model for general feature extraction to obtain the global context encoding result, specifically:

[0015] Encode the embedded features and use the multi-head self-attention mechanism to capture global context features;

[0016] Generate query, key, and value vectors for each embedded feature and calculate the attention weights;

[0017] Process the attention output, output the general features, and obtain the global context encoding result.

[0018] In a preferred embodiment, the global context encoding result is input into the student model, and fine-grained local features are captured through an integrated efficient multi-attention module, specifically:

[0019] Extract the local texture information of the general features through depthwise separable convolution;

[0020] Weight the channel dimension to highlight the significant regions and output the local enhanced features.

[0021] In a preferred embodiment, the global background information and local details are fused through a global-local attention module, specifically:

[0022] Extract the global structure information of the input features, learn the global context, and output the global features;

[0023] Extract the key detail information according to the fine-grained small target features and output the local features;

[0024] Generate the fused features by weighted addition of the global features and the local features.

[0025] In a preferred embodiment, the type and location of the surface defects are output through a classification head module, specifically:

[0026] Perform linear transformation and non-linear activation on the fused features;

[0027] Normalize the classification results to category probabilities and output the final detection results, including the defect type and location annotation.

[0028] On the other hand, the present invention provides a surface defect detection system, including a feature mapping unit, a feature extraction unit, a feature capture and fusion unit, and a defect output unit;

[0029] Feature mapping unit: Input the industrial surface image into the feature mapping module for slicing and vectorization processing to generate embedded features in a unified form;

[0030] Feature extraction unit: Input the embedded features into the teacher model for general feature extraction to obtain the global context encoding result;

[0031] Feature capture and fusion unit: Input the global context encoding result into the student model, capture fine-grained local features through an integrated efficient multi-attention module, and complete the feature fusion of global background information and local details through a global-local attention module;

[0032] Defect output unit: Output the type and location of surface defects through the classification head module.

[0033] Technical effects and advantages of the surface defect detection method and system of the present invention:

[0034] By adopting a unified teacher-student learning network, using the teacher model to extract global context features, and then capturing fine-grained local information through depthwise separable convolution and multi-attention mechanisms in the student model, the efficient fusion of global and local features is achieved, thus taking into account both the overall background and local details, improving the detection accuracy and robustness for small targets and diverse defects, while reducing the model parameters and computational amount, being applicable to surface defect detection in complex industrial environments, and having high generality and practicality. Description of the drawings

[0035] Figure 1 It is a schematic diagram of a surface defect detection method of the present invention;

[0036] Figure 2 It is a schematic structural diagram of a surface defect detection system of the present invention. Detailed implementation manners

[0037] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0038] Embodiment 1

[0039] Figure 1 A surface defect detection method of the present invention is given, which includes the following steps:

[0040] Input the industrial surface image into the feature mapping module for slicing and vectorization processing to generate embedded features in a unified form;

[0041] Input the embedded features into the teacher model for general feature extraction to obtain the global context encoding result;

[0042] Input the global context encoding result into the student model, capture fine-grained local features through an integrated efficient multi-attention module, and complete the feature fusion of global background information and local details through a global-local attention module;

[0043] Output the type and location of surface defects through the classification head module.

[0044] Specifically, input the industrial surface image into the feature mapping module for slicing and vectorization processing to generate embedded features in a unified form, including:

[0045] Slice the entire industrial surface image into multiple non-overlapping image patches of the same size; specifically, the same size means that all image patches have the same number of rows and columns to ensure consistent dimensions; non-overlapping means that there is no overlapping part between these sub-image patches, so as to cover different regions of the entire image; when slicing the image, it is necessary to first determine the size of the image patch; for example, an image of 512×512 can be sliced into patches of 32×32 size, then several image patches can be obtained in each direction; if the size of the image patch is set to 32×32, the 512×512 image can be divided into 16 patches in both the horizontal and vertical directions, with a total of 16×16 = 256 image patches;

[0046] Perform a flattening operation on the image patch to convert the two-dimensional image patch into a one-dimensional feature vector; specifically, each image patch is in the form of a two-dimensional matrix when sliced (such as 32×32 pixels); the image also has a number of channels (such as 3 channels for RGB), so the shape of the image patch is more specifically [32, 32, 3]; flattening means arranging these two-dimensional (or three-dimensional) data in a row or by channel order into a one-dimensional vector; for example, if an image patch is 32×32×3, after flattening, a vector with a length of 3072 can be obtained; for the aforementioned 32×32×3 image patch, after flattening, it becomes a one-dimensional vector of 3072 dimensions. For 256 image patches, a feature matrix of 256×3072 or a combination of several vectors is obtained.

[0047] Specifically, input the embedded features into the teacher model for general feature extraction to obtain the global context encoding result, including:

[0048] Encode the embedded features and use the multi - head self - attention mechanism to capture global context features. Specifically, encoding the embedded features means sending each one - dimensional vector into the Transformer encoder in the teacher model. The multi - head self - attention mechanism is the core of Transformer, which is used to learn the correlation relationships between all input vectors, thereby capturing the context information of the image globally. It allows the model to pay attention to the content of other vectors while processing a certain vector. The global context features refer to the model's overall understanding of the entire image, including the associations between its parts. Exemplarily, if there are 256 vectors, the multi - head self - attention mechanism of Transformer will calculate the attention for the remaining 255 vectors based on each vector, obtain their dependency relationships and synthesize them, so as to better acquire the global information of the entire picture. "Multi - head" means using multiple attention heads simultaneously to capture the correlations in different sub - spaces or from different angles, which can often enhance the model's representation ability for complex scenes.

[0049] Generate query, key, and value vectors for each embedded feature and calculate the attention weights. Specifically, the query, key, and value vectors are three core vectors in the multi - head self - attention mechanism, and each image patch corresponds to a set of query, key, and value. Calculate the similarity through the dot - product of the query and the key, and then weight the value to achieve information fusion. Suppose there are N image patches, and each image patch is a D - dimensional vector. In the multi - head self - attention mechanism, query, key, and value matrices of N×D dimensions will be obtained, and then the attention distribution (N×N) will be calculated. This can help the model understand which image patches are more relevant to each other.

[0050] Process the attention output, output general features, and obtain the global context encoding result. Specifically, the result of the attention output is further passed through a feed - forward network or other layers (such as residual connection, layer normalization, etc.) to further enhance the feature expression ability, and finally obtain the global context encoding result or general features. These general features carry information about the overall structure of the image, the background environment, etc. on a large scale and can be used for subsequent finer - grained judgment of defects.

[0051] Specifically, input the global context encoding result into the student model, and capture fine - grained local features through an integrated efficient multi - attention module, including:

[0052] The local texture information of the common features is extracted through the depthwise separable convolution. Specifically, the depthwise separable convolution is an efficient convolution method that can retain the effective spatial features as much as possible while reducing the number of model parameters. Through this convolution operation, each channel can be filtered separately and then merged to obtain clearer fine-grained texture information. For example, the traditional convolution will convolve the input features in both the channel and space dimensions. The depthwise separable convolution first performs convolution in the channel direction and then completes the 1×1 convolution in the channel merging stage, thereby separating the calculation and reducing the amount of calculation.

[0053] The channel dimension is weighted to highlight the salient area and output the local enhanced features. Specifically, based on the convolution to capture the local texture, the response of each channel is weighted through mechanisms such as Squeeze-and-Excitation (SE) to highlight the most useful or significant areas for defect detection. This can be understood as further focusing the attention at the channel level to enhance the local salient information. The local enhanced features finally outputted contain higher resolution texture information and the weighting of important channels.

[0054] Specifically, the feature fusion of global background information and local details is completed through the global-local attention module, including:

[0055] Extract the global structural information of the input features, learn the global context, and output the global features. Specifically, in the global-local attention module, a part of the attention or network operation focuses on capturing the overall pattern and large-scale context of the image (such as where the defect appears, the main color of the background, etc.). This part of the output is called the global feature. The global feature can usually help locate the approximate area of ​​the defect or determine the contrast relationship between the defect and the background.

[0056] According to the fine-grained small target features and extracting key detail information, local features are output; specifically, another part of the attention or network operation focuses more on the local small area, texture edge or the specific shape of the small target itself, etc., to obtain local features; combined with the output of the global local attention module used above, it further emphasizes the ability to distinguish small defect areas that are difficult to detect;

[0057] The global features and local features are weighted added to generate fused features. Specifically, the global features and local features are synthesized with a certain weighting or fusion strategy. This strategy can be a simple addition, splicing, or a more complex attention fusion. Through fusion, both large-scale background information and local fine information can be taken into account at the same time, thereby obtaining a feature representation with both breadth and depth.

[0058] Specifically, the classification head module outputs the type and location of surface defects, including:

[0059] Perform linear transformation and nonlinear activation on the fused features. Specifically, a multi-layer perceptron or a fully connected layer is usually used to further transform the fused features. Linear transformation mainly refers to matrix multiplication, and nonlinear activation refers to ReLU, GELU or other activation functions. This allows the model to learn higher-dimensional mapping relationships and transform the fused features into representations that are easier to distinguish between different defect types.

[0060] The classification results are normalized into category probabilities, and the final detection results are output, including defect types and position annotations. Specifically, the Softmax or Sigmoid function is used on the final output results to obtain the probability distribution of different defect categories. The position annotation can simultaneously regress the coordinates of the defects (e.g., pixel-level coordinates) during network design, or an additional position regression branch can be included in the classification head to output the corresponding coordinate information. The final result includes categories and positions, such as scratches in the (x1, y1, x2, y2) region of the image, spots in another region, etc., to meet the needs of industrial detection. For example, in terms of classification, if three common defect types are supported, such as scratches, cracks, and discolored spots, the output may be [scratches: 0.8, cracks: 0.15, discolored spots: 0.05]; in terms of position, a rectangular box coordinate (upper left corner, lower right corner) or mask form can be returned to indicate the specific position of the defect in the image.

[0061] Example 2

[0062] The difference between Example 2 of the present invention and Example 1 is that this example introduces a surface defect detection system.

[0063] Figure 2 A structural schematic diagram of a surface defect detection system of the present invention is given, wherein the surface defect detection system comprises a feature mapping unit, a feature extraction unit, a feature capture fusion unit and a defect output unit;

[0064] Feature mapping unit: The industrial surface image is input into the feature mapping module for slicing and vectorization to generate embedded features in a unified form;

[0065] Feature extraction unit: inputs the embedded features into the teacher model for general feature extraction to obtain the global context encoding result;

[0066] Feature capture and fusion unit: The global context encoding result is input into the student model, and fine-grained local features are captured by integrating efficient multi-attention modules, and the feature fusion of global background information and local details is completed through the global-local attention module;

[0067] Defect output unit: Outputs the type and location of surface defects through the classification head module.

[0068] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data and performing software simulations to get a formula closest to the actual situation. The preset parameters and threshold selections in the formulas are set by those skilled in the art according to the actual situation.

[0069] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0070] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0071] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and modules described above can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.

[0072] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or modules can be in electrical, mechanical, or other forms.

[0073] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules. They can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0074] In addition, in each embodiment of the present application, the functional modules can be integrated into a processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.

[0075] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0076] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0077] Finally, the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A surface defect detection method, characterized in that, It includes the following steps: Input the industrial surface image into the feature mapping module for slicing and vectorization processing to generate embedded features in a unified form; Input the embedded features into the teacher model for general feature extraction to obtain the global context encoding result; Input the global context encoding result into the student model, capture fine-grained local features through an integrated efficient multi-attention module, and complete the feature fusion of global background information and local details through a global-local attention module; Output the type and location of surface defects through the classification head module.

2. The surface defect detection method according to claim 1, wherein Input the industrial surface image into the feature mapping module for slicing and vectorization processing to generate embedded features in a unified form, specifically: Slice the entire industrial surface image into multiple non-overlapping image patches of the same size; Flatten the image patches to convert the two-dimensional image patches into one-dimensional feature vectors.

3. The surface defect detection method according to claim 2, wherein, Input the embedded features into the teacher model for general feature extraction to obtain the global context encoding result, specifically: Encode the embedded features and capture global context features using the multi-head self-attention mechanism; Generate query, key, and value vectors for each embedded feature and calculate the attention weights; Process the attention output to output general features and obtain the global context encoding result.

4. A surface defect detection method according to claim 3, characterized in that, Input the global context encoding result into the student model, capture fine-grained local features through an integrated efficient multi-attention module, specifically: Extract the local texture information of general features through depthwise separable convolution; Weight the channel dimension to highlight the significant regions and output local enhanced features.

5. A surface defect detection method according to claim 4, characterized in that, Complete the feature fusion of global background information and local details through a global-local attention module, specifically: Extract the global structure information of the input features, learn the global context, and output global features; Extract key detail information based on the fine-grained small target features and output local features; Generate a fusion feature by adding the global features and local features through weighted summation.

6. The surface defect detection method according to claim 5, wherein Output the type and location of surface defects through the classification head module, specifically: Perform linear transformation and non-linear activation on the fusion features; Normalize the classification result to category probabilities and output the final detection result, including defect type and location annotation.

7. A surface defect detection system for implementing the surface defect detection method according to any one of claims 1-6, characterized in that, It includes a feature mapping unit, a feature extraction unit, a feature capture and fusion unit, and a defect output unit; Feature mapping unit: Input the industrial surface image into the feature mapping module for slicing and vectorization processing to generate embedded features in a unified form; Feature extraction unit: Input the embedded features into the teacher model for general feature extraction to obtain the global context encoding result; Feature capture and fusion unit: Input the global context encoding result into the student model, capture fine-grained local features through an integrated efficient multi-attention module, and complete the feature fusion of global background information and local details through a global-local attention module; Defect output unit: Output the type and location of surface defects through the classification head module.