Metalearning-based double-flow perception cross-flattening Transform defect detection method

By employing a meta-learning-based dual-stream sensing cross-flattening Transformer method, the feature extraction and location information preservation of both query and support images are enhanced, solving the problem of low defect detection accuracy with small sample sizes and achieving efficient and low-cost defect detection.

CN121074508APending Publication Date: 2025-12-05SHANDONG UNIV OF TECH +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511233739.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-01
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing technologies cannot effectively solve the problem of surface defect detection methods with small sample sizes, especially for surface defect detection methods with low accuracy, high cost and poor adaptability.

Method used

A meta-learning-based dual-stream sensing cross-flattening Transformer defect detection method is adopted. This method enhances the joint feature extraction capability between the query image and the support image, preserves key location information, explicitly preserves directional spatial information using a location-aware module, and improves feature extraction capability through a dual-stream adaptive module, thereby improving detection accuracy.

Benefits of technology

It achieves high-precision defect detection with a very small number of samples, reduces reliance on large amounts of labeled data, and reduces labeling costs and time, making it suitable for detecting diverse defect types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121074508A_ABST
    Figure CN121074508A_ABST
Patent Text Reader

Abstract

The invention discloses a double-flow perception cross-flattening Transform defect detection method based on meta learning. The method comprises the following steps: 1) sending a query image and a support image to a cross-flattening Transform to obtain a corresponding embedding; 2) carrying out joint feature learning through asymmetric cross-flattening attention; 3) reserving directional information of the query features and the support features through position perception enhancement; 4) performing specific task fine tuning by using double-flow self-adaption; 5) generating a candidate box through a region proposal network, and further fusing information in the query image and the support image; and 6) matching targets in the query image and the support image to obtain a final detection result. The invention provides a new solution for tasks with irregular surfaces and diversified defect types and dependence on independent feature alignment and fusion modules in surface defect detection. Especially in a few-sample scene, the detection precision under limited samples is effectively improved by improving the combined learning ability between the query image and the support image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of defect detection technology, and in particular to a surface defect detection method applicable to situations with a limited number of samples. Background Technology

[0002] Product quality inspection is a core component of industrial production. Traditional surface defect detection technologies rely on training with a large number of defect samples to achieve efficient and accurate defect detection. However, this method, which depends on large-scale labeled samples, faces multiple challenges. First, new or rare defect types may be difficult to detect effectively due to insufficient sample size, leading to a high false negative rate. Second, with the rapid iteration and updates of defect designs, existing detection methods often need to be retrained to adapt to new defect features, which is not only time-consuming but also costly. In few-shot learning scenarios, meta-learning can quickly adjust to identify new types of defects with minimal data support. This approach is particularly crucial for dealing with the frequent emergence of new defect types in industrial production. Although meta-learning has demonstrated strong potential in other fields, its application in surface defect detection still faces technical challenges. These include how to design an architecture that can effectively learn from a small number of samples and how to handle the high diversity and complexity of surface defects. Therefore, there is an urgent need for a new few-shot surface defect detection method designed to overcome these limitations and address the rapidly changing surface defect design and production requirements. The meta-learning-based dual-stream sensing cross-flattening Transformer defect detection method in this invention improves task generalization by enhancing the multi-level interaction capability between the query image and the supporting image, rather than reimplementing existing discovery methods, thereby improving detection accuracy in small sample scenarios. Summary of the Invention

[0003] The purpose of this invention is to overcome the shortcomings and deficiencies of the prior art and provide a meta-learning-based dual-stream sensing cross-flattening Transformer defect detection method. This method breaks through the traditional surface defect detection method that relies on a large number of labeled samples. By enhancing the joint feature extraction capability between query images and support images and retaining key location information, it can quickly learn and accurately identify unknown defects based on a very small number of samples.

[0004] To achieve the above objectives, the present invention provides a method for detecting defects in a dual-stream sensing cross-flattening Transformer based on meta-learning, comprising the following steps:

[0005] 1) Input the query image and supporting images and send them to the first three stages of the cross-flattening Transformer to obtain the corresponding embeddings;

[0006] 2) The query and supporting image embeddings obtained in step 1) are fed into the asymmetric cross-flattening attention for joint feature learning;

[0007] 3) Based on the embedding obtained in step 1), the position-aware module explicitly preserves directional spatial information to enhance discriminative features;

[0008] 4) Based on step 3), the enhanced query features and supporting features are obtained. Through the dual-stream adaptive module, the feature extraction capability for different defect targets is further improved, and the adaptability of Transformer in different tasks is enhanced.

[0009] 5) Based on the features obtained in the first three stages obtained in step 4), the network is fed into the proposal network to generate corresponding candidate boxes, and the relevant information is further integrated in the fourth stage.

[0010] 6) Based on step 5), the integrated features are fed into the detection head to complete the corresponding classification and regression tasks, obtain the final detection results, and calculate the loss.

[0011] In step 1), two sets of categories are given, namely the base class. and new categories ,and The datasets corresponding to the base class and the new class are respectively denoted as... and ,in and This represents an image containing a defective target. and This represents the corresponding target category and its bounding box coordinates; specifically, we follow a task-oriented data organization approach, starting from... and Constructing task sequences Each task consists of a query set. and each category Support set of each instance Composition; by supporting a finite set of samples, the ability to identify targets in the query set is learned, thereby acquiring knowledge of a general structure; subsequently, knowledge will be gained from the base class tasks. Transferring the acquired meta-knowledge to new types of tasks This improves the detection performance for new target classes; the image is fed into PatchEmbedding to obtain the processed query features and support feature embeddings. and .

[0012] In step 2), the embeddings of the query and support images obtained in step 1) are multiplied by the corresponding learnable matrices through multi-head attention to obtain the corresponding Q, K and V matrices; by introducing a mapping function, the attention is changed to linear attention to reduce computational complexity and ensure the performance of attention; the batch of query and support images is aligned, and then the attention of continuous interaction is calculated to effectively emphasize the co-occurrence region;

[0013] a. The formulas for converting query features and supporting features into Q, K, and V matrices are as follows:

[0014]

[0015]

[0016]

[0017] in Represents the number of heads. This represents a space reduction operation;

[0018] b. The calculation formula for introducing the mapping function as the focusing enhancement function is as follows:

[0019]

[0020] in This represents the element-wise exponentiation operation. This indicates a focus enhancement factor, used to increase the weight of important regions;

[0021] c. The formula for batch aligning the query image with the supporting images is as follows:

[0022]

[0023]

[0024]

[0025] in The representative will use tensors repeat Second-rate;

[0026] d. The formula for calculating the attention between query features and supporting features is as follows:

[0027]

[0028]

[0029]

[0030] in It is the projection matrix that maps back to the original feature space. This represents a splicing operation.

[0031] In step 3), based on the feature embeddings of the query features and supporting features obtained in step 1), spatial information in the horizontal and vertical directions is explicitly preserved through location-aware enhancement to obtain the final enhanced query features. and supporting features This is to ensure that the spatial representation of defects can be fully preserved;

[0032] a. Taking query features as an example, the average pooling calculation for query features in the horizontal and vertical directions is as follows:

[0033]

[0034] in express , and Represents height and width, ;

[0035] b. Concatenate the feature vectors in the horizontal and vertical directions and feed them into a shared cross-dimensional interaction layer. Dimensionality reduction reduces computation, and complementary information between different channels is learned through channel interaction. The calculation process is as follows:

[0036]

[0037] in This indicates a shared cross-dimensional interaction layer. This represents splicing operations along spatial dimensions;

[0038] c. Split the features back into horizontal and vertical directions to avoid loss of orientation information, and combine... The function learns adaptive weights to more accurately adjust the spatial information of the features. The calculation process is as follows:

[0039]

[0040] in This represents splitting along spatial dimensions;

[0041] d. Multiply the original feature map by the weights in the horizontal and vertical directions to obtain the enhanced features, calculated as follows:

[0042]

[0043] e. Similarly, max pooling is performed in both the horizontal and vertical directions to obtain significant feature information, calculated as follows:

[0044]

[0045]

[0046]

[0047]

[0048] f. Fuse the features from average pooling and max pooling to obtain the final enhanced query features. The calculation is as follows:

[0049]

[0050] In step 4), the enhanced support features and query features are combined according to steps 2) and 3), which can simultaneously capture global dependencies and local spatial information; the combined query features and supporting features The method improves its adaptability to different tasks by using a two-branch adaptive part for task-specific fine-tuning.

[0051] a. Using a bottleneck structure for task fine-tuning, the calculation method is as follows:

[0052]

[0053]

[0054] in Representative level normalization, This represents the scaling factor.

[0055] b. The fine-tuned features are fused into the MLP block via residual connections, calculated as follows:

[0056]

[0057]

[0058] In step 5), the features from the third stage are fed into the proposal network to generate candidate boxes. The relevant information is integrated and the representation is adjusted through the cross-flattening Transformer block in the fourth stage to adapt to the object detection task.

[0059] In step 6), the integrated features from step 5) are fed into the detection head, and the corresponding ROI features are extracted using ROI Align. The average of the supporting image features is then used to match the query image. Finally, the prediction result of the target is obtained, and the corresponding loss is calculated, as shown below:

[0060] a. To reduce computational complexity, an average operation is performed on all supporting images, calculated as follows:

[0061]

[0062] in , and These represent the height and width of the feature map after ROI Alignment. The number of channels in the feature map;

[0063] b. The final loss calculation is as follows:

[0064]

[0065] in This represents the bounding box regression loss. This represents the binary cross-entropy loss.

[0066] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0067] 1. This invention is the first to realize the modeling of long-distance dependency relationship between query image and support image in meta-learning framework. While enhancing the interaction capability of multi-branch features, it retains key location information, effectively solves the defect detection problem in scenarios with irregular surfaces and diverse defect types, and eliminates the dependence on independent feature alignment and fusion modules.

[0068] 2. This invention can combine small sample learning, allowing for effective learning and defect detection even with limited labeled samples. This learning method reduces the dependence on a large amount of labeled data and significantly reduces labeling costs and time.

[0069] 3. The method of the present invention has a wide range of applications in surface defect detection, is highly practical, and has broad prospects in many fields such as medical imaging and remote sensing target detection. Attached Figure Description

[0070] Figure 1 This is a schematic diagram of the logic flow of the present invention.

[0071] Figure 2 This is a framework diagram of the method of the present invention.

[0072] Figure 3 This is a schematic diagram of the asymmetric cross-flattening attention of the present invention.

[0073] Figure 4 This is a schematic diagram of the position sensing function of the present invention.

[0074] Figure 5 This is a schematic diagram of the dual-stream adaptive method of the present invention. Detailed Implementation

[0075] The present invention will be further described in detail below with reference to the embodiments and accompanying drawings, but the embodiments of the present invention are not limited thereto.

[0076] This embodiment discloses a meta-learning-based dual-stream sensing cross-flattening Transformer defect detection method, such as... Figure 1 As shown; the method is logically implemented as follows: Figure 2 As shown, this method feeds the query image and support images into a cross-flattening Transformer block for joint feature extraction; it emphasizes co-occurrence regions through asymmetric cross-flattening attention, thereby fully utilizing the information in the support images; considering the Transformer's insufficient ability to capture local information, it performs position-aware enhancement to preserve key directional information; and it uses dual-stream adaptive fine-tuning to improve the method's adaptability to different defect detection tasks, further improving the final detection accuracy. Specifically, it includes the following steps:

[0077] 1) In Figure 2 Given a query image and a set of support images, the images are fed into Patch Embedding to obtain the processed query features and support feature embeddings. and ;like Figure 3 As shown, the embedded query and supporting images are multiplied by the corresponding learnable matrix through multi-head attention to obtain the corresponding Q, K and V matrices; by introducing a mapping function, the attention is optimized into linear attention to reduce computational complexity and ensure the performance of attention; the batch of query and supporting images is aligned to calculate the attention of continuous interaction, effectively emphasizing the co-occurrence region;

[0078] a. The formulas for converting query features and supporting features into Q, K, and V matrices are as follows:

[0079]

[0080]

[0081]

[0082] in Represents the number of heads. Represents space reduction operation; specific implementation methods, for example... Figure 3 Given the features of the query image and supporting features , Represents the number of channels. Represents the height of the image. Represents the width of the image;

[0083] After this step, the Q, K, and V matrices corresponding to the query and supporting images are obtained, in order to Figure 3 middle and For example, the corresponding Q, K, and V matrices are respectively , , and , , ;

[0084] b. The calculation formula for introducing the mapping function as the focusing enhancement function is as follows:

[0085]

[0086] in This represents the element-wise exponentiation operation. This represents the focus enhancement factor, used to increase the weight of important regions; Q and K after focus enhancement are respectively , and , ;

[0087] c. The formula for batch aligning the query image with the supporting images is as follows:

[0088]

[0089]

[0090]

[0091] in The representative will use tensors repeat Secondly, query branches and support branches typically have different batch sizes, and directly performing cross-branch attention computation can lead to dimensionality mismatch issues. Specifically, query branches often perform detection on a single image basis, while support branches may contain multiple sample images from a single category. To align the batch sizes of the query branches, the key and value of the support branches are pooled to obtain the result. and Then, the key-value pairs of the two branches are concatenated to obtain the interaction features of the query branch. and In addition, to align with the batch size of supported branches, the key and value of the query branch are repeated. Next, the key-value pairs of the two branches are concatenated to obtain the interaction features supporting the branches. and ;

[0092] d. The formula for calculating the attention between query features and supporting features is as follows:

[0093]

[0094]

[0095]

[0096] in It is the projection matrix that maps back to the original feature space. The concatenation operation represents the final query features and support features obtained after asymmetric cross-flattening attention. and .

[0097] 2) The process of preserving directional information through location-aware enhancement of query features and supporting features is as follows:

[0098] a. Taking query features as an example, the average pooling calculation for query features in the horizontal and vertical directions is as follows:

[0099]

[0100] in express , and Represents height and width, The average pooling results in the horizontal and vertical directions are respectively and ;

[0101] b. Concatenate the feature vectors in the horizontal and vertical directions and feed them into a shared cross-dimensional interaction layer. Dimensionality reduction reduces computation, and complementary information between different channels is learned through channel interaction. The calculation process is as follows:

[0102]

[0103] in This indicates a shared cross-dimensional interaction layer. This represents a splicing operation along the spatial dimension; the features after sharing the cross-dimensional interaction layer are... ;

[0104] c. Split the features back into horizontal and vertical directions to avoid loss of orientation information, and combine... The function learns adaptive weights to more accurately adjust the spatial information of the features. The calculation process is as follows:

[0105]

[0106] in This represents splitting along the spatial dimension; the resulting horizontal and vertical features are respectively... and ;

[0107] d. Multiply the original feature map by the weights in the horizontal and vertical directions to obtain the enhanced features, calculated as follows:

[0108]

[0109] The result of query features after average pooling and max pooling is ;

[0110] e. Similarly, max pooling is performed in both the horizontal and vertical directions to obtain significant feature information, calculated as follows:

[0111]

[0112]

[0113]

[0114]

[0115] f. Fuse the features from average pooling and max pooling to obtain the final enhanced query features. The calculation is as follows:

[0116]

[0117] Taking image queries as an example, the enhanced query features ;

[0118] 3) The process of fine-tuning the enhanced features for specific tasks is as follows:

[0119] a. Using a bottleneck structure for task fine-tuning, the calculation method is as follows:

[0120]

[0121]

[0122] in Representative level normalization, Represents the scaling factor; the fine-tuned query features and supporting features are respectively and .

[0123] b. The fine-tuned features are fused into the MLP block via residual connections, calculated as follows:

[0124]

[0125]

[0126] The fused query features and supporting features are respectively and ;

[0127] 4) The features from the third stage are fed into the proposal network to generate candidate boxes. The relevant information is integrated and the representation is adjusted through the cross-flattening Transformer block in the fourth stage to adapt to the object detection task.

[0128] 5) The integrated features are fed into the detection head, and the corresponding ROI features are extracted through ROI Align to complete the matching between the supporting image and the query image. The process is as follows:

[0129] a. To reduce computational complexity, an average operation is performed on all supporting images, calculated as follows:

[0130]

[0131] in , and These represent the height and width of the feature map after ROI Alignment. The number of channels in the feature map;

[0132] b. The final loss calculation is as follows:

[0133]

[0134] in This represents the bounding box regression loss. This represents the binary cross-entropy loss; the detection result of the final image obtained by the detection head includes the category of the detected object and the category score. For example, given a defect image, we detect that the type of defect at the current location is short circuit, and the corresponding score is 0.88.

[0135] In summary, by adopting the above scheme, this invention provides a new method for solving the problem of diverse defect types and feature alignment in existing surface defect detection with limited samples. It can effectively retain orientation-sensitive information by extracting joint features, thereby improving the detection accuracy in small sample scenarios. It has practical application value and is worth promoting.

[0136] The above-described embodiments are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Therefore, any changes made in accordance with the shape and principle of the present invention should be covered within the protection scope of the present invention.

Claims

1. A meta-learning-based dual-flow perception cross-flattening Transformer defect detection method, comprising the following modules and functional steps: 1) Input query image and support image into the first three stages of cross-flattening Transformer to obtain corresponding embeddings; 2) Send the query and support image embeddings obtained in step 1) to asymmetric cross-flattening attention for joint feature learning; 3) Use the position perception module to explicitly retain directional spatial information and enhance discriminative features based on the embeddings obtained in step 1); 4) Further improve the feature extraction capability for different defect targets and enhance the adaptability of the Transformer in different tasks through the dual-flow adaptive module based on the enhanced query and support features obtained in step 3); 5) Send the features obtained in step 4) into the proposal network to generate the corresponding candidate box, and further integrate the corresponding information through the fourth stage; 6) Send the integrated features in step 5) to the detection head to complete the corresponding classification and regression tasks and calculate the loss.

2. The dual-aware cross-attention flattening Transformer defect detection method based on meta-learning according to claim 1, characterized in that: In step 1), two sets of categories are given, namely the base class. and new categories ,and The datasets corresponding to the base class and the new class are denoted as follows: and ,in and This represents an image containing a defective target. and This represents the corresponding target category and its bounding box coordinates; specifically, we follow a task-oriented data organization approach, starting from... and Constructing task sequences Each task consists of a query set. and each category Support set of each instance Composition; by supporting a finite set of samples, the ability to identify targets in the query set is learned, thereby acquiring knowledge of a general structure; subsequently, knowledge will be gained from the base class tasks. Transferring the acquired meta-knowledge to new types of tasks This improves the detection performance for new target classes; the image is fed into Patch Embedding to obtain the processed query features and support feature embeddings as follows: and .

3. The dual-aware cross-attention flattening Transformer defect detection method based on meta-learning according to claim 1, characterized in that: In step 2), the embeddings of the query and support images obtained in step 1) are multiplied by the corresponding learnable matrix through multi-head attention to obtain the corresponding Q, K and V matrices; the mapping function is introduced to optimize the attention to linear attention to reduce the computational complexity and ensure the performance of the attention; the query and support image batches are aligned, and then the attention of the continuous interaction is calculated to effectively emphasize the co-occurrence region; a, the calculation formula for converting the query and support features into Q, K and V matrices is as follows: ; ; ; wherein representing the number of heads, representing a spatial downscaling operation; b, the calculation formula for introducing the mapping function as the focus enhancement function is as follows: ; wherein represents an element-wise power operation, represents a focus enhancement factor to boost the weight of important regions; c, the calculation formula for aligning the query image and support image batches is as follows: ; ; ; wherein represents the tensor repeat times; d, the formula for attention calculation of query and support features is as follows: ; ; ; wherein is a projection matrix mapping back to the original feature space, represents a stitching operation.

4. The dual-aware cross-attention flattening Transformer defect detection method based on meta-learning according to claim 1, characterized in that: In step 3), the feature embeddings of the query feature and the support feature obtained according to step 1) are enhanced by position awareness, and the spatial information in the horizontal direction and the vertical direction is explicitly reserved to obtain the final enhanced query feature and support feature and support feature , so as to ensure that the spatial representation of the defect can be completely reserved. a, taking the query feature as an example, the horizontal and vertical direction average pooling calculation of the query feature is as follows: ; wherein represents , and represent height and width, ; b, the horizontal and vertical direction feature vectors are spliced and sent to the shared cross-dimensional interaction layer, which reduces the calculation through dimension reduction processing and learns the complementary information between different channels through channel interaction, and the calculation process is as follows: ; wherein represents a shared cross-dimension interaction layer, represents a stitching operation along the spatial dimension; c. Split the features back to horizontal and vertical directions to avoid the loss of direction information, combined with Function learning adaptive weights to adjust the spatial information of features more accurately, the calculation process is as follows: ; wherein represents a split along a spatial dimension; d, multiply the original feature map by the horizontal and vertical direction weights to obtain the enhanced feature, and the calculation is as follows: ; e, similarly, maximum pooling is performed in the horizontal and vertical directions to obtain significant feature information, and the calculation is as follows: ; ; ; ; f, the final enhanced query feature is obtained by fusing the average-pooled and max-pooled features , which is calculated as follows: 。 5. The dual-aware cross-attention flattening Transformer defect detection method based on meta-learning according to claim 1, characterized in that: In step 4), the enhanced support features and the query features are combined according to steps 2) and 3), so that global dependency and local spatial information can be captured simultaneously; the combined query features and support features Through task-specific fine-tuning by the double-branch adaptive part, the adaptability of the method to different tasks is improved. a, the bottle-neck structure is used for task fine-tuning, and the calculation method is as follows: ; ; wherein representative layer normalization, representative scaling factor; b, the fine-tuned features are fused into the MLP block through residual connection, and the calculation method is as follows: ; ; wherein representative layer normalization.

6. The dual-aware cross-attention flattening Transformer defect detection method based on meta-learning according to claim 1, characterized in that: In step 5), the features through the third stage in step 4) are sent to the proposal network to generate the candidate box, and the corresponding information is integrated and adjusted through the fourth stage cross-flattening Transformer block to adapt to the target detection task.

7. The dual-aware cross-attention flattening Transformer defect detection method based on meta-learning according to claim 1, characterized in that: In step 6), the features integrated in step 5) are sent to the detection head, the corresponding ROI features are extracted through ROI Align, the support image features are averaged for matching with the query image, and the corresponding loss is calculated, and the calculation expression is as follows: a、To reduce the computational complexity, the average operation is performed on all support images, and the calculation method is as follows: ; wherein , and are the height and width of the ROI Align post-feature map, respectively, is the number of channels of the feature map; b、The final loss is calculated as follows: ; wherein denotes the bounding box regression loss, denotes the binary cross-entropy loss.