Metal part anomaly classification method based on Manhattan related attention network
By applying Manhattan-related attention network in metal component abnormality detection, combined with the Manhattan attention retention module and structural context attention module, the problem of difficulty in detecting small area defects and similar defect areas in the prior art is solved, and efficient and accurate abnormal detection effect is achieved.
Patent Information
- Application Number
- CN202510151664.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to accurately detect small area defects in metal parts and defect areas similar to normal areas, resulting in problems of low detection efficiency and large errors.
Using the method of metal component anomaly classification based on Manhattan's related attention network (MCA-Net), the spatial relationship and contextual structure dependencies of input features are analyzed and subtle anomaly areas are identified through the Manhattan Attention Retention Module (MRA) and the structural context attention module (SCA).
Accurate detection of defects in small and medium-sized metal parts and difficult to distinguish abnormal areas is achieved, which improves detection efficiency and accuracy and reduces errors.
Smart Images

Figure CN119942231A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of metal anomaly classification, and in particular relates to a metal component anomaly classification method based on Manhattan Correlation Attention Network (MCA-Net). Background Art
[0002] Hundreds of trillions of metal parts are produced worldwide each year in the manufacturing industry, serving a wide range of industries. However, the quality of these parts can vary greatly due to production conditions and external factors. Therefore, it is critical to sort metal parts to ensure high-quality product production.
[0003] Currently, metal anomaly classification and detection is usually performed by experienced workers. However, manual methods are inefficient and prone to significant errors, often resulting in inconsistent product quality and significant losses. Therefore, it is crucial to adopt automatic tools to detect anomalies in metal parts.
[0004] To this end, many computer-assisted methods have been proposed for the anomaly classification of metal parts, which can be divided into traditional machine learning methods and deep learning methods. Traditional machine learning techniques, such as support vector machines and random forests, are often not robust enough and are limited to fixed environments. In contrast, deep learning methods have performed well in metal anomaly detection, mainly due to their powerful feature extraction capabilities. For example, a magnetic anomaly detection method based on deep learning neural networks is proposed, which decouples and denoises the comprehensive detection process to accurately extract magnetic anomalies from a single pipe. In addition, a novel defect classification algorithm is proposed, which combines convolutional variational autoencoders and deep convolutional neural networks for metal anomaly detection.
[0005] However, detecting metal anomalies is difficult due to the following issues. Many defects in metal parts are small and may not be visible to the naked eye and cannot be detected by standard inspection methods, such as Figure 1 (a) and (b). These small-scale anomalies are particularly difficult to identify when the detection technology has limited resolution or sensitivity. In addition, these defects may not significantly affect the overall appearance or function of the metal part, making them even more difficult to detect. Another challenge is that the abnormal areas are often similar to normal areas of the metal. For example, a seemingly uniform metal surface may actually contain underlying defects, or very small anomalies, such as subtle changes in density, structure, or material properties, such as Figure 1 Therefore, in order to distinguish between normal and abnormal areas, highly specialized detection methods, such as advanced imaging techniques, are required to more accurately identify these differences. Summary of the invention
[0006] In view of the above-mentioned deficiencies in the prior art, the metal part anomaly classification method based on Manhattan correlation attention network provided by the present invention solves the problem that the existing metal part defect detection method is difficult to accurately detect small area defects and defect areas similar to normal areas.
[0007] In order to achieve the above-mentioned invention object, the technical solution adopted by the present invention is: a method for classifying metal parts anomalies based on Manhattan-related attention network, comprising the following steps:
[0008] S1, build a metal parts dataset;
[0009] S2. Construct a Manhattan-related attention network and train it using the metal parts dataset to obtain a metal parts anomaly classification network.
[0010] The Manhattan-related attention network includes an input module, a residual module, a Manhattan attention retention module, a structural context attention module, a feedforward neural network module and an output module connected in sequence;
[0011] S3. Use the metal parts anomaly classification network to identify the metal parts image to be detected and obtain the anomaly classification result.
[0012] Furthermore, in the Manhattan-correlated attention network, the input module is used to receive an input metal parts image, which includes a convolutional layer, a batch normalization layer, a ReLU activation function layer and a maximum pooling layer connected in sequence.
[0013] Furthermore, in the Manhattan-related attention network, the residual module is used to extract a feature image of an input image, which includes four ResBlock residual blocks connected in sequence.
[0014] Furthermore, in the Manhattan-related attention network, the Manhattan attention retention module is used to analyze the spatial relationship of the input features and retain the temporal or spatial contextual relationship, and then search for small abnormal areas through global information modeling.
[0015] Furthermore, the method for processing the input features by the Manhattan attention retention module is specifically as follows:
[0016] S21, generate the query q, key k and value v of the input feature through a linear layer respectively;
[0017] S22, based on the calculated query q and key k, calculate the corresponding Manhattan distance;
[0018] S23, calculating the attention weight based on the Manhattan distance product of the query q and the key k;
[0019] S24, combining the attention weight and value v, and calculating the output of each position by weighted sum;
[0020] S25. Introduce a memory retention mechanism to integrate the memory information obtained based on the value v into the current output to obtain an output feature map containing small abnormal areas.
[0021] Furthermore, in step S22, the Manhattan distances corresponding to query q and key k are respectively:
[0022] d manhattan (q) = q·cos(i) + q·sin(i);
[0023] d manhattan (k) = k·cos(i)+k·sin(i);
[0024] Where, d manhattan (q) represents the Manhattan distance of query q, i represents the index coordinate vector, q = {q1, q2, ..., q d} and k={k1,k2,...,k d} denote the vector representation of query and key respectively, i denotes the index coordinate vector, and s is the dimension of the sequence;
[0025] In step S23, the attention weight α of the i-th position i for:
[0026]
[0027] Where λ represents the scaling factor, f i represents the Manhattan distance product of the i-th position, represents the Manhattan distance product of the j-th position;
[0028] In step S24, the output y of each position j ' is expressed as:
[0029]
[0030] In the formula, v j represents the jth position value v, α j represents the attention weight of the jth position;
[0031] In step S25, the output feature value y of the small abnormal area is included j It is expressed as:
[0032]
[0033] In the formula, β represents the learning parameter, m i Represents the memory information from the previous time step.
[0034] Furthermore, in the Manhattan-related attention network, the structural context attention module is used to distinguish subtle abnormal areas from normal areas by modeling local features and spatial dependencies, and aggregate context information to determine the position of the area in the overall structure, thereby obtaining a more refined abnormal area detection feature map.
[0035] Furthermore, the structural context attention module processes the input feature map as follows:
[0036] S26, determining the input feature map x and the feature map set X to which it belongs;
[0037] Among them, the feature map set includes the input feature map x and its adjacent feature maps {x1, x2, ..., x8};
[0038] S27, respectively generating a query feature map, a key feature map, and a value feature map of the input feature map through a convolutional layer, and generating context information based on the feature map set;
[0039] S28, introducing a context-aware mechanism, and performing weighted calculations on the query feature graph, the key feature graph, and the value feature graph respectively in combination with the context information;
[0040] S29, calculating context-enhanced attention based on the context-weighted query feature graph, the key feature graph, and the value feature graph;
[0041] S210. According to the enhanced attention described below, a more refined abnormal region detection feature map is output through a linear layer.
[0042] Further, in step S27, the context information is obtained by performing a convolution operation on the corresponding feature map set; in step S28, the context-weighted query feature map, key feature map and value feature map are respectively expressed as:
[0043] q'=q+r;
[0044] k'=k+s;
[0045] v'=v+t;
[0046] Where r' = W r r,k'=W k k,v'=W v v,W r , W k and W v Represent the learnable parameter matrices corresponding to the query feature map, key feature map, and value feature map respectively;
[0047] In step S29, the contextual enhanced attention ContextualAttenion(q,k,v) is expressed as:
[0048]
[0049] Where, d k represents the kth scaling factor, and T represents the transposition symbol.
[0050] The beneficial effects of the present invention are:
[0051] The method of the present invention proposes a Manhattan Correlation Attention Network (MCA-Net) for metal part anomaly classification, in which a Manhattan Preserving Attention (MRA) module is designed to search for small abnormal areas through global information modeling, and it is good at identifying subtle anomalies that are difficult to detect through local feature analysis; and a Structural Context Attention (SCA) model is designed to distinguish abnormal areas that are visually similar to normal areas by aggregating contextual structural dependencies. It not only considers the characteristics of the region itself, but also its surrounding context and structural relationships, which helps the system to accurately classify abnormal areas. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 Schematic diagram of abnormal metal parts; (a) and (b) represent small abnormal areas; (c) and (d) represent situations where the abnormal areas are similar to the surrounding normal areas.
[0053] Figure 2 A metal parts anomaly classification method based on Manhattan correlation attention network.
[0054] Figure 3 Schematic diagram of the Manhattan-related attention network structure.
[0055] Figure 4 The receiver operating characteristic (ROC) and precision-recall (PR) curves of different models on the metal parts dataset; AUC and AP represent the area under the curve and average precision, respectively.
[0056] Figure 5 Category attention maps of different methods on part datasets; (a) original image; (b) RMT; (c) MetaFormer; (d) EfficientNet-b4; (e) BiFormer; (f) VGG19; (g) ResNet50; (h) GhostNet; (i) DenseNet161; (j) MCA-Net.
[0057] Figure 6Schematic diagram of the ablation results on the parts database; where: (a) original image; (b) true label; (c) benchmark; (d) benchmark + MRA; (e) benchmark + SCA; (f) benchmark + MRA + SCA (MCA-Net). DETAILED DESCRIPTION
[0058] The specific implementation modes of the present invention are described below so that those skilled in the art can understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation modes. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations utilizing the concept of the present invention are protected.
[0059] The embodiment of the present invention provides a method for classifying metal parts anomalies based on Manhattan correlation attention network, such as Figure 2 As shown, the following steps are included:
[0060] S1, build a metal parts dataset;
[0061] S2. Construct a Manhattan-related attention network and train it using the metal parts dataset to obtain a metal parts anomaly classification network.
[0062] The Manhattan-related attention network includes an input module, a residual module, a Manhattan attention retention module, a structural context attention module, a feedforward neural network module and an output module connected in sequence;
[0063] S3. Use the metal parts anomaly classification network to identify the metal parts image to be detected and obtain the anomaly classification result.
[0064] In the embodiment of the present invention, Figure 3 As shown, in the Manhattan correlation attention network, the input module is used to receive the input metal parts image, which includes a convolution layer, a batch normalization layer, a ReLU activation function layer and a maximum pooling layer connected in sequence.
[0065] In the embodiment of the present invention, Figure 3 As shown, in the Manhattan-related attention network, the residual module is used to extract the feature image of the input image, which includes four ResBlock residual blocks connected in sequence.
[0066] In the embodiment of the present invention, Figure 3 As shown, in the Manhattan-related attention network, the Manhattan attention retention module is used to analyze the spatial relationship of the input features and retain the temporal or spatial contextual relationship, and then search for small abnormal areas through global information modeling.
[0067] Specifically, in the manufacturing and quality control of metal parts, detecting defects is critical to ensuring structural integrity and safety. Metal parts, especially in high-stress applications such as automotive, aerospace, and machinery, may develop defects during the production process that can affect performance and durability. While some defects are large and easy to detect, others are small and difficult to identify by traditional recognition methods. Understanding these small defects and effective detection techniques are key to maintaining quality and safety. In this embodiment, the Manhattan Attention Retention Module (MRA) is designed to search for small abnormal areas through global information. It can analyze the input data to capture spatial relationships and retain context in time or space to avoid missing small deviations in distant areas.
[0068] In this embodiment, based on Figure 3 The Manhattan attention retention module structure shown in the figure processes the input features as follows:
[0069] S21, generate the query q, key k and value v of the input feature through a linear layer respectively;
[0070] S22, based on the calculated query q and key k, calculate the corresponding Manhattan distance;
[0071] S23, calculating the attention weight based on the Manhattan distance product of the query q and the key k;
[0072] S24, combining the attention weight and value v, and calculating the output of each position by weighted sum;
[0073] S25. Introduce a memory retention mechanism to integrate the memory information obtained based on the value v into the current output to obtain an output feature map containing small abnormal areas.
[0074] In step S22, the Manhattan distances corresponding to query q and key k are:
[0075] d manhattan (q) = q·cos(i) + q·sin(i);
[0076] d manhattan (k) = k·cos(i)+k·sin(i);
[0077] Where, d manhattan (q) represents the Manhattan distance of query q, i represents the index coordinate vector, q = {q1, q2, ..., q d} and k={k1,k2,...,k d} denote the vector representation of query and key respectively, i denotes the index coordinate vector, and s is the dimension of the sequence;
[0078] In step S23, the Manhattan distance product of the query q and the key k is expressed as:
[0079] f=d manhattan (q)*d manhattan (k);
[0080] The attention weight α of the i-th position i for:
[0081]
[0082] Where λ represents the scaling factor, f i represents the Manhattan distance product of the i-th position, represents the Manhattan distance product of the j-th position;
[0083] In step S24, the output y at each position j ' is expressed as:
[0084]
[0085] In the formula, v j represents the jth position value v, α j represents the attention weight of the jth position;
[0086] In step S25, the output feature value y containing the small abnormal area j It is expressed as:
[0087]
[0088] In the formula, β represents the learning parameter, m i Represents the memory information from the previous time step.
[0089] The Manhattan attention retention module (MRA) provided in this embodiment combines the Manhattan distance-based attention mechanism with the memory retention mechanism to improve the model's ability to capture long-term dependencies (such as global context perception and retention attention); unlike traditional attention mechanisms that focus on local areas, MRA captures the global context by considering the entire spatial layout of the input, and is able to detect patterns and relationships between different regions, which is crucial for identifying subtle anomalies; the retention aspect of MRA enables the model to retain contextual information over time or space, and by remembering previous regions or time steps, it is able to detect anomalies that appear over long distances or long time intervals, even if these anomalies are small or scattered in the input.
[0090] In the embodiment of the present invention, Figure 3As shown in the figure, in the Manhattan correlation attention network, the structural context attention module is used to distinguish subtle abnormal areas from normal areas by modeling local features and spatial dependencies, and aggregate context information to determine the position of the area in the overall structure, thereby obtaining a more refined abnormal area detection feature map.
[0091] Specifically, the challenge begins when the abnormal metal area is very similar to the normal area and difficult to distinguish. This similarity may appear in attributes such as texture, color or structure, and the anomaly only exhibits subtle changes. Traditional detection methods rely on obvious contrasts or predefined patterns and often fail to identify these subtle anomalies, especially when the anomalies are localized or exhibit small deviations, which requires the use of advanced techniques that can capture subtle differences. To address this issue, a structural contextual attention module (SCA) is designed in this embodiment to distinguish subtle anomalies from normal areas by modeling local features and broader spatial dependencies, and aggregate contextual information to understand the location of the region in the overall structure, so that missed anomalies in isolated analysis can be detected.
[0092] In this embodiment, based on Figure 3 The structural context attention module shown in FIG. 1 processes the input feature map as follows:
[0093] S26, determining the input feature map x and the feature map set X to which it belongs;
[0094] Among them, the feature map set includes the input feature map x and its adjacent feature maps {x1, x2, ..., x8};
[0095] S27, respectively generating a query feature map, a key feature map, and a value feature map of the input feature map through a convolutional layer, and generating context information based on the feature map set;
[0096] S28, introducing a context-aware mechanism, and performing weighted calculations on the query feature graph, the key feature graph, and the value feature graph respectively in combination with the context information;
[0097] S29, calculating context-enhanced attention based on the context-weighted query feature graph, the key feature graph, and the value feature graph;
[0098] S210. According to the enhanced attention described below, a more refined abnormal region detection feature map is output through a linear layer.
[0099] In step S27, the context information is obtained by performing a convolution operation on the corresponding feature map set, which is expressed as:
[0100] q=Conv(x),k=Conv(x),v=Conv(x);
[0101] In step S28, the context-weighted query feature graph, key feature graph, and value feature graph are respectively expressed as:
[0102] q'=q+r;
[0103] k'=k+s;
[0104] v'=v+t;
[0105] Where r' = W r r,k'=W k k,v'=W v v,W r , W k and W v Represent the learnable parameter matrices corresponding to the query feature map, key feature map, and value feature map respectively;
[0106] In step S29, the contextual enhanced attention ContextualAttenion(q,k,v) is expressed as:
[0107]
[0108] Where, d k represents the kth scaling factor, and T represents the transposition symbol.
[0109] In the embodiment of the present invention, an example of verifying the effect of the above method is provided.
[0110] In this embodiment, MCA-Net is evaluated on a parts dataset, where images are collected from different public parts datasets. Specifically, the dataset contains 13 part categories, with a total of 4898 images, including bolts, black brackets, brown brackets, white brackets, cables, connectors, metal nuts, metal plates, thumb tacks, screws, cup head bolts, transistors, and sleeves, with 400, 368, 262, 170, 374, 172, 335, 151, 889, 480, 761, 313, and 223 images, respectively. During the experiment, the dataset was divided into a training set and a test set, with 3878 and 1020 images, respectively, as shown in Table 1 for details;
[0111] Table 1: Details of the training and test sets of the parts dataset
[0112]
[0113] In order to verify the effectiveness of the proposed MCA-Net, accuracy (Acc), recall (Rec), precision (Pre) and F1 score (F1) are used as evaluation indicators, and receiver operating characteristic (ROC) and precision-recall (PR) curves are also used to show the superiority of the proposed MCA-Net.
[0114] In this example, in order to verify the performance of the proposed MCA-Net, five classic state-of-the-art models, namely VGG19, ResNet50, GhostNet], EfficientNet-b4 and DenseNet161, and three Transform-based variants, namely RMT, MetaFormer and BiFormer, were tested on the parts dataset. The results are shown in Tables 2 and Figure 4 .
[0115] Table 2: Performance of different deep learning models on the parts dataset
[0116] method Accuracy (%) Accuracy (%) recall(%) F1 score (%) RMT 78.92 78.92 83.40 70.45 MetaFormer 79.41 79.41 82.38 71.68 EfficientNet-b4 80.58 80.58 78.52 76.87 BiF0rmer 82.06 82.06 82.52 77.67 VGG19 86.27 86.27 85.56 85.65 ResNet50 86.86 86.86 86.41 85.67 GhostNet 87.54 87.54 87.43 86.29 DenseNet161 88.43 88.43 87.97 88.03 MCA-Net 91.76 91.76 91.69 91.35
[0117] As shown in Table 2, the proposed MCA-Net outperforms all other methods in metal part anomaly detection on the parts dataset with an accuracy of 91.76%, precision of 91.76%, recall of 91.69%, and F1 score of 91.35%. Figure 4 It can be inferred that the proposed MCA-Net performs best in terms of area under the curve (AUC) and average precision (AP). These results demonstrate the effectiveness of MCA-Net in detecting anomalies in metal parts. Compared with Transformer-based variants such as RMT, MetaFormer, and BiFormer, MCA-Net consistently shows superior performance, highlighting its advantage in leveraging the Transformer attention mechanism. In addition, MCA-Net outperforms the classic models, with improvements of 11.18%, 11.18%, 13.17%, and 14.48% in accuracy, precision, recall, and F1 score, respectively, compared to EfficientNet-b4. This shows that MCA-Net's ability to capture global information and contextual dependencies enhances its performance in metal part anomaly detection. In summary, these results confirm that MCA-Net is an effective and efficient tool for detecting anomalies in metal parts.
[0118] at last, Figure 5 The visual results in highlight the superiority of the proposed MCA-Net. Figure 5As shown, some methods, such as GhostNet (h), EfficientNet-b4 (d), and DenseNet161 (i), cannot detect abnormal areas. Other methods, including RMT (b), MetaFormer (c), and BiFormer (e), can identify small anomalies in complex backgrounds, but they often miss similar anomalies. In contrast, the proposed MCA-Net (j) can accurately detect small and similar metal anomalies on the surface. These results show that MCA-Net outperforms other methods in anomaly detection on metal parts datasets.
[0119] In this embodiment of the present invention, an ablation experiment is also performed on the Manhattan preservation attention module and the structural context attention module in the above-mentioned MCA-Net model.
[0120] In this embodiment, the Manhattan Retention Attention (MRA) module is designed to detect small abnormal regions by modeling global information. In order to evaluate the impact of the MRA module, several experiments were conducted on the part dataset, and the results are shown in Table 3;
[0121] Table 3: Performance of Manhattan Preservation Attention Module on Different Datasets
[0122] Serial number Baseline MRA SCA Accuracy (%) Accuracy (%) recall(%) F1(%) 1 √ 87.25 87.25 87.13 87.19 2 √ √ 90.68 90.68 90.64 90.66 3 √ √ 90.19 90.19 90.44 90.31 4 √ √ √ 91.76 91.76 91.69 91.35
[0123] Among them, No. 1 represents the performance of the baseline model. As shown in No. 2, after incorporating the MRA module into the baseline model, the accuracy is improved by 3.43%.
[0124] The precision is 3.51%, the recall is 3.51%, and the F1 score is 3.47%, indicating that the MRA module enhances the ability to capture metal surface anomalies. In addition, the results of No. 3 and No. 4 show that combining the MRA and SCA modules outperforms the MRA module alone in all indicators, with an increase of 1.57%, 1.57%, 1.25%, and 1.04% in accuracy, precision, recall, and F1 score, respectively. This shows that the MRA module is a valuable component for metal anomaly detection.
[0125] In order to further verify the effectiveness of the MRA module in capturing small anomalies, we Figure 6 (b) and (c) show the quantitative segmentation results of the baseline model and the model with the MRA module. These results show that the MRA module enhances the ability of the baseline model to detect metal anomalies, especially in the case of small anomalies. Figure 6 In (b) and (c) of the first row, we can see that the baseline model missed some small metal anomalies and generated incorrect attention. In contrast, the MRA module effectively detected these small anomalies and provided accurate predictions, such as Figure 6 (c) This confirms that the MRA module is able to detect small anomalies by modeling global information. Figure 6 In (d) and (e), we observe that the MRA module is able to shift attention to small defects, while the baseline model with the SCA module has difficulty detecting anomalies in complex backgrounds. These results show that the MRA module effectively exploits the global information on the feature map to more accurately locate small anomalies in metal parts.
[0126] In this embodiment, the structural context attention module (SCA) aims to distinguish abnormal areas from normal areas by aggregating contextual structural dependencies. In order to explore the effect of the SCA module, several experiments were conducted on some datasets, and the results are listed in Table 3. It is worth noting that compared with the SCA (No.3) of the baseline (No.1), the performance of accuracy, precision, recall and F1 score is improved to 90.19%, 90.19%, 90.44% and 90.31%, respectively. Similarly, from No.2 and No.4 in Table 3, it can be seen that the proposed SCA module improves the accuracy, precision, recall and F1 score by 1.08%, 1.08%, 1.05% and 0.69%, respectively, which verifies the effectiveness of the proposed SCA module in metal anomaly detection.
[0127] In addition, several class attention maps of the SCA module are Figure 6 Specifically, the comparison Figure 6 In the last four rows (b) and (d), we can see that the network is able to further focus on the abnormal regions rather than other normal regions. Figure 6 As shown in (c) and (e) in Figure 3, the attention map is transferred from the background area to the anomaly area. These excellent visualization results provide an interpretable explanation for the effectiveness of the proposed SCA module and further improve the performance of MCA-Net in metal anomaly detection.
[0128] The present invention uses specific embodiments to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
[0129] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific variations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed by the present invention, and these variations and combinations are still within the protection scope of the present invention.
Claims
1. A method for classifying abnormal metal parts based on Manhattan-related attention network, characterized in that: The following steps are involved: S1, build a metal parts dataset; S2. Construct a Manhattan-related attention network and train it using the metal parts dataset to obtain a metal parts anomaly classification network. The Manhattan-related attention network includes an input module, a residual module, a Manhattan attention retention module, a structural context attention module, a feedforward neural network module and an output module connected in sequence; S3. Use the metal parts anomaly classification network to identify the metal parts image to be detected and obtain the anomaly classification result.
2. The method for classifying metal parts anomalies based on Manhattan-correlated attention network according to claim 1 is characterized in that: In the Manhattan-correlated attention network, the input module is used to receive an input metal parts image, which includes a convolutional layer, a batch normalization layer, a ReLU activation function layer and a maximum pooling layer connected in sequence.
3. The method for classifying metal parts anomalies based on Manhattan-correlated attention network according to claim 1, characterized in that: In the Manhattan-related attention network, the residual module is used to extract the feature image of the input image, which includes four ResBlock residual blocks connected in sequence.
4. The method for classifying metal parts anomalies based on Manhattan-correlated attention network according to claim 1, characterized in that: In the Manhattan-related attention network, the Manhattan attention retention module is used to analyze the spatial relationship of input features and retain the temporal or spatial contextual relationship, and then search for small abnormal areas through global information modeling.
5. The method for classifying metal parts anomalies based on Manhattan-correlated attention network according to claim 4 is characterized in that: The method for processing input features by the Manhattan attention retention module is specifically as follows: S21, generate the query q, key k and value v of the input feature through a linear layer respectively; S22, based on the calculated query q and key k, calculate the corresponding Manhattan distance; S23, calculating the attention weight based on the Manhattan distance product of the query q and the key k; S24, combining the attention weight and value v, and calculating the output of each position by weighted sum; S25. Introduce a memory retention mechanism to integrate the memory information obtained based on the value v into the current output to obtain an output feature map containing small abnormal areas.
6. The method for classifying metal parts anomalies based on Manhattan-correlated attention network according to claim 5, characterized in that: In step S22, the Manhattan distances corresponding to query q and key k are: d manhattan (q)=q·cos(i)+q·sin(i); d manhattan (k)=k·cos(i)+k·sin(i); Where, d manhattan (q) represents the Manhattan distance of query q, i represents the index coordinate vector, q = {q1, q2, ..., q d } and k={k1,k2,...,k d } denote the vector representation of query and key respectively, i denotes the index coordinate vector, and s is the dimension of the sequence; In step S23, the attention weight α of the i-th position i for: Where λ represents the scaling factor, f i represents the Manhattan distance product of the i-th position, represents the Manhattan distance product of the j-th position; In step S24, the output y of each position j ' is expressed as: In the formula, v j represents the jth position value v, α j represents the attention weight of the jth position; In step S25, the output feature value y of the small abnormal area is included j It is expressed as: In the formula, β represents the learning parameter, m i Represents the memory information from the previous time step.
7. The method for classifying metal parts anomalies based on Manhattan-correlated attention network according to claim 1, characterized in that: In the Manhattan-correlated attention network, the structural context attention module is used to distinguish subtle abnormal areas from normal areas by modeling local features and spatial dependencies, and aggregate context information to determine the position of the area in the overall structure, thereby obtaining a more refined abnormal area detection feature map.
8. The method for classifying metal parts anomalies based on Manhattan-correlated attention network according to claim 7, characterized in that: The structural context attention module processes the input feature map as follows: S26, determining the input feature map x and the feature map set X to which it belongs; Among them, the feature map set includes the input feature map x and its adjacent feature maps {x1, x2, ..., x8}; S27, respectively generating a query feature map, a key feature map, and a value feature map of the input feature map through a convolutional layer, and generating context information based on the feature map set; S28, introducing a context-aware mechanism, and performing weighted calculations on the query feature graph, the key feature graph, and the value feature graph respectively in combination with the context information; S29, calculating context-enhanced attention based on the context-weighted query feature graph, the key feature graph, and the value feature graph; S210. According to the enhanced attention described below, a more refined abnormal region detection feature map is output through a linear layer.
9. The method for classifying metal parts anomalies based on Manhattan-correlated attention network according to claim 8, characterized in that: In the step S27, the context information is obtained by performing a convolution operation on the corresponding feature map set; In step S28, the context-weighted query feature graph, key feature graph, and value feature graph are respectively expressed as: q'=q+r; k'=k+s; v'=v+t; Where r' = W r r,k'=W k k,v'=W v v,W r , W k and W v Represent the learnable parameter matrices corresponding to the query feature map, key feature map, and value feature map respectively; In step S29, the contextual enhanced attention ContextualAttenion(q,k,v) is expressed as: Where, d k represents the kth scaling factor, and T represents the transposition symbol.