Aero-engine blade defect detection method based on multi-view learning
The multi-view learning method enhances aircraft engine blade defect detection by integrating multiple angles' information, addressing high computational costs and improving accuracy through cross-view feature interaction.
Patent Information
- Application Number
- CN202510776273.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-15
AI Technical Summary
The existing aero engine blade defect detection methods rely on single-view data, which are prone to imaging blind spots, difficult to capture subtle defects, and are costly to train and test models and have low detection accuracy.
The multi-view learning method is adopted, by obtaining the aero engine blade images from different perspectives, using the ViT feature extractor and multi-layer perception bottleneck layer for feature compression and dimensionality reduction, combining the multi-view feature decoder and the cross-view feature learning module, the loss function is calculated to generate defect positioning results.
It improves the accuracy of aircraft engine blade defect detection, effectively integrates multi-view information, captures surface details, reduces calculation costs, and improves detection efficiency.
Smart Images

Figure CN120318222A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of aeroengine blade defect detection, and particularly to a method for aeroengine blade defect detection based on multi-view learning. Background Art
[0002] Currently, aeroengine blade defect detection methods are mainly based on the unsupervised paradigm, which learns the defect distribution pattern from defect-free images and locates the defect area. This method reduces the difficulty of collecting and labeling aeroengine blade detection images in complex industrial environments. However, existing methods are generally set for single-category defects, and dedicated models need to be trained and tested for each defect category, resulting in a significant increase in time cost and memory overhead. On the other hand, existing methods usually rely on single-view data, which is prone to imaging blind spots, and single data sources are more likely to make it difficult to distinguish subtle defects, unable to capture the subtle and complex defect features on the surface of aeroengine blades, directly affecting the accuracy of detection. Therefore, how to improve the accuracy of aeroengine blade defect detection while ensuring low computational cost is a problem that urgently needs to be solved. Summary of the Invention
[0003] Based on this, it is necessary to provide a method for aeroengine blade defect detection based on multi-view learning, which includes: S1: Obtain surface images of several aeroengine blades from different perspectives, take the surface images of the same aeroengine blade from different perspectives as a group, and divide them into multiple groups of surface images; S2: Input any group of surface images into a pre-trained ViT feature extractor, take any one perspective as the main perspective, and extract different-level semantic features of the surface images from each perspective; calculate the semantic features between the image features in the main perspective and the image features in other perspectives; S3: Use a multi-layer perceptron bottleneck layer with perturbation to compress and reduce the dimension of different-level semantic features to obtain compressed and dimension-reduced image features; input the image features into a multi-view feature decoder to calculate the multi-view features between the image features in the main perspective and the image features in other perspectives; S4: Calculate a loss function based on different-level semantic features and multi-view features, and minimize the loss function to enable the defect detection model to learn the feature space and its distribution characteristics of normal images of aeroengine blades, and generate defect location results.
[0004] Preferably, the process of obtaining multi-view features from image features through a multi-view feature decoder includes: In any set of surface images, taking any perspective as the main perspective, calculate the attention score between the image features under the main perspective and the image features under other perspectives, and perform normalization on the attention score to obtain the weight coefficient; filter the weight coefficient based on a preset retention ratio to obtain the correlation weight; Based on the correlation weight and the value matrix in the self-attention mechanism, calculate the enhanced feature; enhance the image features under the main perspective based on the enhanced feature to obtain the fused feature; construct the multi-view feature based on the fused feature and the image features under the main perspective.
[0005] Preferably, calculating the attention score between the image features under the main perspective and the image features under other perspectives includes: Multiply the image features under the main perspective by the query weight matrix to obtain the query matrix; Multiply the image features under other perspectives by the key weight matrix to obtain the key matrix; Multiply the query matrix by the transposed matrix of the key matrix to obtain the first product; Multiply the L1 norm of the query matrix by the L1 norm of the key matrix to obtain the second product; Divide the first product by the second product to obtain the attention score between the image features under the main perspective and the image features under other perspectives.
[0006] Preferably, performing normalization on the attention score to obtain the weight coefficient includes: In any set of surface images, based on the attention score between the image features under any main perspective and the image features under other perspectives, and the attention scores corresponding to all main perspectives, use the softmax function to perform normalization on the attention score corresponding to this main perspective to obtain the weight coefficient.
[0007] Preferably, filtering the weight coefficient based on a preset retention ratio to obtain the correlation weight includes: Among all the attention scores of any set of surface images, retain the top preset retention ratio of attention scores, and use the weight coefficients corresponding to the retained attention scores as the correlation weights corresponding to the corresponding main perspectives; Assign the weight coefficients corresponding to the remaining attention scores to zero, and use zero as the correlation weight corresponding to the corresponding main perspective.
[0008] Preferably, calculating the enhanced feature based on the correlation weight and the value matrix in the self-attention mechanism includes: Multiply the image features under other perspectives by the value weight matrix to obtain the value matrix; In any set of surface images, sum the products of the correlation weights corresponding to all main perspectives and the value matrix to obtain the enhanced feature.
[0009] Preferably, based on the fusion features and the image features from the main perspective, constructing multi-view features includes: Define a training parameter in any set of surface images; Pass the training parameter through the sigmoid function to obtain the first coefficient; Multiply the first coefficient by the fusion features to obtain the first result; Pass the image features from the main perspective through the self-attention mechanism and multiply them by the balance number of the first coefficient with respect to 1 to obtain the second result; Concatenate the first result and the second result to obtain the multi-view features.
[0010] Preferably, the loss function is denoted as: ; where L represents the loss function; m represents the number of image features in each group; represents the top p % of the features selected by cosine similarity sorting; represents the cosine similarity; represents the semantic features output by the pre-trained ViT feature extractor for the i-th image; represents the multi-view features reconstructed by the multi-view feature decoder for the i-th image feature; represents the scaling function; represents the top p % of the features selected by cosine similarity sorting, and represents the cosine similarity of the i-th feature; represents the scaling factor.
[0011] Beneficial effects: This method effectively integrates multi-angle visual information, captures richer surface details of aero-engine blades, and improves the accuracy of aero-engine blade defect detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0013] Figure 1 FIG. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0014] To make the above objects, features, and advantages of the present application more obvious and understandable, the following will describe the specific embodiments of the present application in detail with reference to the accompanying drawings. Many specific details are set forth in the following description to fully understand the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present application. Therefore, the present application is not limited by the specific embodiments disclosed below.
[0015] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present application, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0016] As Figure 1 shown, this embodiment provides a method for detecting defects in aero-engine blades based on multi-view learning. S1: Obtain surface images of several aero-engine blades from 5 different perspectives (resolution: 448×448 pixels), and take the surface images of the same aero-engine blade from different perspectives as a group to divide multiple groups of surface images.
[0017] This step further includes: constructing a multi-view learning-based aero-engine blade defect detection model, mainly including a pre-trained ViT feature extractor, a perturbed multi-layer perceptron bottleneck layer, and a multi-view feature decoder with multi-view feature learning capabilities.
[0018] In this embodiment, the feature extractor uses the ViT model, which can extract multi-dimensional features in the aero-engine blade image and output semantic features at different levels for further processing by the multi-view feature decoder; in order to effectively aggregate the outputs of each layer of the ViT feature extractor, this embodiment introduces a perturbed multi-layer perceptron bottleneck layer for feature compression and dimensionality reduction, reducing the computational complexity and extracting more compact core features for defect detection. At the same time, the Dropout block is used to prevent the model from overfitting, enabling it to have stronger generalization ability when facing a large number of aero-engine blade images; the multi-view feature decoder integrates a trainable cross-view feature learning module for decoding the features output by the multi-layer perceptron bottleneck layer. This module can enhance the robustness of the model in complex scenarios by fusing information from different views, and finally reconstruct the multi-view features to accurately detect and locate the defects in the aero-engine blades and generate defect location results.
[0019] By collecting aero-engine blade images from different perspectives and grouping all the collected results, with each group containing images of the same aero-engine blade from different perspectives, it is ensured that a structured multi-view relationship prior is provided for the cross-view feature learning module during model training.
[0020] S2: Input any group of surface images into the pre-trained ViT feature extractor. Taking any one perspective as the main perspective, extract the semantic features of different levels of the surface images from each perspective; calculate the semantic features between the image features in the main perspective and the image features in other perspectives.
[0021] In this embodiment, the ViT-B / 32 version of the model is selected as the feature extractor and pre-trained on a large-scale dataset through the DINO self-supervised learning method. The embedding vector dimension of this ViT feature extractor is set to 768, which is responsible for extracting multi-dimensional feature representations from the input aero-engine blade surface images, and passing the image features extracted from its 2nd, 3rd, 4th, 5th, 6th, 7th, 8th, and 9th layers to the multi-view feature decoder for further processing.
[0022] The working process of the feature extractor is as follows: The input surface image is first segmented into non-overlapping image patches of a fixed size. Each image patch is linearly projected into a high-dimensional embedding vector respectively, and a learnable position encoding is superimposed to retain spatial information; subsequently, the encoded sequence undergoes global feature interaction through multiple layers of Transformer encoders, where each layer contains a multi-head self-attention mechanism and a feed-forward neural network, and outputs the features of different semantic levels solved by each layer.
[0023] S3: Use a multi-layer perceptron bottleneck layer with perturbations to compress and reduce the dimension of the semantic features of different levels to obtain the compressed and dimension-reduced image features; input the image features into the multi-view feature decoder and calculate the multi-view features between the image features in the main perspective and the image features in other perspectives.
[0024] Based on the semantic features output by the ViT feature extractor, the multi-layer perceptron bottleneck layer gradually compresses the dimension through two fully-connected networks. Each layer sequentially performs linear transformation, GELU non-linear activation, and normalization operations, so as to enhance the discriminability of the features while reducing redundancy, and finally outputs a low-dimensional and compact representation, that is, the compressed and dimension-reduced image features.
[0025] To effectively learn multi-view features and improve the defect detection performance, this embodiment provides a cross-view feature learning module, enabling the main view to focus on the highly relevant regions of other auxiliary views, effectively utilizing multi-view information to achieve more accurate and comprehensive feature learning. Specifically, the process of obtaining multi-view features from the image features through the multi-view feature decoder includes: To effectively establish the associative interaction relationship between the main view features and the auxiliary view features, in any set of surface images, taking any perspective as the main perspective, calculate the attention scores between the image features under the main perspective and the image features under other perspectives, specifically: Multiply the image features under the main perspective by the query weight matrix to obtain the query matrix; Multiply the image features under other perspectives by the key weight matrix to obtain the key matrix; Multiply the query matrix by the transposed matrix of the key matrix to obtain the first product; Multiply the L1 norm of the query matrix by the L1 norm of the key matrix to obtain the second product; Divide the first product by the second product to obtain the attention scores between the image features under the main perspective and the image features under other perspectives.
[0026] To ensure the effectiveness and consistency of the attention mechanism between different views, normalize the attention scores to obtain the weight coefficients, specifically: In any set of surface images, based on the attention scores between the image features under any main perspective and the image features under other perspectives, as well as the attention scores corresponding to all main perspectives, use the softmax function to normalize the attention scores corresponding to the main perspective to obtain the weight coefficients.
[0027] Through the above processing, it is ensured that in the case of taking different perspectives as the main perspective for the same object, the model can adaptively adjust the weights, so as to accurately extract the features most relevant to the main perspective image in all auxiliary perspective images, and realize cross-view information sharing and fusion.
[0028] Then, this embodiment proposes a focus-enhanced attention mechanism. To effectively learn the most relevant features and suppress irrelevant features, the focus-enhanced attention mechanism focuses on retaining the most significant part of the weights, and filters the weight coefficients based on a preset retention ratio to obtain the correlation weights, specifically: Among all the attention scores of any set of surface images, retain the top preset retention ratio (in this embodiment, the preset retention ratio is 30%) of the attention scores, and use the weight coefficients corresponding to the retained attention scores as the correlation weights under the corresponding main perspective; Assign the weight coefficients corresponding to the remaining attention scores to zero, and use zero as the correlation weight under the corresponding main perspective.
[0029] The focus-enhanced attention mechanism ensures that in multi-view feature learning, the model only focuses on the most important features, thereby improving the learning efficiency and reducing the dependence on redundant information.
[0030] Based on the correlation weights and the value matrix in the self-attention mechanism, enhanced features are calculated as follows: Multiply the image features from other perspectives by the value weight matrix to obtain the value matrix; In any set of surface images, sum the products of the correlation weights corresponding to all principal perspectives and the value matrix to obtain the enhanced features.
[0031] Finally, perform cross-view adaptive feature enhancement. According to the correlation weights between the principal view and the auxiliary views, the cross-view feature learning module can effectively focus on the feature information of the auxiliary views most relevant to the current principal view, thereby obtaining richer context information. Enhance the image features in the principal view based on the enhanced features to obtain the fused features; based on the fused features and the image features in the principal view, construct multi-view features as follows: Define a training parameter in any set of surface images; Pass the training parameter through the sigmoid function to obtain the first coefficient; Multiply the first coefficient by the fused features to obtain the first result; Pass the image features in the principal view through the self-attention mechanism and multiply by the balance number of the first coefficient with respect to 1 to obtain the second result; Concatenate the first result and the second result to obtain the multi-view features.
[0032] This process ensures that the model can adaptively adjust the features according to the relationships between different views, thereby optimizing the effect of multi-view feature learning. In this embodiment, the cross-view feature learning module runs multiple times, and the number of times is the same as the number of layers of the ViT feature extractor, to output the multi-view features corresponding to the corresponding layers.
[0033] In this example, the multi-layer perceptron bottleneck layer with perturbations compresses and reduces the dimension of the 8-layer semantic features output by the ViT feature extractor by taking the mean value, in order to reduce the computational complexity and extract a more compact core feature representation for defect detection. At the same time, a Dropout block with a dropout rate of 0.3 is used to prevent overfitting. The multi-view feature decoder in this embodiment integrates a trainable cross-view feature learning module to decode the image features output by the multi-layer perceptron bottleneck layer. By fusing different view features, this module enhances the robustness of the model in complex scenarios, can more accurately reconstruct the final multi-view features, and further improves the ability to detect and locate defects in aero-engine blades, and generates defect location results.
[0034] S4: Calculate the loss function based on the semantic features and multi-view features at different layers. Minimize the loss function to enable the defect detection model to learn the feature space and its distribution characteristics of the normal images of aero-engine blades, and generate defect location results.
[0035] In this embodiment, the loss function is denoted as: ; where L represents the loss function; m represents the number of image features in each group (m = 5 in this embodiment); represents the top p % of the feature quantities selected by cosine similarity ranking ( p = 90 in this embodiment); represents the cosine similarity; represents the semantic feature output by the pre-trained ViT feature extractor for the i-th image feature; represents the multi-view feature reconstructed by the multi-view feature decoder corresponding to the i-th image feature; represents a scaling function used to adjust the feature weights to amplify the contribution ratio of the top p % features to the loss function; represents the cosine similarity of the i-th feature among the top p % features selected by cosine similarity ranking. The higher the cosine similarity of a feature, the larger it is, and the greater the weight of this feature in the loss function; represents a scaling factor, which is adjustable and used to control the overall intensity of feature weight amplification.
[0036] The goal of this loss function is to minimize the cosine similarity between the corresponding feature vectors in the output of the ViT feature extractor and the output of the multi-view feature decoder, and dynamically select the detection area by identifying the most relevant features ranked by feature similarity, so as to accurately represent the features of aero-engine blades.
[0037] In this embodiment, an unsupervised learning method is adopted. Only the normal aero-engine blade surface images are used to train the model. The total number of training iterations is 50,000 times. The number of image groups in a single training is 4 groups, the batch size is 20, and the training device is a single NVIDIA RTX-A6000 GPU. Based on the semantic features of the specified layer obtained by the ViT feature extractor and the corresponding layer multi-view features reconstructed by the multi-view feature decoder, the loss function is jointly calculated to gradually learn the feature space and its distribution characteristics of normal aero-engine blade images, so as to obtain the ability to identify abnormal images deviating from this pattern. After the model training is completed, the defect detection model can be applied to the aero-engine blade production line for defect detection. When detecting the surface images of aero-engine blades with defects, since the model has not encountered defect features during training, the reconstruction error of the abnormal area increases significantly, thus realizing the accurate detection of defects and the positioning of the defect area.
[0038] The aero-engine blade defect detection method based on multi-view learning provided in this embodiment has the following beneficial effects: This method constructs an aero-engine blade defect detection model that includes a pre-trained ViT feature extractor, a multi-layer perceptron bottleneck layer with perturbations, and a multi-view feature decoder, overcoming the imaging blind spot problem in single-view methods. By combining the advantages of multi-view learning and unsupervised defect detection, it captures richer defect details. Through the proposed cross-view feature learning module, it can effectively learn information from different views and the correlation interaction relationships between views, thereby guiding the main view to focus on the highly relevant regions in the auxiliary views. After enhancing weak defect features, it significantly improves the accuracy of aero-engine blade defect detection. Finally, by combining unsupervised learning for model training and calculating the reconstruction loss between the ViT feature extractor and the multi-view feature decoder, it achieves accurate detection and localization of aero-engine blade defects. This method effectively integrates multi-view feature information, can capture richer details on the surface of aero-engine blades, and realizes accurate and efficient detection of aero-engine blade defects.
[0039] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0040] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A method for detecting defects of aero-engine blades based on multi-view learning, characterized in that, Including: S1: Obtain surface images of several aeroengine blades from different perspectives. Take the surface images of the same aeroengine blade from different perspectives as a group, and divide them into multiple groups of surface images. S2: Input any group of surface images into a pre-trained ViT feature extractor. Take any one perspective as the main perspective, and extract different-level semantic features of the surface images from each perspective. Calculate the semantic features between the image features in the main perspective and the image features in other perspectives. S3: Use a multi-layer perceptron bottleneck layer with perturbations to compress and reduce the dimension of the different-level semantic features to obtain the compressed and dimension-reduced image features. Input the image features into a multi-view feature decoder to calculate the multi-view features between the image features in the main perspective and the image features in other perspectives. S4: Calculate the loss function based on the different-level semantic features and multi-view features. Minimize the loss function to enable the defect detection model to learn the feature space and its distribution characteristics of the normal images of aeroengine blades, and generate defect localization results.
2. The method for detecting defects of aero-engine blades based on multi-view learning according to claim 1, characterized in that The process of obtaining multi-view features from image features through a multi-view feature decoder includes: In any group of surface images, take any one perspective as the main perspective, calculate the attention scores between the image features in the main perspective and the image features in other perspectives, and perform normalization processing on the attention scores to obtain weight coefficients. Filter the weight coefficients based on a preset retention ratio to obtain correlation weights. Calculate enhanced features based on the correlation weights and the value matrix in the self-attention mechanism. Enhance the image features in the main perspective based on the enhanced features to obtain fused features. Construct multi-view features based on the fused features and the image features in the main perspective.
3. The method for detecting defects of aero-engine blades based on multi-view learning according to claim 2, wherein, Calculating the attention scores between the image features in the main perspective and the image features in other perspectives includes: Multiply the image features in the main perspective by the query weight matrix to obtain a query matrix. Multiply the image features in other perspectives by the key weight matrix to obtain a key matrix. Multiply the query matrix by the transpose matrix of the key matrix to obtain a first product. Multiply the L1 norm of the query matrix by the L1 norm of the key matrix to obtain a second product. Divide the first product by the second product to obtain the attention scores between the image features in the main perspective and the image features in other perspectives.
4. The method for detecting defects of aero-engine blades based on multi-view learning according to claim 2, wherein, Performing normalization processing on the attention scores to obtain weight coefficients includes: In any group of surface images, based on the attention scores between the image features in any one main perspective and the image features in other perspectives, and the attention scores corresponding to all main perspectives, use the softmax function to perform normalization processing on the attention scores corresponding to this main perspective to obtain weight coefficients.
5. The method for detecting defects of aeroengine blades based on multi-view learning according to claim 2, wherein, Filtering the weight coefficients based on a preset retention ratio to obtain correlation weights includes: Among all the attention scores of any group of surface images, retain the top preset retention ratio of attention scores, and use the weight coefficients corresponding to the retained attention scores as the correlation weights corresponding to the corresponding main perspective. Assign the weight coefficients corresponding to the remaining attention scores to zero, and use zero as the correlation weights corresponding to the corresponding main perspective.
6. The method for detecting defects of aero-engine blades based on multi-view learning according to claim 2, wherein Calculating enhanced features based on the correlation weights and the value matrix in the self-attention mechanism includes: Multiply the image features from other perspectives by the value weight matrix to obtain a value matrix; In any set of surface images, sum the products of the correlation weights corresponding to all principal perspectives and the value matrix to obtain enhanced features.
7. The method for detecting defects of aero-engine blades based on multi-view learning according to claim 2, characterized in that Based on the fused features and the image features from the principal perspective, construct multi-view features including: In any set of surface images, define a training parameter; Pass the training parameter through the sigmoid function to obtain a first coefficient; Multiply the first coefficient by the fused features to obtain a first result; Pass the image features from the principal perspective through the self-attention mechanism and multiply by the balance number of the first coefficient with respect to 1 to obtain a second result; Concatenate the first result and the second result to obtain multi-view features.
8. The method for detecting defects of aero-engine blades based on multi-view learning according to claim 1, characterized in that The loss function is denoted as: ; Among them, L represents the loss function; m represents the number of image features in each group; represents the top p % of the number of features selected by cosine similarity ranking; represents the cosine similarity; represents the semantic feature output by the pre-trained ViT feature extractor for the i-th image; represents the multi-view feature reconstructed by the multi-view feature decoder for the i-th image feature; represents the scaling function; represents the top p % of the cosine similarity of the i-th feature among the selected features; represents the scaling factor.
Citation Information
Patent Citations
Organ image segmentation method and system based on multi-scale multi-view network
CN116109822A
Aero-engine blade defect detection method based on multi-scale DETR
CN117173449A