Dual-channel decoupling and gated memory fusion method for detecting damage in borescope images
Through the hole-detection image damage detection method fusion with dual-channel decoupling and gated memory, the problems of high-frequency details loss and miss detection of damage recognition in the hole-detection image of civil aviation engines are solved, and efficient identification and accurate positioning of local microstructure damage is achieved.
Patent Information
- Application Number
- CN202510629149.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-16
AI Technical Summary
The prior art is susceptible to background noise and residual signals in the detection of hole-detection images of civil aviation engines, resulting in the loss of high-frequency details, making it difficult to effectively identify damage to local fine structures, and traditional methods are prone to missed or missed detection, which cannot fully cover the feature distribution of normal samples.
The hole-detection image damage detection method is adopted with dual-channel decoupling and gating memory fusion. Feature decoupling is performed through dual backbone encoder, gradient inversion layer and mutually exclusive losses are introduced, and feature fusion modules of adaptive gating and memory guidance are combined. The structure-aware reconstruction decoder is used for image reconstruction, and the feature fusion weight is dynamically adjusted to improve the model's sensitivity and discrimination ability to damage features.
It effectively improves the intelligence level of detection and consistency of interpretation, can accurately identify small damage in complex backgrounds, reduce false alarm rates, and improves the model's response ability and positioning accuracy to the damaged area.
Smart Images

Figure CN120147320B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of complex equipment fault identification and processing, and specifically to a method for detecting damage to borescope images by dual-channel decoupling and gated memory fusion, which can effectively improve the intelligent level of detection and the consistency of interpretation. Background Art
[0002] First, a reconstruction anomaly detection method based on an encoder-decoder architecture. This method constructs a mapping network for normal samples and realizes pixel-level positioning by using the reconstruction residual anomaly value of the damaged area. Such methods have achieved outstanding results in the field of industrial product surface defect detection. However, in the scenario of civil aviation engine borescope inspection, due to the influence of image-irrelevant information such as background noise and residual signals during the reconstruction of the internal blade texture of civil aviation engines, high-frequency details are easily lost, resulting in confusion between the residual anomaly value of the damaged area and the normal texture fluctuation, thus causing missed detection or false detection. Second, an anomaly detection method based on adversarial training. This method establishes an implicit distribution of normal samples through adversarial training, and the discriminant network is highly sensitive to the pattern deviation of the damaged area. These methods also have extensive applications in the field of industrial product surface defect detection. However, for the damage identification of civil aviation engines, due to the characteristics of small scale and irregular shape of the damage features of civil aviation engines, the damage identification of engine borescope images often relies on local fine structure information. If the form of adversarial training is adopted, it is easy to have a mode collapse problem, that is, the generator tends to learn global features and ignore local details, making it difficult to effectively capture and reconstruct the fine damage details existing in the borescope images, and also unable to fully cover the feature distribution of normal samples, thus causing an increase in the false alarm rate. Third, an anomaly detection method based on contrastive learning. The feature decoupling framework of this method constructs positive and negative sample pairs through data augmentation to amplify the distinction between the damaged area and the normal texture in the feature space. Such methods rely too much on data augmentation strategies. At the same time, traditional contrastive learning assumes that positive and negative samples have symmetric or alignable feature distributions. However, for civil aviation engine borescope images, the geometric and texture characteristics of the damaged area and the normal structure are significantly different, and this asymmetry may lead to the loss of key discriminant information during the feature decoupling process, making it difficult for the model to effectively separate damaged and normal features.. Summary of the Invention
[0003] Aiming at the shortcomings and deficiencies existing in the prior art, the present invention proposes a method for detecting damage to borescope images by dual-channel decoupling and gated memory fusion, which can break through the category limitations of labeled samples, realize open recognition of unknown anomaly types, and effectively improve the intelligent level of detection and the consistency of interpretation.
[0004] The present invention is achieved by the following measures:
[0005] A method for detecting damage in borescope images with dual-channel decoupling and gated memory fusion, characterized by establishing a borescope image damage detection model with dual-channel decoupling and gated memory fusion, and training only with unlabeled normal images. The establishment of the borescope image damage detection model with dual-channel decoupling and gated memory fusion includes feature decoupling, feature fusion, and feature reconstruction. In the feature decoupling stage, a dual-backbone encoder is used, serving as an effective information encoding branch and a redundant information encoding branch respectively;
[0006] In the feature fusion stage, there is a feature fusion module with adaptive gating and memory guidance. Among them, the gating mechanism adaptively adjusts the feature fusion weights of the dual-backbone encoding branches according to the similarity between the features extracted from the current input in the effective information encoding branch and the features of normal images stored in the memory bank. The feature content in the memory bank will be dynamically updated according to the training progress of the effective information encoder to adapt to the continuous evolution of the model's representation ability;
[0007] In the feature reconstruction stage, a structure-aware reconstruction decoder is established. The fused representation features are sent to the structure-aware reconstruction decoder for image reconstruction, and the structure-aware reconstruction decoder uses a jointly optimized loss function.
[0008] In the feature decoupling stage during the establishment process of the borescope image damage detection model with dual-channel decoupling and gated memory fusion of the present invention, the effective information encoding branch is used to extract effective features with high semantic consistency in the image, and the redundant information encoding branch, by introducing a gradient reversal layer, is used to extract redundant or non-discriminative features that are mutually exclusive with the backbone information in the image. A mutual exclusion loss is also introduced. By minimizing the correlation between the output features of different levels of the two encoding branches, the process of feature decoupling is promoted, ensuring that the effective information branch focuses on learning task-related features, while the redundant information branch focuses on capturing features irrelevant to anomaly detection.
[0009] The present invention uses DenseNet-121 as the backbone network of the dual-backbone encoder. The overall structure of DenseNet-121 includes an initial convolutional layer, dense blocks, and transition layers. Among them, in the initial stage of the network, the initial convolutional layer is used to extract basic features of the input image. Subsequently, the obtained initial features are connected to the dense block part. The dense block promotes feature reuse by concatenating the outputs of all previous layers in the channel dimension as the input of the current layer to alleviate the problem of gradient disappearance, as shown in the following formula:
[0010] (1),
[0011] In the formula, Consists of Batch Normalization, ReLU, and convolution;
[0012] Next, a transition layer is designed between each dense block and the next dense block to downsample the feature map and compress the channels. 1×1 convolution and average pooling operations are used to reduce the spatial size and the number of feature maps, so that the computational cost and memory occupancy of the entire network can be effectively controlled.
[0013] In the dual backbone encoder with feature decoupling of the present invention, the effective information encoding backbone adopts the standard DenseNet-121 architecture. Let the input image be , and the feature representation extracted by the effective information encoding backbone is:
[0014] (1), where represents the input image, represents the effective information encoder constructed based on the DenseNet-121 architecture, represents the output effective feature map, , represents the height, width, and number of channels of the encoded feature map;
[0015] The redundant information encoding backbone is also based on DenseNet-121, but a gradient reversal layer GRL is introduced at the output end of the backbone. GRL is an identity mapping during forward propagation, while it reverses the gradient direction during backpropagation, thereby guiding this backbone to learn features mutually exclusive with the effective information:
[0016] (3),
[0017] where represents the output redundant feature map, , represents the redundant information encoder constructed based on the DenseNet-121 architecture, represents the gradient reversal layer, and the hyperparameter controlling the gradient reversal intensity is , where the definition of the gradient reversal layer is as follows, where represents the identity matrix, indicating that the gradient is multiplied by during backpropagation to achieve the reversal of the training direction of the redundant branch.
[0018] (4), (5), For the setting of the gradient reversal intensity hyperparameter , a dynamic adjustment strategy is adopted to balance the stability and adversarial nature in the learning process, as shown in the following formula:
[0019] (6), where represents the training progress, .
[0020] In the present invention, the adaptive gating and memory-guided feature fusion module in the feature fusion part dynamically integrates the features output by different backbone encoders during the reconstruction process to achieve differential modeling of different types of inputs. The goal of the feature fusion part is to perform weighted fusion on the features extracted by the dual backbone encoders of the feature decoupling module at each spatial position The fused features are expressed as:
[0021] (7),
[0022] In the formula, represents the fusion weight map generated by the gating mechanism, which is used to control the proportion of the main branch features at each spatial position, , in order to enable the gating weight to be dynamically adjusted according to the standard of normal samples, a memory bank module is introduced. The adaptive gating and memory-guided feature fusion module is used to establish a memory bank and a gating attention mechanism. The construction of the memory bank dynamically adjusts the number of features in the memory bank through the long-term change trend of the reconstruction error, so as to more comprehensively capture the diversity of normal samples and improve the judgment accuracy of the gating mechanism, including the following content: First, at the th training stage, record the current reconstruction loss value as , and use the exponential moving average (EMA) method to record the trend of the smoothed loss:
[0023] (8),
[0024] In the formula, represents the smoothed average loss at the th moment, represents the smoothing parameter, with a value of 0.9,
[0025] Next, let the size of the initial memory bank be , and the current number of features in the memory bank be . The update goal is to gradually and dynamically update the number of features in the memory bank when the model training gradually stabilizes and the loss significantly decreases:
[0026] (9), in the formula, represents the smoothed average loss at the th moment, represents the smoothed average loss at the -1th moment, represents the Sigmoid function, which is used to normalize the difference to the [0,1] interval, represents the growth rate, ,
[0027] (10),
[0028] wherein, is used to control the slope of the Sigmoid function, i.e., the sensitivity, , is used to control the midpoint position of the function (determining the starting point of the memory bank growth), which is set to 0.5 here to trigger the expansion of the memory bank scale when the loss starts to significantly decrease; among them, the number of prototypes stored in the memory bank in the initial stage needs to ensure that too much noise or unstable features will not be written into the memory bank during the initial training, so a relatively moderate scale is selected, with the value K 0 = 128, which can cover sufficient sample diversity and avoid introducing unstable features prematurely,
[0029] Finally, the memory bank is regarded as a dynamically updated feature set:
[0030] (11), wherein, represents the number of prototypes in the memory bank at a training stage, represents the typical feature of a certain type of valid information in the normal samples;
[0031] After each training, the image features output by each valid information encoder are sent to the decoder one by one for reconstruction. Subsequently, according to the current available capacity of the memory bank, several features with the smallest reconstruction error are selected from the current batch for update. Through this form, it is ensured that the stored features have a high reconstruction ability, that is, a stronger expression ability for the normal mode, thereby improving the discrimination accuracy of the gating mechanism for anomalies. Specifically: First, in the current training iteration, a set of candidate features are extracted through batch training; subsequently, the reconstruction error of each feature is calculated ; finally, the first features with the smallest error (note that here will be adjusted with the number of iterations) are written into the memory bank; at the same time, to avoid frequently disturbing the stable memory content in the later stage of training, a decay coefficient is also provided to ensure that as the training process progresses, the update rate of the features in the memory bank gradually slows down. The exponential weighted average (EMA) update strategy is adopted for the update of the prototypes in the memory bank for the selected valid information features:
[0032] (12),
[0033] wherein, represents the number of prototypes in the memory bank, represents the dynamic update decay coefficient, ,
[0034] Meanwhile, a dynamic update decay coefficient is set, and the formula is: (13),
[0035] In the formula, represents the EMA decay coefficient at the -th iteration, represents the smaller value at the initial moment, taking 0.2, represents the larger value at the initial moment, taking 0.8 here, represents the hyperparameter for controlling the growth rate, taking 1e-2 here, represents the current iteration number.
[0036] The goal of the gating mechanism in the present invention is to perform dynamic weighted fusion on the features output by the dual-encoding branches to achieve adaptive fusion regulation of effective information and redundant information, including:
[0037] First, similarity calculation is performed. For the effective features at a certain spatial position in the input image and each prototype in the memory bank, the similarity is calculated. Considering that for unsupervised anomaly detection tasks, the scale information of features often does not have direct discriminative significance, and more attention is paid to the pattern (direction) rather than the absolute value. Therefore, cosine similarity is used for calculation, and its formula is as follows: (14),
[0038] In the formula, represents the typical feature of a certain type of effective information in the normal samples, represents the feature extracted by the effective encoding branch at the spatial position , represents the cosine similarity between the feature at the spatial position and the -th prototype,
[0039] Next, the similarities of all prototypes are normalized by the Softmax function to obtain the normalized matching weights of each prototype feature at this position:
[0040] (15),
[0041] In the formula, represents the normalized matching weight,
[0042] Then, the overall similarity score at the position is defined as:
[0043] (16),
[0044] To normalize the score to the interval [0, 1], a gating weight is generated through the Sigmoid function:
[0045] (17), which can be regarded as a gating coefficient reflecting the "importance" or "selective transmission" of features at the position. If is relatively high, it means that when the effective information matches well with the memory bank prototype, a higher weight is assigned; if is relatively low, there may be an anomaly, and it is necessary to enhance the influence of redundant information, and the contribution of the redundant information encoding branch is relatively enhanced, resulting in a large error in the abnormal area during reconstruction.
[0046] In the present invention, the structure-aware reconstruction decoder network in the feature reconstruction part includes an upsampling module (UpBlock) and an output module. The upsampling module is mainly used to gradually restore the spatial resolution, and the output module is to adjust the number of channels of the feature map to be the same as that of the input image to generate a reconstructed image. Among them, the upsampling module is composed of a transposed convolution (Transposed Convolution, TC), batch normalization (Batch Normalization, BN), an activation function, and a skip connection part. Let the input feature map of the th upsampling module of the decoder be , and be the height and width of the feature map respectively. After the feature map enters the upsampling module, first, its spatial size is enlarged through the transposed convolution, and the transposed feature map is obtained after the transposed convolution:
[0047] (18), where represents the rd layer, represents the transposed feature map, and its mathematical expression should be:
[0048] (19),
[0049] where represents the kernel function of the transposed convolution, represents the stride of the transposed convolution, which takes a value of 2 here, represents the spatial position in represents the spatial coordinates in
[0050] Among them, hyperparameters such as the convolution kernel size and stride need to be set according to specific tasks;
[0051] Next, for the output of the transposed convolution Perform batch normalization and activation functions. The purpose of using batch normalization is to slow down the change of feature distribution, make the input distribution change less during the training process of each layer, ensure the stability of the input distribution of each layer, and alleviate the problems of gradient disappearance or gradient explosion. For the input feature map For each channel of the mean and variance will be calculated, and then through the learnable scaling parameter and translation parameter (20),
[0052] wherein, represents the transposed feature map, represents the batch normalization operation, represents a constant of a smaller value used to prevent division by zero in the formula, and here 1e-4 is taken;
[0053] Use ReLU (Rectified Linear Unit) as the activation function to provide the network with non-linear transformation ability. Finally, the feature map after batch normalization and activation function is expressed as follows:
[0054] (21),
[0055] wherein, represents the batch normalization operation, represents the rectified linear unit activation function, which sets negative values to zero;
[0056] Then, through skip connections, the low-level features of the corresponding spatial dimensions obtained from the effective information branch backbone of the encoder need to be concatenated with the processed features in the channel dimension to obtain the concatenated features :
[0057] (22),
[0058] wherein, represents the concatenation operation in the channel dimension, represents the skip feature of the i th layer of the effective information branch, represents the number of channels of the skip connection of this layer,
[0059] Subsequently, further convolutional fusion is performed on the concatenated feature map , and after passing through batch normalization and activation functions again, the output feature of the upsampling module is obtained:
[0060] (23),
[0061] Wherein, represents the batch normalization operation, represents the rectified linear unit activation function, which sets negative values to zero, represents the convolution process of
[0062] The output module consists of a final output layer and an image segmentation output part. The final output layer uses convolution to adjust the number of channels of the feature map with the same size as the original input obtained through multiple upsampling modules to the target output number of channels, specifically as follows:
[0063] (24),
[0064] Wherein, represents the activation function, represents the feature map output by the last upsampling module, represents the reconstructed image, represents the convolution process of . After the generation of the reconstructed image is completed, the sizes of the original image and the reconstructed image should be the same, and the pixel positions should correspond one by one to facilitate the calculation of the error by comparing pixel by pixel.
[0065] The present invention also includes a model training stage. In the model training stage, the total objective loss function consists of a mutual exclusion loss and a reconstruction loss, which are used to jointly constrain the learning process of the model, as shown in Equation (25).
[0066] (25),
[0067] Wherein, represents the mutual exclusion loss function, represents the reconstruction loss function, represents an adjustable hyperparameter.
[0068] The present invention uses the Euclidean distance as a measurement method for the mutual exclusion difference of the model. The reason is that this distance directly reflects the geometric difference between two groups of feature vectors, can quantify the relative independence between the backbone of effective information and redundant information, and by maximizing the Euclidean distance between the effective information features and the redundant information features, it can ensure that the model better separates these two types of information in the feature space to improve the modeling ability of the encoder for discriminative features. Its calculation formula is as follows: (26),
[0069] Wherein, represents the effective information feature of the layer of the effective information branch, Representing the redundant information features of the n-th layer of the redundant information branch, representing the effective information features of the output layer of the effective information branch, representing the redundant information features of the output layer of the redundant information branch, , representing the adjustable hyperparameter.
[0070] The present invention also has a reconstruction loss to measure the difference between the reconstructed image and the original image. By minimizing this loss, it can ensure that the model restores the result closest to the normal image. The mean squared error loss and the structural similarity loss are used as the reconstruction loss part of this method. Among them, the mean squared error loss mainly calculates the squared difference of pixel values between the input image and the reconstructed image, and is used to evaluate the degree of detail recovery of the image. Its calculation formula is as follows: (27),
[0071] In the formula, represents the pixel value at the position of the input image, represents the pixel value at the position of the reconstructed image, represents the height and width of the input and reconstructed images. The structural similarity loss (SSIM) measures the similarity between the brightness, contrast, and structure of two images, and is used to evaluate the ability of the model to maintain the semantic and structural consistency of the image. By minimizing this loss function, it can ensure that the reconstructed image is visually consistent with the original image and avoid distortion of the reconstructed image caused by over-optimizing the pixel-level loss. The calculation formula is as follows:
[0072] (28),
[0073] In the formula, , represents the pixel mean of the input image and the reconstructed image, , represents the pixel variance of the input image and the reconstructed image, represents the covariance of the input image and the reconstructed image, , represents constants for stable calculation, taking 1e-4 and 9e-4 respectively here. Finally, the calculation formula for the reconstruction loss of the model is:
[0074] (29),
[0075] In the formula, , represents the hyperparameter for controlling the weights of each loss term.
[0076] The present invention also includes anomaly detection of the model. In the anomaly detection stage of the model, it is necessary to infer the input image to be detected. The core objective is to compare the difference between the input image and its reconstructed image to discover potential abnormal regions. Specifically, assuming that the input image to be detected is , the trained network is used to encode and decode the input image to obtain the reconstructed image . The pixel-level reconstruction error is calculated to achieve anomaly detection and segmentation. The calculation method is:
[0077] (30),
[0078] In the formula, represents the reconstruction error at the spatial position ; represents the value of the original image to be detected at the spatial position; represents the value of the reconstructed image to be detected at the spatial position ;
[0079] After obtaining the reconstruction error at each pixel position, the anomaly score of the image can be further calculated. The anomaly score can be divided into two categories: local anomaly score and global anomaly score, which are used to locate pixel-level anomalies and judge image-level anomalies respectively. The calculation formula of the local anomaly score is:
[0080] (31),
[0081] In the formula, represents the reconstruction error of the pixel point ;
[0082] The calculation formula of the global anomaly score is:
[0083] (32),
[0084] In the formula, represents the global anomaly score of the input detection image, represents the height and width of the input detection image. The global anomaly score can be used as a rough screening index at the image level to quickly judge whether the image is abnormal.
[0085] According to the calculated local or global anomaly scores, whether it is abnormal is judged by setting local or global judgment thresholds. Among them, the global judgment threshold is used to evaluate the overall abnormal degree of the image and is statistically set based on the reconstruction error distribution of the normal images in the training set:
[0086] (33),
[0087] In the formula, represents the mean of the normal images in the training set, represents the standard deviation of the normal images in the training set, represents the coefficient for adjusting the sensitivity, with a value of 2.6, represents the overall judgment threshold. If the overall anomaly score of the image , then the image is determined to be abnormal;
[0088] Once the image is determined to be overall abnormal, a local judgment threshold can be further introduced to finely discriminate specific pixels in the image.
[0089] Aiming at the limitations of traditional supervised detection techniques in restricting the model performance due to problems such as high annotation cost, data imbalance, and label errors, based on the characteristics of a large amount of redundant information existing in the civil aviation engine borescope images, the present invention proposes a civil aviation engine borescope image damage detection method based on dual-path decoupling and gated memory fusion. This method adopts a dual-encoding - single-decoding architecture, including three stages: feature decoupling, feature fusion, and feature reconstruction. In the feature decoupling stage, a dual-branch feature decoupling architecture is constructed, combined with gradient reversal and mutual exclusion loss, to effectively separate the foreground key information from the redundant interference information; in the feature fusion stage, an adaptive gated and memory-guided feature fusion module is designed, using gated attention and a dynamic memory bank to adaptively adjust the fusion weights according to feature similarity, thereby enhancing the model's sensitivity to damage features; in the feature reconstruction stage, through a structure-aware reconstruction decoder, jointly optimizing the reconstruction error and the structure loss, enhancing the model's response ability and localization accuracy to the damaged area. Finally, through experiments, it is verified that this method has good adaptability and effectiveness in the civil aviation engine borescope image damage detection task. Brief Description of the Drawings
[0090] Att Figure 1 is the architecture diagram of the borescope image damage detection model based on dual-path decoupling and gated memory fusion in the present invention.
[0091] Att Figure 2 is the structure diagram of the DenseNet-121 network in the present invention.
[0092] Att Figure 3 is the schematic diagram of the feature fusion module in the present invention.
[0093] Att Figure 4 is the schematic diagram of the D2AM-Net anomaly localization result on the MVTecAD dataset in the embodiment of the present invention.
[0094] Att Figure 5 is the schematic diagram of the D2AM-Net anomaly localization result on the BID-DET dataset in the embodiment of the present invention.
[0095] Att Figure 6It is a curve graph showing the influence of the hyperparameter values of λ and γ of the loss function in the embodiments of the present invention on the image-level AUROC results of the model on BID-DET.
[0096] Appendix Figure 7 It is a curve graph showing the influence of the hyperparameter values of λ and γ of the loss function in the embodiments of the present invention on the image-level AUROC results of the model on MVTec AD.
[0097] Appendix Figure 8 It is a curve graph showing the influence of the hyperparameter values of λ and γ of the loss function in the embodiments of the present invention on the pixel-level AUROC results of the model on MVTec AD. Detailed implementation manners
[0098] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0099] To verify the effectiveness and superiority of the method of the present invention in unsupervised anomaly detection tasks, in this example, the publicly available dataset MVTec AD (MVTec Anomaly Detection Dataset) and the civil aviation engine borescope dataset (BID-DET) will be used to verify the method. The specific descriptions of the datasets are as follows:
[0100] MVTec AD dataset. This dataset is one of the relatively authoritative publicly available datasets in the current industrial vision field, covering 15 categories of common industrial objects and materials, including various types of defects in real scenarios, such as scratches, stains, cracks, and deformations. In the detection task setting, the defect types such as scratches, cracks, and contaminations included in MVTec AD have certain similarities with the common damage characteristics in civil aviation engine components. Therefore, using this dataset as the verification part dataset of the experiment helps to provide a reference for the actual application in the field of civil aviation engines.
[0101] BID-DET dataset. This dataset is collected from borescope images obtained from the borescope reports of major airlines, covering all key areas of the engine. The BID-DET dataset contains 3396 normal images, which usually have problems such as uneven illumination, specular interference, local blurring, and complex backgrounds. The abnormal images mainly have six forms of internal engine damage, such as cracks, ablation, chipping, material loss, pits, and warping. The training set for the model is all composed of normal images, and real damage images will be used in the test stage. The overall setting conforms to the setting of unsupervised anomaly detection tasks. This dataset helps to comprehensively evaluate the generalization ability and detection performance of the model in the borescope inspection environment.
[0102] The experiments conducted in this example were still completed in the Windows 11 operating system environment. The hardware configuration used in the experiments included an Intel i9-12900K processor and an NVIDIA GeForce RTX 4090 graphics card. The training and testing of all models were implemented based on the PyTorch deep learning framework. The PyTorch version used was 1.13.1, and the CUDA version was 11.6.
[0103] Its hyperparameter settings are shown in Table 1. It should be noted that although hyperparameter tuning is regarded as a key task in the field of deep learning and relies on a large number of experiments and empirical explorations, since the hyperparameters of some internal calculation items have relatively limited impact on the overall performance under reasonable default values, and the core of this paper is to construct an anomaly detection model applicable to the borescope inspection process of civil aviation engines, rather than optimizing the hyperparameters of existing models. Therefore, in order to avoid a sharp increase in the tuning complexity caused by overly expanding the hyperparameter search space, this paper chooses to fix some internal hyperparameters, making the model design more concise and easy to reproduce. In subsequent work, this paper will focus on solving the problem of hyperparameter tuning.
[0104] Table 1 Model Hyperparameter Settings
[0105]
[0106] The experimental part of this example mainly covers the following three items: anomaly detection performance comparison experiment, anomaly detection localization comparison experiment, and model ablation experiment. It should be noted that to ensure the consistency and objectivity of the evaluation, except that the performance data of AnoGAN, AE-L2, and AE-SSIM are cited from the experimental reports of Bergmann et al., the performance results of other methods mainly refer to the experimental results disclosed in their original papers, and some methods may also be cited from the reproduction or comparative studies of other subsequent works.
[0107] In addition, since the mutual exclusion loss and the reconstruction loss are the two core modules of the model training strategy, their weights directly determine the balance between information separation and reconstruction error control.
[0108] To verify the effectiveness of the unsupervised anomaly detection method proposed in the present invention, the area under the receiver operating characteristic curve (AUROC) is selected as the main evaluation index. AUROC is an important measure to evaluate the classification performance of the model at different decision thresholds, which can comprehensively reflect the overall ability of the model to effectively distinguish normal samples from abnormal samples without relying on specific threshold settings, so it has strong robustness and generalization ability. The closer the AUROC value is to 1, the stronger the discrimination ability of the model.
[0109] AUROC essentially calculates the area under the ROC curve and can be formally expressed as:
[0110] (34),
[0111] In the formula, represents the true positive rate, represents the false positive rate;
[0112] Among them, the definitions of TPR and FPR are as follows: In the anomaly detection task, the true positive rate (TPR) represents the proportion of samples correctly identified as anomalies among all true anomaly samples; the false positive rate (FPR) represents the proportion of normal samples misidentified as anomalies among all true normal samples.
[0113] In the anomaly localization task, TPR represents the proportion of pixels correctly identified as anomalies among all true anomaly pixels, and FPR represents the proportion of normal pixels misjudged as anomalies among all true normal pixels. The calculation formulas for TPR and FPR are:
[0114] (35), (36),
[0115] In the formula, represents the number of true anomaly samples correctly identified as anomalies, represents the number of true anomaly samples misjudged as normal, represents the number of true normal samples misjudged as anomalies, represents the number of true normal samples correctly identified as normal.
[0116] MVTec AD dataset: Comparative experiments on image-level anomaly detection were conducted on this dataset for this method and seven existing unsupervised image anomaly detection methods, namely AnoGAN, AE-L2, AE-SSIM, RIAD, MKD, P-SVDD, and US. The results are shown in Table 2.
[0117] Table 2 Comparison results of image-level anomaly detection for each method on the MVTec AD dataset
[0118]
[0119] The specific analysis of the experimental results is as follows:
[0120] First, for texture-like samples, D2AM-Net achieved AUC scores of 0.938, 0.982, and 0.998 in the three subcategories of Carpet, Grid, and Leather, respectively, significantly outperforming most of the comparison methods. In particular, it achieved a performance close to 1 in the Leather subcategory. This fully demonstrates that D2AM-Net has extremely strong feature expression and discrimination capabilities in dealing with subtle texture changes and local structural perturbations, especially in the perception of minor damages, which fully meets the demand for being highly sensitive to detail anomalies in civil aviation engine borescope inspections. In this example, it is believed that the manifestation of this advantage mainly stems from the model's focus on effective features during the feature decoupling stage, and the reinforcement of the independence of each encoding branch through the exclusive loss, thus fully avoiding the interference of redundant features. At the same time, through the precise adjustment of the feature fusion strategy by the feature fusion module, the model can quickly respond to anomalies when facing minor perturbations or texture inconsistencies in the image.
[0121] Second, for object-like samples, D2AM-Net also demonstrated excellent detection performance, with an overall average AUC reaching 0.911, which is comparable to the existing relatively excellent P-SVDD model and RIAD model. Among them, in categories such as Toothbrush, Hazelnut, and Zipper, D2AM-Net achieved high scores of 0.946, 0.986, and 0.993 respectively, highlighting the model's strong perception ability for anomalies such as local structural damages or deformations of objects. Especially in categories with rich details and complex structures such as Hazelnut and Zipper, the scores of D2AM-Net are close to the optimal, verifying the effectiveness of the joint optimization loss in the decoder in improving the quality of structural detail reconstruction.
[0122] From the above results, it can be seen that D2AM-Net shows high detection performance in most tasks. Especially for scenarios with complex texture backgrounds and fine-grained structural anomalies, it demonstrates strong robustness and generalization ability, and is suitable for the damage recognition task in the civil aviation engine borescope inspection environment.
[0123] BID-DET dataset: Based on the above detection results, this method compared the image-level damage detection performance with five unsupervised image anomaly detection methods, namely AnoGAN, AE-L2, AE-SSIM, RIAD, and P-SVDD, on the BID-DET dataset. The results are shown in Table 3. It should be noted that these results are all based on the open-source code or the reproduced code provided by the author to reconstruct the model and complete the evaluation in a unified experimental environment, ensuring the consistency of the evaluation conditions and the comparability of the results.
[0124] Table 3 Comparison results of image-level damage detection of each method on the borescope image dataset
[0125]
[0126] The specific analysis of the experimental results is as follows:
[0127] It can be seen from the experimental results on the BID-DET dataset that D2AM-Net demonstrates excellent performance and high robustness in various damage detection tasks. Compared with traditional reconstruction methods such as AnoGAN, AE-L2, and AE-SSIM, D2AM-Net has achieved significant improvements in all categories. Especially in anomalies with obvious structural or texture features such as Crack, Dent, and TBCMissing, it shows extremely high detection accuracy. Among them, the model obtained high scores of 0.998 and 0.983 in the Crack and Dent categories respectively, approaching or even exceeding the performance of some discriminative methods such as P-SVDD, indicating that the model has strong expressive ability in capturing key structural information. At the same time, in relatively more challenging categories such as Burn and Missing material, D2AM-Net also obtained scores of 0.985 and 0.899, showing good perception ability for non-significant anomalies, which is attributed to its clear division of effective and redundant information in the feature decoupling stage and the dynamic weight adjustment mechanism achieved through the gated attention mechanism in the fusion stage.
[0128] Overall, the stable performance of D2AM-Net on different types of damage not only verifies the effectiveness of its structural design but also fully demonstrates that this method has good generalization ability and robustness in complex backgrounds, fully verifying the practicality and foresight of this method in the detection of civil aviation engine borescope damage.
[0129] Anomaly localization experiment: The anomaly localization experiment mainly adopts two methods: quantitative evaluation and qualitative evaluation. Since the civil aviation engine borescope image dataset (BID-DET) obtained in this example lacks pixel-level annotations, it is impossible to directly evaluate the precise localization ability of the model in the anomaly area. Therefore, in this anomaly localization experiment, the MVTec AD dataset with pixel-level anomaly annotations is selected to quantitatively evaluate the model to verify the localization accuracy and generalization ability of the D2AM-Net model; at the same time, a qualitative evaluation is carried out on the BID-DET dataset, and the performance of the model in anomaly localization on the dataset is visually observed through the reconstruction loss heat map.
[0130] MVTec AD Dataset: In this example, pixel-level AUROC is used as the evaluation metric. On this dataset, the proposed method in this example is compared with seven existing unsupervised image anomaly detection methods (including AnoGAN, AE-L2, AE-SSIM, RIAD, MKD, P-SVDD, and US) in terms of anomaly localization performance. The experimental results are shown in Table 4 below and Figure 4 as follows.
[0131] Table 4 Comparison results of image-level damage localization of each method on the MVTec AD dataset
[0132]
[0133] The specific analysis of the experimental results is as follows:
[0134] Through the experimental results on the MVTec AD dataset, it can be proved that the D2AM-Net model still has good performance in the anomaly localization task. Although this dataset covers a large number of industrial defect types with different texture and structure characteristics, D2AM-Net still achieves comparable or even better localization accuracy than mainstream advanced methods in most categories. The overall average pixel-level AUROC reaches 0.949, which is significantly better than traditional autoencoder-based methods (such as 0.809 for AE-L2 and 0.863 for AE-SSIM), and there is only a slight gap compared with the currently strongest-performing P-SVDD (0.956), demonstrating strong comprehensive capabilities and robustness. Specifically, in categories with complex structures or delicate textures such as Leather, Screw, and Zipper, D2AM-Net achieved high scores of 0.995, 0.981, and 0.993 respectively, showing its high sensitivity in capturing subtle anomalies and edge perturbations. In contrast, the performance of traditional methods in these categories is generally unstable, indicating that while maintaining overall performance consistency, D2AM-Net can better adapt to the anomaly characteristics of different types of samples. In addition, in some object categories such as Toothbrush and Capsule, D2AM-Net also demonstrated the ability to approach or exceed the existing best methods, proving its adaptability in the scenarios of anomaly detection for complex structures and non-rigid objects.
[0135] It is worth emphasizing that although the MVTec AD dataset is not specifically designed for the aviation field, the high-resolution images, complex background interference, and various types of industrial defects it contains have certain similarities with the characteristics of civil aviation engine borescope images in terms of image quality and abnormal morphology. Therefore, the excellent localization performance achieved by D2AM-Net on this dataset can to a certain extent reflect its potential applicability and generalization ability in the borescope damage detection task. In particular, the stable performance of D2AM-Net in detail structure perception and under high background noise is the key ability to solve the key difficulties such as complex texture interference and tiny damage localization in borescope images.
[0136] BID-DET dataset: Since the BID-DET dataset does not have pixel-level labels, in this example, the actual effect of the model is verified from a qualitative dimension. It can be seen from the results of the following reconstructed heatmaps that the high-response regions highly coincide with the actual damage regions, indicating that the model has good abnormal feature perception ability and localization accuracy, and can accurately locate various types of abnormal regions, including structural damages such as cracks, ablation, chipping, and material loss. In addition, it can also be seen from Figure 5 that this method demonstrates good generalization ability and robustness in damage scenarios with different morphologies and scales, further verifying the sensitivity and adaptability of the D2AM-Net model to complex defect structures in actual borescope images, and having high engineering application potential, suitable as an auxiliary means for manual borescope damage detection.
[0137] In summary, D2AM-Net has certain advantages in overall localization accuracy and also demonstrates good stability and adaptability in the engine borescope inspection scenario, laying a solid foundation for its unsupervised anomaly detection and localization tasks in civil aviation engine borescope damage.
[0138] (3) Model ablation experiment
[0139] To verify the effectiveness of each module in the method proposed in this example, ablation experiments are designed and carried out in this example, removing or replacing key components in the model one by one to evaluate their impact on the overall performance. Among them, Tables 5 and 6 represent the results of the image-level anomaly detection ablation experiment of the D2AM-Net model on the borescope inspection dataset and the MVTec AD dataset. Table 7 represents the results of the pixel-level anomaly localization ablation experiment of the D2AM-Net model on the MVTec AD dataset.
[0140] Among them, the symbol "–" indicates removing the module or loss function from the complete model. Among them, "–DualEncoder" means removing the redundant information encoding branch, "–GRL" means removing the gradient reversal layer, and "–MutualLoss" means no longer using the mutual exclusion loss; in the feature fusion stage, "–GateAttention" means removing the gated attention mechanism, "–DynamicMemory" means removing the dynamic memory bank module, and "–FusionModule" means removing both the gated mechanism and the memory bank module in the fusion stage; in the feature reconstruction stage, "–StructureLoss" and "–MSELoss" respectively mean only retaining the mean squared error loss or the structure loss.
[0141] As shown in the results, after removing any module or loss function, the performance of the model decreases to varying degrees, indicating that the design of each proposed module and joint loss has a positive effect on the overall anomaly detection performance. The complete D2AM-Net model achieves the best detection effect, verifying the effectiveness and collaborative gain of the multi-stage feature processing and fusion strategy proposed in this example.
[0142] Table 5 Results of ablation experiments on image-level damage detection of D2AM-Net on the borescope inspection damage dataset
[0143]
[0144] Table 6 Results of ablation experiments on image-level damage detection of D2AM-Net on the MVTec AD dataset
[0145]
[0146] Table 7 Results of ablation experiments on pixel-level damage localization of D2AM-Net on the MVTec AD dataset
[0147]
[0148] To enable D2AM-Net to work well in coordination at each stage, this example explores the weight hyperparameters of the mutual exclusion loss and the reconstruction loss in its joint optimization loss function. Given that it focuses on structural guidance rather than directly affecting the reconstruction result, usually only a small loss weight needs to be assigned. Therefore, in this example, its weight value is set in the range of 0.1 to 0.5. The reconstruction loss directly measures the difference between the reconstructed image and the original image and is the main supervision signal supporting the model's reconstruction ability. Therefore, its weight is set between 0.5 and 1.0 to ensure the dominant position of the reconstruction target.
[0149] The experimental results are as Figure 6 、 7, as shown in Figure 8. Through experiments, it can be seen that when the mutual exclusion loss weight is set to 0.3 and the reconstruction loss is set to 0.6, the model can achieve the best performance in both image-level and pixel-level anomaly detection tasks. In this example, it is considered that the reason for the excellent performance under this configuration is that it can well balance the extraction of effective information and the elimination of redundant information. The mutual exclusion loss of 0.3 is sufficient to effectively constrain the feature overlap between the two encoding branches, prompting the effective information branch to focus on extracting high-semantic features related to the task, while ensuring that the redundant information branch captures background and noise features and does not cause too much negative interference to the main task; and the reconstruction loss of 0.6 provides sufficient optimization strength in the pixel-level and structure-level reconstruction processes, enabling subtle anomaly regions to be clearly reflected in the reconstruction error, thereby improving the accuracy of image anomaly detection and the accuracy of anomaly localization.
[0150] The discussion on the rationality of hyperparameter configuration is as follows: From the trend of the curve itself, it can be seen that when γ significantly increases to 0.5, in both image-level and pixel-level detections, some λ values will show a significant decline in performance, indicating that too strong mutual exclusion constraints may inhibit the model's capture of key information; similarly, when λ significantly decreases to below 0.5, the training signal provided by the reconstruction loss is insufficient, and the pixel-level discrimination of the model for anomaly regions is often not ideal. Therefore, considering the overall change trend of the curve, the setting of the mutual exclusion loss γ value between 0.1 - 0.5 and the reconstruction loss λ value between 0.5 - 1.0 is reasonable. In the experiment γ = 0.3, λ = 0.6 configuration achieved relatively stable and excellent detection results, which is consistent with the corresponding theoretical setting.
Claims
1. A method for detecting damage in borescope images through dual-channel decoupling and gated memory fusion, characterized in that, A borehole inspection image damage detection model with dual-path decoupling and gated memory fusion is established and trained only with unlabeled normal images. The establishment of the borehole inspection image damage detection model with dual-path decoupling and gated memory fusion includes feature decoupling, feature fusion, and feature reconstruction. In the feature decoupling stage, a dual-backbone encoder is adopted, which serves as an effective information encoding branch and a redundant information encoding branch respectively; In the feature fusion stage, a feature fusion module with adaptive gating and memory guidance is provided. Among them, the gating mechanism adaptively adjusts the feature fusion weights of the dual-backbone encoding branches according to the similarity between the features extracted from the effective information encoding branch for the current input and the features of normal images stored in the memory bank. The feature content in the memory bank will be dynamically updated according to the training progress of the effective information encoder to adapt to the continuous evolution of the model's representation ability; In the feature reconstruction stage, a structure-aware reconstruction decoder is established, and the fused representation features are fed into the structure-aware reconstruction decoder for image reconstruction. The structure-aware reconstruction decoder adopts a joint optimization loss function.
2. The method for detecting damage in borescope images by dual-channel decoupling and gated memory fusion according to claim 1, wherein, In the feature decoupling stage during the establishment process of the borehole inspection image damage detection model with dual-path decoupling and gated memory fusion, the effective information encoding branch is used to extract effective features with high semantic consistency in the image, and the redundant information encoding branch, by introducing a gradient reversal layer, is used to extract redundant or non-discriminative features that are mutually exclusive with the backbone information in the image. A mutual exclusion loss is also introduced. By minimizing the correlation between the output features of different levels of the two encoding branches, the process of feature decoupling is promoted, ensuring that the effective information branch focuses on learning task-related features, while the redundant information branch focuses on capturing features irrelevant to anomaly detection.
3. The method for detecting damage of borescope images by dual-channel decoupling and gating memory fusion according to claim 2, wherein, DenseNet-121 is used as the backbone network of the dual-backbone encoder. The overall structure of DenseNet-121 includes an initial convolutional layer, dense blocks, and transition layers. Among them, in the initial stage of the network, the initial convolutional layer is used to extract basic features from the input image. Subsequently, the obtained initial features are connected to the dense block part. The dense block uses channel concatenation of all previous outputs of each layer as the input of the current layer to promote feature reuse and alleviate the problem of gradient disappearance, as shown in the following formula: (1) , In the formula, is composed of Batch Normalization, ReLU, and convolution; Next, a transition layer is designed between each dense block and the next dense block to perform downsampling and channel compression on the feature map, using 1×1 convolution and average pooling operations to reduce the spatial size and the number of feature maps.
4. The method for detecting damage in borescope images by dual-channel decoupling and gated memory fusion according to claim 3, wherein, In the feature decoupled dual backbone encoder, the effective information encoding backbone adopts the standard DenseNet-121 architecture. Let the input image be , and the feature representation extracted by the effective information encoding backbone is: (2), In the formula, represents the input image, represents the effective information encoder constructed based on the DenseNet-121 architecture, represents the output effective feature map, , represents the height, width, and number of channels of the encoded feature map; The redundant information encoding backbone is based on DenseNet-121, and a gradient reversal layer GRL is introduced at the output end of the backbone. GRL is an identity mapping during forward propagation, while it reverses the gradient direction during backward propagation, thereby guiding this backbone to learn features that are mutually exclusive with the effective information: (3), In the formula, represents the output redundant feature map, , represents the redundant information encoder constructed based on the DenseNet-121 architecture, represents the gradient reversal layer, and the hyperparameter controlling the gradient reversal intensity is , Among them, the definition of the gradient reversal layer is as follows, where represents the identity matrix, indicating that the gradient is multiplied during backpropagation by to achieve the reversal of the training direction of the redundant branch. (4), (5), For the hyperparameter of gradient reversal strength a dynamic adjustment strategy is adopted to balance stability and adversariality in the learning process, as shown in the following formula: (6), In the formula, represents the training progress, .
5. A method for detecting damage in borescope images by dual-channel decoupling and gated memory fusion according to claim 4, characterized in that, The adaptive gating and memory-guided feature fusion module in the feature fusion part dynamically integrates the features output by different backbone encoders during the reconstruction process to achieve differential modeling of different types of inputs. The goal of the feature fusion part is to perform weighted fusion on the features extracted by the dual backbone encoders of the feature decoupling module at each spatial position and its fused features are expressed as: (7), wherein, represents the fusion weight map generated by the gating mechanism, which is used to control the proportion of the features of the main branch at each spatial position, , in order to enable the gating weight to be dynamically adjusted according to the standard of normal samples, a memory bank module is introduced, including the following: First, in the th training stage, record the current reconstruction loss value as , and record the trend of the smoothed loss by using exponential moving average (EMA): (8), In the formula, represents the loss smoothed mean at the -th moment, represents the smoothing parameter, and its value is 0.
9. Next, assume that the size of the initial memory bank is , and the current number of features in the memory bank is . The update goal is to gradually and dynamically update the number of features in the memory bank when the model training gradually stabilizes and the loss significantly decreases: (9), In the formula, represents the smoothed mean of losses at the -th moment, represents the smoothed mean of losses at the -1-th moment, represents the Sigmoid function, which is used to normalize the difference to the interval [0, 1], represents the growth rate, , (10), In the formula, is used to control the slope, i.e., the sensitivity, of the Sigmoid function, , is used to control the midpoint position of the function, which is set to 0.5 to facilitate triggering the expansion of the memory bank scale when the loss starts to decrease significantly; among them, the number of prototypes stored in the memory bank in the initial stage needs to ensure that too much noise or unstable features are not written into the memory bank during the initial training, so a relatively moderate scale is selected, with the value K 0 = 128, Finally, the memory bank is regarded as a dynamically updated set of features: (11), In the formula, represents the number of prototypes in the training phase memory bank, represents the typical feature of a certain type of valid information in the normal samples.
6. The method for detecting damage of borescope images by dual-channel decoupling and gated memory fusion according to claim 5, characterized in that, After each training, the image features output by each valid information encoder are sent to the decoder one by one for reconstruction. Subsequently, according to the current available capacity of the memory bank, several features with the smallest reconstruction error are selected from the current batch for updating. In this way, it is ensured that the stored features have a high reconstruction ability, that is, they have a stronger expression ability for the normal mode, thereby improving the discrimination accuracy of the gating mechanism for anomalies. Specifically: First, in the current training iteration, a set of candidate features are extracted through batch training ; Subsequently, the reconstruction error of each feature is calculated separately ; Finally, the first features with the smallest error are written into the memory bank. To avoid frequently disturbing the stable memory content in the later stage of training, a decay coefficient is also provided to ensure that as the training process is updated, the update rate of the features in the memory bank gradually slows down. The exponential weighted average (EMA) update strategy is adopted for the update of the prototypes in the memory bank for the selected valid information features (12), In the formula, represents the number of prototypes in the memory bank, represents the dynamic update decay coefficient, , At the same time, a dynamic update decay coefficient is set, and the formula is: (13), Wherein, represents the EMA decay coefficient at the th iteration, taking 0.2, taking 0.8, taking 1e-2, represents the current iteration number.
7. A method for detecting damage in borescope images by dual-channel decoupling and gated memory fusion according to claim 6, characterized in that, The goal of the gating mechanism is to perform dynamic weighted fusion on the features output by the dual-encoding branches to achieve adaptive fusion regulation of effective information and redundant information, including: First, perform similarity calculation. Calculate the similarity between the effective feature at a certain spatial position in the input image and each prototype in the memory bank. The cosine similarity is used for calculation, and the formula is as follows: between them. The calculation is carried out using cosine similarity, and the formula is as follows: (14), In the formula, represents the typical feature of a certain type of valid information in the normal sample, represents the feature extracted by the valid coding branch at the spatial position , represents the cosine similarity between the feature at the spatial position and the th prototype. Next, the similarities of all prototypes are normalized by the Softmax function to obtain the normalized matching weights of each prototype feature at this position: (15), In the formula, represents the normalized matching weight, Then, define the position The score of the overall similarity is as follows: (16), To normalize the scores to the interval [0, 1], gating weights are generated through the Sigmoid function: (17), can be regarded as a gating coefficient reflecting the "importance" or "selective transmission" of the features of the image at the position. If is high, it means that when the effective information matches well with the memory bank prototype, a higher weight is given; if is low, there may be an anomaly, and it is necessary to enhance the influence of redundant information. The contribution of the redundant information coding branch is relatively enhanced, resulting in a large error in the abnormal area during reconstruction.
8. A method for detecting damage in borescope images by dual-channel decoupling and gated memory fusion according to claim 7, characterized in that, The structure-aware reconstruction decoder network in the feature reconstruction part includes an upsampling module and an output module. The upsampling module consists of a transposed convolution, batch normalization, an activation function, and a skip connection part. Let the input feature map of the -th upsampling module of the decoder be , and be the height and width of the feature map respectively. After the feature map enters the upsampling module, first, its spatial dimensions are enlarged through a transposed convolution, and the transposed feature map is obtained after the transposed convolution: (18), In the formula, represents the layer, represents the transposed feature map, and its mathematical expression should be: (19), In the formula, represents the kernel function of the transposed convolution, represents the stride of the transposed convolution, and the value here is 2. represents the spatial position in represents the spatial coordinates in Among them, hyperparameters such as the convolution kernel size and stride need to be set according to specific tasks; Next, perform batch normalization and activation function on the transposed convolution output The purpose of using batch normalization is to slow down the change of feature distribution, make the input distribution change less during the training process of each layer, ensure the stability of the input distribution of each layer, and alleviate the problems of gradient disappearance or gradient explosion. For each channel of the input feature map Batch normalization operation will calculate the mean and variance , and then through the learnable scaling parameter and translation parameter , which are specifically described as follows: (20), In the formula, represents the feature map after transposition, represents the batch normalization operation, takes 1e-4; Using ReLU (Rectified Linear Unit) as the activation function to provide the network with the ability of non-linear transformation, and finally through batch normalization processing and the feature map of the activation function is expressed as follows: (21), In the formula, represents the batch normalization operation, represents the rectified linear unit activation function, which sets negative values to zero; Through skip connections, the low-level features of the corresponding spatial dimensions obtained from the effective information branch backbone of the encoder are concatenated with the processed features in the channel dimension to obtain concatenated features : (22), In the formula, represents the concatenation operation in the channel dimension, represents the i skip feature of the effective information branch of the th layer, and represents the number of channels of the skip connection of this layer. Subsequently, further perform convolution fusion on the spliced feature map and again pass through batch normalization and activation function to obtain the output feature of the upsampling module : (23), In the formula, represents the batch normalization operation, represents the rectified linear unit activation function, which sets negative values to zero, represents the convolution operation of The output module consists of a final output layer and an image segmentation output part. The final output layer uses convolution to adjust the number of channels of the feature map with the same size as the original input obtained through multiple upsampling modules to the target output number of channels, specifically as follows: (24), In the formula, represents the activation function, represents the feature map output by the last upsampling module, represents the reconstructed image, represents the convolution process of. After the generation of the reconstructed image is completed, the sizes of the original image and the reconstructed image should be the same, and the pixel positions should correspond one by one to facilitate the calculation of the error by comparing pixel by pixel.
9. A method for detecting damage to borescope images by dual-channel decoupling and gated memory fusion according to claim 8, characterized in that, It also includes a model training stage, in which the total objective loss function is composed of a mutual exclusion loss and a reconstruction loss, which are used to jointly constrain the learning process of the model, as shown in Equation (25), (25), In the formula, represents the mutual exclusion loss function, represents the reconstruction loss function, represents an adjustable hyperparameter.
Citation Information
Patent Citations
Multi-layer feature decoupling optical remote sensing image building extraction method
CN115731461A
Collection, calculation and inspection integrated bridge damage detection method and system
CN116626059A