Two-way decoupling and gated memory fused borescope image damage detection method

Through the hole-detection image damage detection method fusion with dual-channel decoupling and gated memory, the problems of key discriminative information loss in high-frequency detail loss, pattern collapse and feature decoupling in civil aviation engine hole-detection image damage detection are solved, and damage detection with high accuracy and reliability is achieved.

CN120147320AActive Publication Date: 2025-06-13HARBIN INST OF TECH AT WEIHAI

Patent Information

Application Number
CN202510629149.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-06-13
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

The prior art has problems such as loss of high-frequency details, pattern collapse and loss of key discrimination information during the process of feature decoupling in civil aviation engine hole image damage detection, resulting in missed or misdetection.

Method used

A hole-detection image damage detection method is proposed for dual-channel decoupling and gated memory fusion. By establishing a dual backbone encoder, a feature fusion module for adaptive gating and memory guidance, and a structure-aware reconstruction decoder, it realizes feature decoupling, fusion and reconstruction, and improves the intelligence level of detection and consistency of interpretation.

Benefits of technology

It effectively improves the intelligence level of detection and consistency of interpretation, realizes open recognition of unknown abnormal types, reduces missed and missed detection, and improves the accuracy and reliability of civil aviation engine hole detection image damage detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147320A_ABST
    Figure CN120147320A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of complex equipment fault identification processing, in particular to a two-way decoupling and gated memory fused borescope image damage detection method, which is characterized in that a two-way decoupling and gated memory fused borescope image damage detection model is established, and training is performed only by adopting a normal image without a label; establishment of the two-way decoupling and gated memory fused borescope image damage detection model comprises feature decoupling, feature fusion and feature reconstruction, and in the feature decoupling stage, double-trunk encoders are adopted to serve as an effective information encoding branch and a redundant information encoding branch respectively; in the feature fusion stage, an adaptive gating and memory guiding feature fusion module is arranged, a structure perception reconstruction decoder is established in the feature reconstruction stage, fused representation features are sent to the structure perception reconstruction decoder for image reconstruction, and the structure perception reconstruction decoder adopts a joint optimization loss function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of complex equipment fault identification and processing, and specifically to a detection method for damage of borescope images with dual-channel decoupling and gated memory fusion, which can effectively improve the intelligent level of detection and the consistency of interpretation. Background Art

[0002] First, a reconstruction anomaly detection method based on an encoder-decoder architecture. This method constructs a mapping network of normal samples and uses the reconstruction residual anomaly value of the damaged area to achieve pixel-level positioning. Such methods have achieved outstanding results in the field of industrial product surface defect detection. However, in the scenario of civil aviation engine borescope inspection, due to the influence of image-irrelevant information such as background noise and residual signals during the reconstruction of the internal blade texture of civil aviation engines, high-frequency details are easily lost, resulting in confusion between the anomaly value of the damage area residual and the normal texture fluctuation, thus causing missed detection or false detection. Second, an anomaly detection method based on adversarial training. This method establishes an implicit distribution of normal samples through adversarial training, and the discriminant network is highly sensitive to the pattern deviation of the damaged area. These methods also have extensive applications in the field of industrial product surface defect detection. However, for the damage identification of civil aviation engines, due to the characteristics of small scale and irregular shape of the damage features of civil aviation engines, the damage identification of engine borescope images often relies on local fine structure information. If the form of adversarial training is adopted, it is easy to have a mode collapse problem, that is, the generator tends to learn global features and ignore local details, making it difficult to effectively capture and reconstruct the fine damage details existing in the borescope images, and unable to fully cover the feature distribution of normal samples, thus causing an increase in the false alarm rate. Third, an anomaly detection method based on contrastive learning. The feature decoupling framework of this method constructs positive and negative sample pairs through data augmentation to amplify the distinguishability between the damaged area and the normal texture in the feature space. Such methods rely too much on data augmentation strategies. At the same time, traditional contrastive learning assumes that positive and negative samples have symmetric or alignable feature distributions. However, for civil aviation engine borescope images, the geometric and texture characteristics of the damaged area and the normal structure are significantly different, and this asymmetry may lead to the loss of key discriminant information during the feature decoupling process, making it difficult for the model to effectively separate damaged and normal features.. Summary of the Invention

[0003] The present invention aims at the disadvantages and deficiencies existing in the prior art, and proposes a detection method for damage of borescope images with dual-channel decoupling and gated memory fusion, which can break through the category limitations of labeled samples, achieve open recognition of unknown anomaly types, and effectively improve the intelligent level of detection and the consistency of interpretation.

[0004] The present invention is achieved by the following measures: A method for detecting damage in borescope images with dual-path decoupling and gated memory fusion, characterized by establishing a borescope image damage detection model with dual-path decoupling and gated memory fusion, and only using unlabeled normal images for training. The establishment of the borescope image damage detection model with dual-path decoupling and gated memory fusion includes feature decoupling, feature fusion, and feature reconstruction. In the feature decoupling stage, a dual-backbone encoder is adopted, serving as the effective information encoding branch and the redundant information encoding branch respectively; In the feature fusion stage, there is a feature fusion module guided by adaptive gating and memory. The gating mechanism adaptively adjusts the feature fusion weights of the dual-backbone encoding branches according to the similarity between the features extracted from the effective information encoding branch of the current input and the features of normal images stored in the memory bank. The feature content in the memory bank will be dynamically updated according to the training progress of the effective information encoder to adapt to the continuous evolution of the model's representation ability; In the feature reconstruction stage, a structure-aware reconstruction decoder is established, and the fused representation features are sent to the structure-aware reconstruction decoder for image reconstruction. The structure-aware reconstruction decoder adopts a joint optimization loss function.

[0005] In the feature decoupling stage of the establishment process of the borescope image damage detection model with dual-path decoupling and gated memory fusion of the present invention, the effective information encoding branch is used to extract effective features with high semantic consistency in the image, and the redundant information encoding branch, by introducing a gradient reversal layer, is used to extract redundant or non-discriminative features in the image that are mutually exclusive with the backbone information. A mutual exclusion loss is also introduced. By minimizing the correlation between the output features of different levels of the two encoding branches, the process of feature decoupling is promoted, ensuring that the effective information branch focuses on learning task-related features, while the redundant information branch focuses on capturing features irrelevant to anomaly detection.

[0006] The present invention uses DenseNet-121 as the backbone network of the dual-backbone encoder. The overall structure of DenseNet-121 includes an initial convolutional layer, dense blocks, and transition layers. Among them, in the initial stage of the network, the initial convolutional layer is used to extract basic features of the input image. Subsequently, the obtained initial features are connected to the dense block part. The dense block promotes feature reuse by taking the concatenation of all previous outputs of each layer as the input of the current layer to alleviate the problem of gradient disappearance, as shown in the following formula: (1), where, is composed of Batch Normalization, ReLU, and convolution; Next, a transition layer is designed between each dense block and the next dense block to downsample the feature map and compress the channels, using 1×1 convolution and average pooling operations to reduce the spatial size and the number of feature maps, thereby effectively controlling the computational amount and memory occupancy of the entire network.

[0007] In the dual backbone encoder with feature decoupling of the present invention, the effective information encoding backbone adopts the standard DenseNet-121 architecture. Let the input image be , and the feature representation extracted by the effective information encoding backbone is: (1), where represents the input image, represents the effective information encoder constructed based on the DenseNet-121 architecture, represents the output effective feature map, , represents the height, width, and number of channels of the encoded feature map; The redundant information encoding backbone is also based on DenseNet-121, but a gradient reversal layer GRL is introduced at the output end of the backbone. GRL is an identity mapping during forward propagation, while it reverses the gradient direction during backpropagation, thereby guiding this backbone to learn features mutually exclusive with the effective information: (3), where represents the output redundant feature map, , represents the redundant information encoder constructed based on the DenseNet-121 architecture, represents the gradient reversal layer, and the hyperparameter controlling the gradient reversal intensity is , where the definition of the gradient reversal layer is as follows, where represents the identity matrix, indicating that the gradient is multiplied by during backpropagation to achieve the reversal of the training direction of the redundant branch.

[0008] (4), (5), for the hyperparameter of the gradient reversal intensity, a dynamic adjustment strategy is adopted to balance the stability and adversarial nature in the learning process, as shown in the following formula: (6), where represents the training progress, .

[0009] In the present invention, the adaptive gating and memory-guided feature fusion module in the feature fusion part dynamically integrates the features output by different backbone encoders during the reconstruction process to achieve differential modeling of different types of inputs. The goal of the feature fusion part is to perform weighted fusion on the features extracted by the dual-backbone encoder of the feature decoupling module at each spatial position for weighted fusion, and its fused features are expressed as: (7), wherein, represents the fusion weight map generated by the gating mechanism and is used to control the proportion of the main branch features at each spatial position, , in order to enable the gating weight to be dynamically adjusted according to the standard of normal samples, a memory bank module is introduced, and an adaptive gating and memory-guided feature fusion module is used to establish a memory bank and a gating attention mechanism. The construction of the memory bank dynamically adjusts the number of features in the memory bank through the long-term change trend of the reconstruction error, so as to more comprehensively capture the diversity of normal samples and improve the judgment accuracy of the gating mechanism, including the following contents: First, at the th training stage, record the current reconstruction loss value as , and use the exponential moving average (EMA) method to record the trend of the smoothed loss: (8), wherein, represents the smoothed average loss at the th moment, represents the smoothing parameter, with a value of 0.9, Next, assume that the size of the initial memory bank is , and the current number of features in the memory bank is . The update goal is to gradually and dynamically update the number of features in the memory bank when the model training gradually stabilizes and the loss significantly decreases: (9), wherein, represents the smoothed average loss at the th moment, represents the smoothed average loss at the -1th moment, represents the Sigmoid function, which is used to normalize the difference to the [0,1] interval, represents the growth rate, , (10), wherein, is used to control the slope of the Sigmoid function, i.e., the sensitivity, , To control the midpoint position of the function (determining the starting point of the memory bank growth), which is set to 0.5 here to facilitate triggering the expansion of the memory bank scale when the loss starts to significantly decline; among them, the number of prototypes stored in the memory bank in the initial stage needs to ensure that too much noise or unstable features are not written into the memory bank during the initial training. Therefore, a relatively moderate scale is selected, with the value K 0 = 128, which can cover sufficient sample diversity and avoid introducing unstable features prematurely. Finally, the memory bank is regarded as a dynamically updated feature set: (11), where represents the number of prototypes in the memory bank during a training stage, represents the typical features of a certain type of valid information in normal samples; After each training, the image features output by each valid information encoder are sent to the decoder one by one for reconstruction. Subsequently, according to the current available capacity of the memory bank, several features with the smallest reconstruction error are selected from the current batch for update. Through this form, it is ensured that the stored features have a high reconstruction ability, that is, they have a stronger expression ability for normal patterns, thereby improving the discrimination accuracy of the gating mechanism for anomalies. Specifically: First, in the current training iteration, a set of candidate features are extracted through batch training; Subsequently, the reconstruction error of each feature is calculated ; Finally, the first features with the smallest error (note that here will be adjusted with the number of iterations) are written into the memory bank; At the same time, to avoid frequently disturbing the stable memory content in the later stage of training, a decay coefficient is also provided to ensure that as the training process progresses, the update rate of the features in the memory bank gradually slows down. The exponential weighted average (EMA) update strategy is adopted for the update of the prototypes in the memory bank for the selected valid information features: (12), where represents the number of prototypes in the memory bank, represents the dynamic update decay coefficient, , At the same time, the dynamic update decay coefficient is set, and the formula is: (13), where represents the EMA decay coefficient at the th iteration, represents the smaller value at the initial moment, taking 0.2, represents the larger value at the initial moment, taking 0.8 here. It represents a hyperparameter that controls the growth rate, and here it is taken as 1e-2. It represents the current iteration number.

[0010] The goal of the gating mechanism in the present invention is to perform dynamic weighted fusion on the features output by the double-encoding branches to achieve adaptive fusion regulation of effective information and redundant information, including: First, similarity calculation is performed. For the effective features at a certain spatial position in the input image and each prototype in the memory bank, the similarity is calculated. Considering that for unsupervised anomaly detection tasks, the scale information of features often does not have direct discriminative significance, and more attention is paid to the pattern (direction) rather than the absolute value. Therefore, cosine similarity is used for calculation, and its formula is as follows: (14), In the formula, represents the typical feature of a certain type of effective information in normal samples, represents the feature extracted by the effective encoding branch at the spatial position , represents the spatial position feature and the th prototype, Next, the similarities of all prototypes are normalized through the Softmax function to obtain the normalized matching weights of each prototype feature at this position: (15), In the formula, represents the normalized matching weight, Then, define the score of the overall similarity at the position as: (16), To normalize the score to the interval [0, 1], the gating weight is generated through the Sigmoid function: (17), can be regarded as a gating coefficient that reflects the "importance" or "selective transmission" of the features at the position of the image. If is relatively high, it means that when the effective information matches well with the memory bank prototype, a higher weight is given; if is relatively low, there may be anomalies, and it is necessary to enhance the influence of redundant information, and the contribution of the redundant information encoding branch is relatively enhanced, resulting in a larger error in the abnormal area during reconstruction.

[0011] In the present invention, the structure-aware reconstruction decoder network of the feature reconstruction part includes an upsampling module (UpBlock) and an output module. The upsampling module is mainly used to gradually restore the spatial resolution, and the output module is to adjust the number of channels of the feature map to be the same as that of the input image to generate a reconstructed image. Among them, the upsampling module consists of a transposed convolution (Transposed Convolution, TC), batch normalization (Batch Normalization, BN), an activation function, and a skip connection part. Let the input feature map of the th upsampling module of the decoder be , and be the height and width of the feature map respectively. After the feature map enters the upsampling module, first, its spatial size is enlarged through the transposed convolution, and the transposed feature map is obtained after the transposed convolution: (18), where represents the th layer, represents the transposed feature map, and its mathematical expression should be: (19), where represents the kernel function of the transposed convolution, represents the stride of the transposed convolution, which takes a value of 2 here, represents the spatial position in represents the spatial coordinates in Among them, hyperparameters such as the convolution kernel size and stride need to be set according to specific tasks; Next, batch normalization processing and an activation function are performed on the output of the transposed convolution. The purpose of using batch normalization is to slow down the change of the feature distribution, so that the input distribution changes less during the training process of each layer, ensure the stability of the input distribution of each layer, and alleviate the problem of gradient disappearance or gradient explosion. For each channel of the input feature map , the batch normalization operation will calculate the mean and variance , and then through the learnable scaling parameter and the translation parameter , the specific description is as follows: (20), where represents the transposed feature map, represents the batch normalization processing operation, A constant representing a smaller value, used to prevent division by zero in the formula, and here 1e-4 is taken; The ReLU (Rectified Linear Unit) is used as the activation function to provide the network with the ability of nonlinear transformation. Finally, the feature maps of batch normalization processing and the activation function are expressed as follows: (21), In the formula, represents the batch normalization processing operation, represents the rectified linear unit activation function, which sets negative values to zero; Then, through skip connections, the low-level features of the corresponding spatial dimensions obtained from the effective information branch backbone of the encoder need to be concatenated with the processed features in the channel dimension to obtain the concatenated features : (22), In the formula, represents the concatenation operation in the channel dimension, represents the i skip features of the effective information branch of the layer, represents the number of channels of the skip connection of this layer, Subsequently, the concatenated feature map is further convolutionally fused, and again passes through batch normalization and the activation function to obtain the output features of the upsampling module : (23), In the formula, represents the batch normalization processing operation, represents the rectified linear unit activation function, which sets negative values to zero, represents convolution processing of; The output module consists of a final output layer and an image segmentation output part. The final output layer uses convolution to adjust the number of channels of the feature map with the same size as the original input obtained through multiple upsampling modules to the target output number of channels, specifically as follows: (24), In the formula, represents the activation function, represents the feature map output by the last upsampling module, represents the reconstructed image, represents For the convolution process, after the reconstructed image is generated, the sizes of the original image and the reconstructed image should be the same, and the pixel positions should correspond one by one to facilitate the calculation of the error by comparing each pixel.

[0012] The present invention also includes a model training stage. During the training stage of the model, the total objective loss function is composed of a mutual exclusion loss and a reconstruction loss, which are used to jointly constrain the learning process of the model, as shown in Equation (25).

[0013] (25), In the formula, represents the mutual exclusion loss function, represents the reconstruction loss function, represents an adjustable hyperparameter.

[0014] The present invention uses the Euclidean distance as a measurement method for the mutual exclusion difference of the model. The reason is that this distance directly reflects the geometric difference between two sets of feature vectors, can quantify the relative independence between the backbone of effective information and redundant information, and by maximizing the Euclidean distance between the effective information features and the redundant information features, it can ensure that the model better separates these two types of information in the feature space to improve the modeling ability of the encoder for discriminative features. Its calculation formula is as follows: (26), In the formula, represents the effective information feature of the th layer of the effective information branch, represents the redundant information feature of the th layer of the redundant information branch, represents the effective information feature of the output layer of the effective information branch, represents the redundant information feature of the output layer of the redundant information branch, , represents an adjustable hyperparameter.

[0015] The present invention also has a reconstruction loss to measure the difference between the reconstructed image and the original image. By minimizing this loss, it can ensure that the model restores the result closest to the normal image. The mean square error loss and the structural similarity loss are used as the reconstruction loss part of this method. Among them, the mean square error loss mainly calculates the square difference of the pixel values between the input image and the reconstructed image, which is used to evaluate the degree of detail recovery of the image. Its calculation formula is as follows: (27), In the formula, represents the pixel value at the th position of the input image, represents the pixel value at the The pixel value at the position, which represent the height and width of the input and reconstructed images. The structural similarity loss (SSIM) is used to evaluate the model's ability to maintain the semantic and structural consistency of the image by measuring the similarity between the brightness, contrast, and structure of the two images. By minimizing this loss function, it can ensure that the reconstructed image is visually consistent with the original image and avoid distortion of the reconstructed image caused by over-optimizing the pixel-level loss. The calculation formula is as follows: (28), where, , represents the pixel mean of the input image and the reconstructed image, , represents the pixel variance of the input image and the reconstructed image, represents the covariance of the input image and the reconstructed image, , represents constants used to stabilize the calculation, taking 1e-4 and 9e-4 respectively here. Finally, the calculation formula for the reconstruction loss of the model is: (29), where, , represents a hyperparameter used to control the weights of each loss term.

[0016] The present invention also includes anomaly detection of the model. In the anomaly detection stage of the model, it is necessary to infer the input image to be detected. Its core goal is to compare the difference between the input image and its reconstructed image to discover potential abnormal regions. Specifically, assume that the given input image to be detected is , and use the trained network to encode and decode the input image to obtain the reconstructed image . Calculate the pixel-level reconstruction error to achieve anomaly detection and segmentation. The calculation method is: (30), where, represents the reconstruction error at the spatial position , represents the value of the original image to be detected at the spatial position, represents the value of the reconstructed image to be detected at the spatial position ; After obtaining the reconstruction error at each pixel position, the anomaly score of the image can be further calculated. The anomaly score can be divided into two categories: local anomaly score and global anomaly score, which are used to locate pixel-level anomalies and judge image-level anomalies respectively. The calculation formula for the local anomaly score is: (31), In the formula, represents the reconstruction error of the pixel point ; The calculation formula of the overall anomaly score is: (32), In the formula, represents the overall anomaly score of the input detection image, represents the height and width of the input detection image. The overall anomaly score can be used as a rough screening index at the image level to quickly judge whether the image is abnormal. According to the calculated local or overall anomaly scores, whether it is abnormal is judged by setting local or overall judgment thresholds. The overall judgment threshold is used to evaluate the overall anomaly degree of the image and is statistically set based on the reconstruction error distribution of the normal images in the training set: (33), In the formula, represents the mean value of the normal images in the training set, represents the standard deviation of the normal images in the training set, represents the coefficient for adjusting the sensitivity, with a value of 2.6, represents the overall judgment threshold. If the overall anomaly score of the image , then the image is determined to be abnormal; Once the image is determined to be overall abnormal, a local judgment threshold can be further introduced to finely discriminate specific pixels in the image.

[0017] Aiming at the limitations of traditional supervised detection techniques in restricting the model performance due to problems such as high annotation cost, data imbalance, and label errors, based on the characteristics of a large amount of redundant information existing in civil aviation engine borescope images, the present invention proposes a civil aviation engine borescope image damage detection method based on dual-path decoupling and gated memory fusion. This method adopts a dual-encoding - single-decoding architecture, including three stages: feature decoupling, feature fusion, and feature reconstruction. In the feature decoupling stage, a dual-branch feature decoupling architecture is constructed, combined with gradient reversal and mutual exclusion loss, to effectively separate foreground key information from redundant interference information; in the feature fusion stage, an adaptive gating and memory-guided feature fusion module is designed, using gating attention and a dynamic memory bank to adaptively adjust the fusion weights according to feature similarity, thereby enhancing the model's sensitivity to damage features; in the feature reconstruction stage, through a structure-aware reconstruction decoder, jointly optimizing the reconstruction error and the structure loss, enhancing the model's response ability and localization accuracy for the damaged area. Finally, through experiments, it is verified that this method has good adaptability and effectiveness in the civil aviation engine borescope image damage detection task. Description of the Drawings

[0018] AppendixFigure 1 It is the architecture diagram of the hole detection image damage detection model for dual-channel decoupling and gated memory fusion in the present invention.

[0019] Appendix Figure 2 It is the structure diagram of the DenseNet-121 network in the present invention.

[0020] Appendix Figure 3 It is the schematic diagram of the feature fusion module in the present invention.

[0021] Appendix Figure 4 It is the schematic diagram of the anomaly localization result of D2AM-Net on the MVTec AD dataset in the embodiment of the present invention.

[0022] Appendix Figure 5 It is the schematic diagram of the anomaly localization result of D2AM-Net on the BID-DET dataset in the embodiment of the present invention.

[0023] Appendix Figure 6 It is the influence curve diagram of the values of the hyperparameters λ and γ of the loss function on the image-level AUROC result of the model on BID-DET in the embodiment of the present invention.

[0024] Appendix Figure 7 It is the influence curve diagram of the values of the hyperparameters λ and γ of the loss function on the image-level AUROC result of the model on MVTec AD in the embodiment of the present invention.

[0025] Appendix Figure 8 It is the influence curve diagram of the values of the hyperparameters λ and γ of the loss function on the pixel-level AUROC result of the model on MVTec AD in the embodiment of the present invention. Detailed implementation manners

[0026] The following further describes the present invention in conjunction with the accompanying drawings and embodiments.

[0027] To verify the effectiveness and superiority of the method of the present invention in unsupervised anomaly detection tasks. In this example, the publicly available dataset MVTec AD (MVTec Anomaly Detection Dataset) and the civil aviation engine borescope dataset (BID-DET) will be used to verify the method. The specific descriptions of the datasets are as follows: MVTec AD dataset. This dataset is one of the relatively authoritative publicly available datasets in the field of industrial vision at present, covering 15 categories of common industrial objects and materials, including various types of defects in real scenarios, such as scratches, stains, cracks, and deformations. In the setting of the detection task, the defect types such as scratches, cracks, and contaminations included in MVTec AD have certain similarities with the common damage characteristics in civil aviation engine components. Therefore, using this dataset as the verification part dataset for the experiment helps to provide a reference for the actual application in the field of civil aviation engines.

[0028] BID-DET dataset. This dataset is collected from the borescope images obtained from major airlines, covering all key areas of the engine. There are 3396 normal images in the BID-DET dataset, which usually have problems such as uneven illumination, specular interference, local blurring, and complex backgrounds. The abnormal images mainly have six forms of internal engine damage, namely cracks, ablation, chipping, material loss, pits, and warping. The training set for the model consists entirely of normal images, and real damage images will be used in the test phase. The overall setting conforms to the setting of unsupervised anomaly detection tasks. This dataset helps to comprehensively evaluate the generalization ability and detection performance of the model in the borescope inspection environment.

[0029] The experiments carried out in this example are still completed in the Windows 11 operating system environment. The hardware configuration used in the experiments includes an Intel i9-12900K processor and an NVIDIA GeForce RTX 4090 graphics card. The training and testing of all models are implemented based on the PyTorch deep learning framework. The PyTorch version used is 1.13.1, and the CUDA version is 11.6.

[0030] Its hyperparameter settings are shown in Table 1. It should be noted that although hyperparameter tuning is regarded as a key task in the field of deep learning, which depends on a large number of experiments and empirical explorations, since the hyperparameters of some internal calculation items have relatively limited impact on the overall performance under reasonable default values, and the core of this paper is to construct a damage anomaly detection model suitable for the borescope inspection process of civil aviation engines, rather than optimizing the hyperparameters of existing models. Therefore, in order to avoid the sharp increase in the tuning complexity caused by over-expanding the hyperparameter search space, this paper chooses to fix some internal hyperparameters, making the model design more concise and easy to reproduce. In subsequent work, this paper will focus on solving the problem of hyperparameter tuning.

[0031] Table 1 Model hyperparameter settings

[0032] The experimental part of this example mainly covers the following three items: anomaly detection performance comparison experiment, anomaly detection localization comparison experiment, and model ablation experiment. It should be noted that to ensure the consistency and objectivity of the evaluation, except for the performance data of AnoGAN, AE-L2, and AE-SSIM which are cited from the experimental reports of Bergmann et al., the performance results of other methods mainly refer to the experimental results published in their original papers, and some methods may also be cited from the replication or comparative studies of other subsequent works.

[0033] In addition, as the two core modules of the model training strategy, the weights of the mutual exclusion loss and the reconstruction loss directly determine the balance between information separation and reconstruction error control.

[0034] To verify the effectiveness of the unsupervised anomaly detection method proposed in the present invention, the area under the receiver operating characteristic curve (AUROC) is selected as the main evaluation index. AUROC is an important measure to evaluate the classification performance of the model under different decision thresholds, which can comprehensively reflect the overall ability of the model to effectively distinguish normal samples from abnormal samples without relying on specific threshold settings, so it has strong robustness and generalization ability. The closer the AUROC value is to 1, the stronger the discrimination ability of the model.

[0035] AUROC is essentially the calculation of the area under the ROC curve, which can be formally expressed as: (34), In the formula, represents the true positive rate, represents the false positive rate; Among them, the definitions of TPR and FPR are as follows: In the anomaly detection task, the true positive rate (TPR) represents the proportion of samples correctly identified as abnormal among all true abnormal samples; the false positive rate (FPR) represents the proportion of normal samples misidentified as abnormal among all true normal samples.

[0036] In the anomaly localization task, TPR represents the proportion of pixels correctly identified as abnormal among all true abnormal pixels, and FPR represents the proportion of normal pixels misjudged as abnormal among all true normal pixels. The calculation formulas of TPR and FPR are: (35), (36), In the formula, represents the number of true abnormal samples correctly identified as abnormal, represents the number of true abnormal samples misjudged as normal, represents the number of true normal samples misjudged as abnormal, represents the number of true normal samples correctly identified as normal.

[0037] MVTec AD Dataset: Comparative experiments on image-level anomaly detection were conducted on this dataset with seven existing unsupervised image anomaly detection methods, namely AnoGAN, AE-L2, AE-SSIM, RIAD, MKD, P-SVDD, and US. The results are shown in Table 2.

[0038] Table 2 Comparison results of image-level anomaly detection methods on the MVTec AD dataset

[0039] The specific analysis of the experimental results is as follows: First of all, for texture-class samples, D2AM-Net achieved AUC scores of 0.938, 0.982, and 0.998 in the three subclasses of Carpet, Grid, and Leather respectively, significantly outperforming most of the comparison methods. Especially in the Leather subclass, it reached a performance close to 1. This fully demonstrates that D2AM-Net has extremely strong feature expression and discrimination capabilities in dealing with subtle texture changes and local structural perturbations, especially in perceiving tiny damages, which fully meets the requirements of being sensitive to detail anomalies in civil aviation engine borescope inspections. This example believes that the manifestation of this advantage mainly stems from the model's focus on effective features in the feature decoupling stage, and the independence of each encoding branch is strengthened through the mutually exclusive loss, thus fully avoiding the interference of redundant features. At the same time, through the precise adjustment of the feature fusion strategy by the feature fusion module, the model can quickly respond to anomalies when facing tiny perturbations or texture inconsistencies in the image.

[0040] Secondly, for object-class samples, D2AM-Net also showed excellent detection performance, with an overall average AUC reaching 0.911, which is comparable to the existing relatively excellent P-SVDD model and RIAD model. Among them, in categories such as Toothbrush, Hazelnut, and Zipper, D2AM-Net achieved high scores of 0.946, 0.986, and 0.993 respectively, highlighting the model's strong perception ability for anomalies such as local structural damages or deformations of objects. Especially in categories with rich details and complex structures such as Hazelnut and Zipper, the scores of D2AM-Net are close to the optimal, verifying the effectiveness of the joint optimization loss in the decoder in improving the quality of structural detail reconstruction.

[0041] From the above results, it can be seen that D2AM-Net shows high detection performance in most tasks. Especially for scenarios with complex texture backgrounds and fine-grained structural anomalies, it shows strong robustness and generalization ability, and is suitable for damage recognition tasks in the civil aviation engine borescope inspection environment.

[0042] BID-DET Dataset: Based on the above detection results, this method compared the image-level damage detection performance with five unsupervised image anomaly detection methods, namely AnoGAN, AE-L2, AE-SSIM, RIAD, and P-SVDD, on the BID-DET dataset. The results are shown in Table 3. It should be noted that these results are all based on the evaluation completed by reconstructing the model in a unified experimental environment using open-source code or the reproduced code provided by the authors, ensuring the consistency of evaluation conditions and the comparability of results.

[0043] Table 3 Comparison results of image-level damage detection of each method on the borescope image dataset

[0044] The specific analysis of the experimental results is as follows: It can be seen from the experimental results on the BID-DET dataset that D2AM-Net demonstrates excellent performance and high robustness in various damage detection tasks. Compared with traditional reconstruction methods such as AnoGAN, AE-L2, and AE-SSIM, D2AM-Net has achieved a significant improvement in all categories. Especially in anomalies with obvious structural or texture features such as Crack, Dent, and TBCMissing, it shows extremely high detection accuracy. Among them, the model obtained high scores of 0.998 and 0.983 in the Crack and Dent categories respectively, approaching or even exceeding the performance of some discriminative methods such as P-SVDD, indicating that the model has strong expressive ability in capturing key structural information. At the same time, in relatively more challenging categories such as Burn and Missing material, D2AM-Net also obtained scores of 0.985 and 0.899, demonstrating good perception ability for non-significant anomalies. This is attributed to its clear division of effective and redundant information in the feature decoupling stage and the dynamic weight adjustment mechanism achieved through the gated attention mechanism in the fusion stage.

[0045] Overall, the stable performance of D2AM-Net on different types of damage not only verifies the effectiveness of its structural design but also fully demonstrates that this method has good generalization ability and robustness in complex backgrounds, fully verifying the practicality and foresight of this method in the borescope damage detection of civil aviation engines.

[0046] Anomaly localization experiment: The anomaly localization experiment mainly adopts two methods: quantitative evaluation and qualitative evaluation. Since the civil aviation engine borescope image dataset (BID-DET) obtained in this example lacks pixel-level annotations, it is impossible to directly evaluate the precise localization ability of the model in the anomaly area. Therefore, in the anomaly localization experiment of this example, the MVTec AD dataset with pixel-level anomaly annotations is selected to quantitatively evaluate the model to verify the localization accuracy and generalization ability of the D2AM-Net model. At the same time, a qualitative evaluation is carried out on the BID-DET dataset, and the performance of the model in anomaly localization on the dataset is visually observed through the reconstructed loss heat map.

[0047] MVTec AD dataset: In this example, pixel-level AUROC is used as the evaluation index. On this dataset, the proposed method is compared with seven existing unsupervised image anomaly detection methods (including AnoGAN, AE-L2, AE-SSIM, RIAD, MKD, P-SVDD, and US) in terms of anomaly localization performance. The experimental results are shown in Table 4 below and Figure 4 as follows.

[0048] Table 4 Comparison results of image-level damage localization of each method on the MVTec AD dataset

[0049] The specific analysis of the experimental results is as follows: Through the experimental results on the MVTec AD dataset, it can be demonstrated that the D2AM-Net model still exhibits good performance in the anomaly localization task. Although this dataset covers a large number of industrial defect types with different texture and structural characteristics, D2AM-Net still achieves comparable or even better localization accuracy than mainstream advanced methods in most categories. The overall average pixel-level AUROC reaches 0.949, significantly outperforming traditional autoencoder-based methods (such as 0.809 for AE-L2 and 0.863 for AE-SSIM), and there is only a slight gap compared with the currently most powerful P-SVDD (0.956), reflecting its strong comprehensive ability and robustness. Specifically, in categories with complex structures or delicate textures such as Leather, Screw, and Zipper, D2AM-Net achieved high scores of 0.995, 0.981, and 0.993 respectively, demonstrating its high sensitivity in capturing subtle anomalies and edge perturbations. In contrast, the performance of traditional methods in these categories is generally unstable, indicating that while maintaining overall performance consistency, D2AM-Net can better adapt to the anomaly characteristics of different types of samples. In addition, in some object categories such as Toothbrush and Capsule, D2AM-Net also demonstrated the ability to approach or exceed the existing best methods, proving its adaptability in complex structure and non-rigid object anomaly detection scenarios.

[0050] It is worth emphasizing that although the MVTec AD dataset is not specifically designed for the aviation field, the high-resolution images, complex background interference, and various types of industrial defects it contains have certain similarities with the characteristics of civil aviation engine borescope images in terms of image quality and anomaly morphology. Therefore, the excellent localization performance achieved by D2AM-Net on this dataset can, to a certain extent, reflect its potential applicability and generalization ability in the borescope damage detection task. In particular, the stable performance of D2AM-Net in detail structure perception and high background noise is the key ability to solve the key difficulties such as complex texture interference and tiny damage localization in borescope images.

[0051] BID-DET dataset: Since the BID-DET dataset does not have pixel-level labels, the actual effect of the model was verified from a qualitative dimension in this example. From the results of the following reconstructed heatmaps, it can be seen that the high-response regions highly coincide with the actual damage regions, indicating that the model has good anomaly feature perception ability and localization accuracy, and can accurately locate various types of anomaly regions, including structural damages such as cracks, ablation, spalling, and material loss. In addition, from Figure 5It can also be seen that this method demonstrates good generalization ability and robustness in damage scenarios of different forms and scales, further verifying the sensitivity and adaptability of the D2AM-Net model to complex defect structures in actual borescope images. It has high engineering application potential and is suitable as an auxiliary means for manual borescope damage detection.

[0052] In summary, D2AM-Net has certain advantages in overall positioning accuracy and also demonstrates good stability and adaptability in the engine borescope inspection scenario, laying a solid foundation for its unsupervised anomaly detection and localization tasks in civil aviation engine borescope damage.

[0053] (3) Model ablation experiment To verify the effectiveness of each module in the method proposed in this example, ablation experiments were designed and carried out in this example. Key components in the model were removed or replaced one by one to evaluate their impact on the overall performance. Among them, Tables 5 and 6 represent the results of the image-level anomaly detection ablation experiment of the D2AM-Net model on the borescope inspection dataset and the MVTec AD dataset. Table 7 represents the results of the pixel-level anomaly localization ablation experiment of the D2AM-Net model on the MVTec AD dataset.

[0054] Among them, the symbol “–” indicates removing the module or loss function from the complete model. Among them, “–DualEncoder” means removing the redundant information encoding branch, “–GRL” means removing the gradient reversal layer, and “–MutualLoss” means no longer using the mutual exclusion loss; in the feature fusion stage, “–GateAttention” means removing the gated attention mechanism, “–DynamicMemory” means removing the dynamic memory bank module, and “–FusionModule” means removing both the gating mechanism and the memory bank module in the fusion stage; in the feature reconstruction stage, “–StructureLoss” and “–MSELoss” respectively mean only retaining the mean square error loss or the structure loss.

[0055] As the results show, after removing any module or loss function, the model performance decreases to varying degrees, indicating that the design of each proposed module and the combined loss all have a positive effect on the overall anomaly detection performance. The complete D2AM-Net model achieves the best detection effect, verifying the effectiveness and collaborative gain of the multi-stage feature processing and fusion strategy proposed in this example.

[0056] Table 5 Results of the image-level damage detection ablation experiment of D2AM-Net on the borescope inspection damage dataset

[0057] Table 6 Results of ablation experiments on image-level damage detection of D2AM-Net on the MVTec AD dataset

[0058] Table 7 Results of ablation experiments on pixel-level damage localization of D2AM-Net on the MVTec AD dataset

[0059] To enable D2AM-Net to function well in each stage, this example explores the weight hyperparameters of the mutual exclusion loss and the reconstruction loss in its joint optimization loss function. Given that it focuses on structural guidance rather than directly affecting the reconstruction result and usually only requires a relatively small loss weight, the weight value is set in the range of 0.1 to 0.5 in this example. The reconstruction loss directly measures the difference between the reconstructed image and the original image and is the main supervision signal supporting the model's reconstruction ability. Therefore, its weight is set between 0.5 and 1.0 to ensure the dominant position of the reconstruction target.

[0060] The experimental results are as shown in Figure 6 、 7 、8. It can be seen from the experiment that when the mutual exclusion loss weight is set to 0.3 and the reconstruction loss is set to 0.6, the model can achieve the best performance in both image-level and pixel-level anomaly detection tasks. This example believes that the reason for the excellent performance under this configuration is that it can well balance the extraction of effective information and the elimination of redundant information. The mutual exclusion loss of 0.3 is sufficient to effectively constrain the feature overlap between the two encoding branches, prompting the effective information branch to focus on extracting high-semantic features related to the task, while ensuring that the redundant information branch captures background and noise features without causing too much negative interference to the main task; the reconstruction loss of 0.6 provides sufficient optimization strength in the pixel-level and structure-level reconstruction processes, enabling subtle anomaly regions to be clearly reflected in the reconstruction error, thereby improving the accuracy of image anomaly detection and the accuracy of anomaly localization.

[0061] The discussion on the rationality of the hyperparameter configuration is as follows: From the trend of the curve itself, it can be seen that when γ significantly increases to 0.5, in both image-level and pixel-level detections, the performance of some λ values will significantly decline, indicating that too strong mutual exclusion constraints may inhibit the model's capture of key information; similarly, when λ significantly decreases to less than 0.5, the training signal provided by the reconstruction loss is insufficient, and the pixel-level discrimination of the model for anomaly regions is often not ideal enough. Therefore, considering the overall change trend of the curve, the value of the mutual exclusion loss γ is between 0.1–0.5 and the value of the reconstruction loss λThe setting with a value between 0.5 and 1.0 is reasonable. In the experiment γ = 0.3, λ The configuration of = 0.6 achieved relatively stable and excellent detection results, which were consistent with the corresponding theoretical settings.

Claims

1. A borescope image damage detection method combining dual-path decoupling and gated memory, characterized in that: A borescope image damage detection model with dual-path decoupling and gated memory fusion is established, and only unlabeled normal images are used for training. The establishment of the borescope image damage detection model with dual-path decoupling and gated memory fusion includes feature decoupling, feature fusion and feature reconstruction. In the feature decoupling stage, a dual-trunk encoder is used, which serves as an effective information encoding branch and a redundant information encoding branch respectively. In the feature fusion stage, there is an adaptive gating and memory-guided feature fusion module, in which the gating mechanism adaptively adjusts the feature fusion weights of the dual-trunk encoding branch based on the similarity between the features extracted from the effective information encoding branch of the current input and the normal image features stored in the memory bank. The feature content in the memory bank will be dynamically updated according to the training progress of the effective information encoder to adapt to the continuous evolution of the model's representation ability. In the feature reconstruction stage, a structure-aware reconstruction decoder is established, and the fused representation features are sent to the structure-aware reconstruction decoder for image reconstruction. The structure-aware reconstruction decoder adopts a joint optimization loss function.

2. According to claim 1, a borescope image damage detection method combining dual-path decoupling and gated memory fusion is characterized in that: In the feature decoupling stage of the process of establishing the borescope image damage detection model integrating dual-path decoupling and gated memory, the effective information encoding branch is used to extract effective features with high semantic consistency in the image, and the redundant information encoding branch is used to extract redundant or non-discriminative features in the image that are mutually exclusive with the trunk information by introducing a gradient reversal layer. Mutually exclusive loss is also introduced to promote the process of feature decoupling by minimizing the correlation between the output features of different levels of the two encoding branches, ensuring that the effective information branch focuses on the learning of task-related features, while the redundant information branch focuses on capturing features that are not related to anomaly detection.

3. The borescope image damage detection method of dual-path decoupling and gated memory fusion according to claim 2 is characterized in that: DenseNet-121 is used as the backbone network of the dual-trunk encoder. The overall structure of DenseNet-121 includes an initial convolutional layer, a dense block, and a transition layer. The initial convolutional layer is used to extract basic features of the input image at the beginning of the network. Subsequently, the obtained initial features are connected to the dense block part. The dense block uses all the outputs before each layer in the form of channel splicing as the input of the current layer to promote feature reuse and alleviate the problem of gradient disappearance, as shown in the following formula: (1) , In the formula, It consists of Batch Normalization, ReLU and convolution; Next, a transition layer is designed between each dense block and the next dense block to downsample and compress the channels of the feature maps, using 1×1 convolution and average pooling operations to reduce the spatial size and the number of feature maps.

4. The borescope image damage detection method of dual-path decoupling and gated memory fusion according to claim 3 is characterized in that: In the feature decoupled dual-trunk encoder, the effective information encoding backbone adopts the standard DenseNet-121 architecture. Suppose the input image is , the feature representation of the effective information coding backbone extraction is: (2), In the formula, represents the input image, represents an effective information encoder built on the DenseNet-121 architecture, represents the effective feature map of the output, , Represents the height, width and number of channels of the encoded feature map; The redundant information encoding backbone is based on DenseNet-121. A gradient reversal layer GRL is introduced at the output of the backbone. GRL is an identity mapping during forward propagation, but reverses the gradient direction during backward propagation, thereby guiding the backbone to learn features that are mutually exclusive with valid information: (3), In the formula, represents the redundant feature map of the output, , Represents the redundant information encoder built on the DenseNet-121 architecture, represents the gradient reversal layer, and the hyperparameter controlling the gradient reversal strength is , The definition of the gradient reversal layer is as follows, where: Represents the identity matrix, indicating that the gradient is multiplied during back propagation , in order to achieve the reversal of the training direction of the redundant branches, (4), (5), For the gradient reversal strength hyperparameter The setting adopts a dynamic adjustment strategy to balance the stability and adversarial nature during the learning process, as shown in the following formula: (6), In the formula, Represents the training progress, .

5. The borescope image damage detection method of dual-path decoupling and gated memory fusion according to claim 4 is characterized in that: The adaptive gating and memory-guided feature fusion module in the feature fusion part dynamically integrates the features output by different backbone encoders during the reconstruction process to achieve differentiated modeling of different types of inputs. The goal of the feature fusion part is to extract the features extracted by the dual backbone encoders of the feature decoupling module at each spatial position. Perform weighted fusion, and its fusion features It is expressed as: (7), In the formula, Represents the fusion weight map generated by the gating mechanism, which is used to control the proportion of main branch features at each spatial position. , in order to make the gate weight It can be dynamically adjusted according to the standards of normal samples, and a memory library module is introduced, including the following: First, in the training phase, record the current reconstruction loss value , using the exponential moving average (EMA) to record the trend of smoothed losses: (8), In the formula, represents the smoothed mean of the loss at the th moment, represents the smoothing parameter, and its value is 0.9, Next, let the size of the initial memory be , the current number of features in the memory bank is The goal of the update is to gradually and dynamically update the number of features in the memory bank when the model training gradually stabilizes and the loss decreases significantly: (9), In the formula, Representative The smoothed mean loss at each moment is Representative -1 moment loss smoothing mean, Represents the Sigmoid function, which is used to standardize the difference to the [0,1] interval. represents the growth rate, , (10), In the formula, Used to control the slope or sensitivity of the Sigmoid function. , The midpoint of the control function is set to 0.5, which is convenient for triggering the expansion of the memory bank when the loss begins to decrease significantly. The number of prototypes stored in the memory bank at the initial stage needs to ensure that too much noise or unstable features are not written into the memory bank at the beginning of training. Therefore, a moderate scale is selected, and the value is K 0=128, Ultimately, the memory is viewed as a dynamically updated collection of features: (11), In the formula, represent The number of prototypes in the memory bank during each training phase, Represents the typical characteristics of a certain type of valid information in normal samples.

6. The borescope image damage detection method of dual-path decoupling and gated memory fusion according to claim 5 is characterized in that: After each training, each image feature output by the effective information encoder is sent to the decoder for reconstruction one by one. Then, according to the current available capacity of the memory bank, several features with the smallest reconstruction error are selected from the current batch for updating. This form ensures that the stored features have a high degree of reconstruction ability, that is, they have a stronger ability to express normal patterns, thereby improving the accuracy of the gating mechanism in distinguishing abnormalities. Specifically, first, in the current training iteration, a set of candidate features are extracted through batch training. ; Then, the reconstruction error of each feature is calculated separately ; Finally, select the front with the smallest error The features are written into the memory bank. In order to avoid frequent disturbance of stable memory content in the later stage of training, a decay coefficient is also provided to ensure that the update rate of the features in the memory bank is gradually slowed down as the training process is updated. The exponential weighted average (EMA) update strategy is adopted. The effective information features after screening are used to update the prototype in the memory bank: (12), In the formula, represents the number of prototypes in the memory bank, represents the dynamically updated attenuation coefficient, , At the same time, set the dynamic update attenuation coefficient, the formula is: (13), In the formula, Representative EMA decay coefficient at the iteration, Take 0.2, Take 0.8, Take 1e-2, Represents the current iteration number.

7. The borescope image damage detection method of dual-path decoupling and gated memory fusion according to claim 6 is characterized in that: The goal of the gating mechanism is to dynamically weight the features output by the dual encoding branches to achieve adaptive fusion control of effective information and redundant information, including: First, similarity calculation is performed between the effective features at a certain spatial position in the input image and each prototype in the memory library. The similarity between them is calculated using cosine similarity, and the formula is as follows: (14), In the formula, Represents the typical characteristics of a certain type of valid information in normal samples, Represents the position in space The features extracted by the effective encoding branch, Represents the position in space Features and The cosine similarity between prototypes, Next, the similarities of all prototypes are normalized by the Softmax function to obtain the normalized matching weight of each prototype feature at that position: (15), In the formula, represents the normalized matching weight, Then, define the location The overall similarity score is: (16), In order to normalize the score to the interval [0,1], the gating weight is generated by the Sigmoid function: (17), It can be regarded as reflecting the image in The gating coefficient for the feature "importance" or "selective transmission" at the position, if High, which means that when the effective information matches the prototype of the memory bank well, it is given a higher weight; if If it is low, there may be anomalies, and the influence of redundant information needs to be enhanced. The contribution of the redundant information encoding branch is relatively enhanced, resulting in large errors in the abnormal area during reconstruction.

8. The borescope image damage detection method of dual-path decoupling and gated memory fusion according to claim 7 is characterized in that: The structure-aware reconstruction decoder network of the feature reconstruction part includes an upsampling module and an output module. The upsampling module consists of a transposed convolution, batch normalization, an activation function, and a jump connection. The input feature map of the upsampling module is , and are the height and width of the feature map, respectively. After the feature map enters the upsampling module, first, its spatial size is expanded by transposed convolution, and then the transposed feature map is obtained after transposed convolution. : (18), In the formula, Representative layer, Represents the transposed feature map, and its mathematical expression should be: (19), In the formula, Represents the kernel function of the transposed convolution, Represents the stride of the transposed convolution, which is 2 here. represent The spatial position in represent The spatial coordinates in Among them, hyperparameters such as convolution kernel size and stride need to be set according to specific tasks; Next, transpose the convolution output Batch normalization and activation function are performed. The purpose of batch normalization is to slow down the change of feature distribution, so that the input distribution of each layer changes less during the training process, ensure the stability of the input distribution of each layer, and alleviate the problem of gradient disappearance or gradient explosion. For each channel of and variance , followed by a learnable scaling parameter and translation parameters , the specific description is as follows: (20), In the formula, represents the transposed feature map, represents the batch normalization operation, Take 1e-4; ReLU (Rectified Linear Unit) is used as the activation function to provide the network with nonlinear transformation capabilities, and finally the feature map of batch normalization and activation function is processed. It is expressed as follows: (21), In the formula, represents the batch normalization operation, Represents the rectified linear unit activation function, which sets negative values ​​to zero; Through the jump connection, the low-level features of the corresponding spatial size obtained from the effective information branch trunk of the encoder are spliced ​​with the processed features in the channel dimension to obtain the spliced ​​features : (22), In the formula, represents the concatenation operation in the channel dimension, Representative i The jumping characteristics of the effective information branch of the layer, Represents the number of channels of the skip connection in this layer, The concatenated feature map is further processed later. conduct Convolution fusion, and again through batch normalization and activation function, to obtain the output features of the upsampling module : (23), In the formula, represents the batch normalization operation, represents the rectified linear unit activation function, which sets negative values ​​to zero. represent Convolution processing; The output module consists of a final output layer and an image segmentation output part. The final output layer uses Convolution adjusts the number of channels of the feature map of the same size as the original input obtained through multiple upsampling modules to the target output channel number, as follows: (24), In the formula, represents the activation function, Represents the feature map output by the last upsampling module, represents the reconstructed image, represent After the convolution processing is completed, the size of the original image and the reconstructed image should be consistent, and the pixel positions should correspond one to one, so as to facilitate pixel-by-pixel comparison and calculation of errors.

9. The borescope image damage detection method of dual-path decoupling and gated memory fusion according to claim 8 is characterized in that: It also includes the model training phase, in which the total target loss function It is composed of mutually exclusive loss and reconstruction loss, which are used to jointly constrain the learning process of the model, as shown in formula (25): (25), In the formula, represents the mutually exclusive loss function, represents the reconstruction loss function, Represents an adjustable hyperparameter.

Citation Information

Patent Citations

  • Multi-layer feature decoupling optical remote sensing image building extraction method

    CN115731461A

  • Collection, calculation and inspection integrated bridge damage detection method and system

    CN116626059A

  • Abnormity detection method based on twin auto-encoders and bidirectional information depth supervision

    CN116645369A

  • Image anomaly detection method for latent space auto-regression based on memory enhancement

    WO2022095645A1

Cited By

  • Dual-channel wind power generation power prediction method and system

    CN120414534A

  • Product defect detection method and device, medium and program product

    CN120543544A

  • Visual guidance positioning method and system for radiator assembly line

    CN121120783A