A hyperspectral image defect detection method and apparatus

By combining 3D convolution and Transformer coding modules with pyramid pooling, the complexity of defect detection in hyperspectral mural images was solved, achieving multi-scale feature extraction and high-accuracy defect detection.

CN120782773BActive Publication Date: 2025-12-09XIAN INST OF OPTICS & PRECISION MECHANICS CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511285064.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-10
Publication Date
2025-12-09
Estimated Expiration
2045-09-10

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively handle the complex dependencies and multi-scale features in hyperspectral mural images, and traditional methods are insufficient for accurately detecting defects in murals.

Method used

A 3D convolution module is used to perform convolution operations in the spatial and spectral dimensions, combined with a Transformer encoding module for global self-attention modeling, a pyramid pooling module for multi-scale feature extraction, and a fully connected layer to generate a defect probability image.

Benefits of technology

It achieves accurate detection of various defects in mural images, especially in cases with complex spectral dimensions and diverse defect types, demonstrating high accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120782773B_ABST
    Figure CN120782773B_ABST
Patent Text Reader

Abstract

The scheme discloses a hyperspectral image defect detection method and device, the method comprises the following steps: cutting a hyperspectral mural image to be detected into a plurality of subgraphs according to a preset size; inputting the subgraph into a three-dimensional convolution module of a target neural network model to obtain a joint feature graph; weighting the spatial features and spectral features of each position according to the self-attention weight to obtain a weighted feature graph; performing a multi-scale pooling operation on the standard weighted feature graph through a pyramid pooling module of the model to obtain multi-scale features; inputting the spliced feature graph into a fully connected layer of the model to generate an abnormal probability of each pixel point in the subgraph; and splicing the various binary images according to the positions of the various binary images to obtain a defect detection result graph of the hyperspectral mural image. The scheme can accurately detect various defects in the mural through multi-scale feature extraction, and exhibits high accuracy and robustness in the hyperspectral mural image with complex spectral dimensions and various defect types.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present scheme relates to the technical field of computer vision, and particularly relates to a hyperspectral image defect detection method and device. BACKGROUND

[0002] Fresco is a precious cultural heritage. Fresco refers to a kind of painting art form directly painted on the surface of a building, has a long history and contains important cultural value. However, fresco has a large creation area and is exposed to the natural environment for many years, is affected by many factors such as wind, sunlight, humidity change, biological erosion and human damage, and has different degrees of damage on the surface. The abnormal points on the surface of the fresco usually refer to the pigment layer change or damaged part caused by various damage factors. These abnormal points may be color change, brightness change, black dirt, cracking, peeling and the like, which seriously affect the beauty and integrity of the fresco. Through the means of anomaly detection, the defect information of the fresco can be obtained, which has important significance for the repair work of the fresco.

[0003] Hyperspectral imaging technology can capture the reflectance spectrum information of an object in multiple wavebands and is widely used in fields such as remote sensing, medical treatment, agriculture and cultural heritage protection. Hyperspectral images contain rich information in spatial and spectral dimensions. In the process of fresco repair and protection, the hyperspectral imaging technology can be used to identify the subtle degradation and defect information in the fresco that is difficult to identify by the naked eye. However, the types of fresco defect points are various, the sizes are different, and the wavebands of the hyperspectral image are usually up to hundreds. The traditional image processing method and the ordinary convolutional neural network are difficult to effectively process the complex dependence relationship and multi-scale features in the hyperspectral data. SUMMARY

[0004] The present scheme aims to at least solve the technical problems in the prior art. To this end, the first aspect of the present application provides a hyperspectral image defect detection method, which comprises:

[0005] Cutting the hyperspectral fresco image to be detected into a plurality of subgraphs according to a preset size, and recording the position of each subgraph in the image; the image includes spatial dimension information and spectral dimension information;

[0006] Inputting the subgraph into a three-dimensional convolution module of a target neural network model, and performing convolution operation on the subgraph in the spatial dimension and the spectral dimension respectively to obtain a joint feature map, wherein the joint feature map includes spatial features and spectral features;

[0007] The Transformer encoding module of the model is used to globally model self-attention of the joint feature map, to obtain self-attention weights of each position, and the spatial features and the spectral features of each position are weighted processed according to the self-attention weights, to obtain a weighted feature map;

[0008] The weighted feature map is standardized and transformed in the output feature dimension through layer normalization, to obtain a standard weighted feature map;

[0009] The standard weighted feature map is subjected to a multi-scale pooling operation through the pyramid pooling module of the model, to obtain multi-scale features; and the multi-scale features are spliced, to obtain a spliced feature map;

[0010] The spliced feature map is input into the fully connected layer of the model, to generate an anomaly probability of each pixel point in the subgraph, and the subgraph is converted into a binary image according to the anomaly probability and a preset probability threshold; and the binary images are spliced according to the positions thereof, to obtain a defect detection result map of the hyperspectral mural image.

[0011] Optionally, the convolution operation is performed on the subgraph in the spatial dimension and the spectral dimension respectively, to obtain a joint feature map, including:

[0012] A 3D convolution kernel is used to extract spatial features of different scales, to obtain multi-scale spatial features;

[0013] The size of the 3D convolution kernel is changed, to extract multi-scale spectral features from the multi-scale spatial features, to obtain multi-scale spectral features;

[0014] The multi-scale spatial features and the multi-scale spectral features are fused, to obtain a joint feature map.

[0015] Optionally, the Transformer encoding module of the model is used to globally model self-attention of the joint feature map, to obtain self-attention weights of each position, including:

[0016] For each position in the joint feature map, an absolute position encoding is generated in a sine-cosine encoding manner, to obtain an encoded feature map;

[0017] A to-be-analyzed feature of a current pixel region in the encoded feature map is obtained, and global features of other regions except the current pixel region are obtained, to obtain global region features; the to-be-analyzed feature is taken as a query feature, and the global region features are taken as key features;

[0018] Similarity calculation is performed on the query feature and the key feature, to obtain a similarity value;

[0019] linearly mapping the to-be-analyzed feature of the current pixel region to obtain a linear mapping value;

[0020] According to the feature dimension of the similarity value, the linear mapping value, the query feature and the key feature, a self-attention weight of the current pixel region is calculated.

[0021] Optionally, after obtaining the weighted feature map, the method further comprises:

[0022] mapping the low-dimensional feature in the weighted feature map to high dimension through linear transformation of the linear layer of the feedforward network to obtain a high-dimensional feature map;

[0023] filtering the non-key feature of the high-dimensional feature map through the nonlinear activation layer of the feedforward network.

[0024] Optionally, the pyramid pooling module of the model performs a pooling operation of multiple scales on the standard weighted feature map to obtain a multi-scale feature, and the multi-scale feature is spliced to obtain a spliced feature map, comprising:

[0025] According to the dimension of each input feature in the standard weighted feature map, the pyramid pooling module of the model divides the input feature into multiple scales;

[0026] For the features of the multiple scales, different scale pooling windows are used for the pooling operation to obtain different scale pooling features;

[0027] The different scale pooling features are weighted and spliced to obtain a spliced feature map.

[0028] Optionally, before the subgraph is input into the three-dimensional convolution module of the target neural network model, the method further comprises:

[0029] The random noise in the subgraph is removed by using median filtering to obtain a denoised image;

[0030] The normalized image is obtained by normalizing each band in the denoised image;

[0031] The normalized image is processed by dimension reduction using principal component analysis.

[0032] Optionally, the loss function of the target neural network model during training is a cross-entropy loss function.

[0033] The second aspect of the application provides a hyperspectral image defect detection device, the device comprises:

[0034] The cutting module is used for cutting a hyperspectral mural image to be detected into a plurality of subgraphs according to a preset size, and recording the position of each subgraph in the image; the image comprises spatial dimension information and spectral dimension information.

[0035] The convolution module is used for inputting the subgraph into a three-dimensional convolution module of a target neural network model, and performing convolution operation on the subgraph in the spatial dimension and the spectral dimension respectively to obtain a joint feature map, wherein the joint feature map comprises spatial features and spectral features.

[0036] The self-attention module is used for performing global self-attention modeling on the joint feature map through a Transformer encoding module of the model to obtain a self-attention weight of each position, and performing weighted processing on the spatial features and the spectral features of each position according to the self-attention weight to obtain a weighted feature map.

[0037] The standard transformation module is used for performing standardization transformation on the weighted feature map in the output feature dimension through a layer normalization method to obtain a standard weighted feature map.

[0038] The pooling module is used for performing a plurality of scale pooling operations on the standard weighted feature map through a pyramid pooling module of the model to obtain a multi-scale feature, and splicing the multi-scale feature to obtain a spliced feature map.

[0039] The detection module is used for inputting the spliced feature map into a fully connected layer of the model to generate an abnormal probability of each pixel point in the subgraph, converting the subgraph into a binary image according to the abnormal probability and a preset probability threshold, and splicing each binary image according to the position to obtain a defect detection result map of the hyperspectral mural image.

[0040] The third aspect of the present application provides an electronic device, which comprises a processor and a memory, and the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set are loaded and executed by the processor to realize the hyperspectral image defect detection method as described in the first aspect.

[0041] The fourth aspect of the present application provides a computer readable storage medium, which stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set are loaded and executed by the processor to realize the hyperspectral image defect detection method as described in the first aspect.

[0042] The embodiment of the present application has the following beneficial effects:

[0043] The hyperspectral image defect detection method provided by the embodiment of the present application cuts a hyperspectral mural image to be detected into a plurality of subgraphs according to a preset size, records the position of each subgraph in the image; the image includes spatial dimension information and spectral dimension information; the subgraph is input into a three-dimensional convolution module of a target neural network model, and convolution operations are performed on the subgraph in the spatial dimension and the spectral dimension respectively to obtain a joint feature graph, the joint feature graph including spatial features and spectral features; global self-attention modeling is performed on the joint feature graph through a Transformer coding module of the model to obtain self-attention weights of each position, and the spatial features and the spectral features of each position are weighted and processed according to the self-attention weights to obtain a weighted feature graph; the weighted feature graph is standardized and transformed in the output feature dimension through a layer normalization method to obtain a standard weighted feature graph; a plurality of scale pooling operations are performed on the standard weighted feature graph through a pyramid pooling module of the model to obtain multi-scale features; the multi-scale features are spliced to obtain a spliced feature graph; the spliced feature graph is input into a fully connected layer of the model to generate an abnormal probability of each pixel point in the subgraph, and the subgraph is converted into a binary image according to the abnormal probability and a preset probability threshold; the binary images are spliced according to the positions to obtain a defect detection result graph of the hyperspectral mural image. The scheme combines the advantages of a 3D convolutional neural network in spatial-spectral joint feature extraction and the global modeling capability of a Transformer model, and can accurately detect various defects in murals through multi-scale feature extraction. The method performs well in different scale defect detection tasks, especially in hyperspectral mural images with complex spectral dimensions and various types of defects, and exhibits high accuracy and robustness. BRIEF DESCRIPTION OF DRAWINGS

[0044] Figure 1 A step flowchart of a hyperspectral image defect detection method provided by the embodiment of the present application;

[0045] Figure 2 A schematic diagram of a hyperspectral mural image provided by the embodiment of the present application;

[0046] Figure 3 A defect detection result graph provided by the embodiment of the present application;

[0047] Figure 4 A structural block diagram of a hyperspectral image defect detection method provided by the embodiment of the present application. DETAILED DESCRIPTION

[0048] With reference to the drawings of the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application.

[0049] Hereinafter, the terms "first" and "second" are only used for description purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first" and "second" can explicitly or implicitly include one or more features. In the description of the embodiments of the present disclosure, unless otherwise stated, the meaning of "a plurality of" is two or more. In addition, the use of "based on" or "according to" means openness and inclusiveness, because the process, step, calculation or other action "based on" or "according to" one or more stated conditions or values can be based on additional conditions or values beyond the stated values in practice.

[0050] Figure 1 A step flow chart of a hyperspectral image defect detection method provided by the embodiments of the present application.

[0051] As shown in Figure 1 The method comprises the following steps:

[0052] Step 101, cutting the hyperspectral mural image to be detected into a plurality of sub-images according to a preset size, and recording the position of each sub-image in the image; the image comprises spatial dimension information and spectral dimension information.

[0053] Our data is obtained from hyperspectral images acquired by a hyperspectral image acquisition device. Hyperspectral images contain rich information in spatial and spectral dimensions. The data structure of a hyperspectral image is a three-dimensional data cube, containing two spatial dimensions and one spectral dimension.

[0054] Among them, the spatial dimension (x, y): like a normal picture, composed of pixel points, forming a two-dimensional image of an object. The spectral dimension (λ): at each pixel point (x, y), there is a complete and continuous spectral curve, recording the intensity values of hundreds of wavelengths ranging from visible light to infrared light.

[0055] Exemplarily, the number of wavebands of the hyperspectral mural image is 128, and the spatial dimension of the image is 1392x1705.

[0056] The entire hyperspectral mural image to be detected is cut into sub-images according to a fixed size, and the characteristic position of each sub-image in the entire image is recorded, so as to reconstruct after defect detection is completed.

[0057] Figure 2 A schematic diagram of a hyperspectral mural image is provided for an embodiment of the present application.

[0058] As shown in Figure 2 , information of a mural image is collected using a hyperspectral image acquisition device to obtain a hyperspectral mural image. The hyperspectral mural image includes spatial dimension and spectral dimension information.

[0059] As an optional embodiment, before step 102, it further includes:

[0060] Step 201, removing random noise in the sub-image by median filtering to obtain a denoised image;

[0061] Step 202, normalizing each band in the denoised image to obtain a normalized image;

[0062] Step 203, performing dimension reduction processing on the normalized image by principal component analysis.

[0063] In steps 201-203, in order to improve the accuracy of defect detection and reduce the influence of noise on the performance of the model, the present step comprehensively preprocesses the collected hyperspectral mural image, mainly including the following operations: noise removal, normalization and PCA (Principal Component Analysis) dimension reduction.

[0064] Firstly, random noise in the hyperspectral data is removed by traditional image processing methods such as median filtering and mean filtering to ensure the quality of the data. Specifically, the random noise in the hyperspectral image is removed by median filtering, which can preserve the edge and detail information of the image.

[0065] Secondly, the spectral values of each band are normalized to scale the spectral values of all bands to a uniform range [0, 1], ensuring data consistency and helping to improve model training effect. Specifically, since the spectral reflection value range of different bands may differ, each band needs to be normalized. The pixel values of each band are scaled to the [0, 1] interval to ensure the consistency of each band during model training, avoiding the negative impact of some band values being too large or too small on model training. The formula for normalization processing is as follows:

[0066]

[0067] where x, y are pixel positions, and λ is the band value. min denotes the minimum value, and max denotes the maximum value.

[0068] Finally, the PCA dimension reduction technique is used to reduce the redundant features between bands, thereby improving the training efficiency and preventing overfitting. Specifically, the principal component analysis method is used to reduce the dimension of the image to eliminate the correlation between different bands. The principal component analysis technique can effectively reduce redundant features, reduce computational complexity, and prevent multicollinearity problems between features. The number of dimensions after dimension reduction is adjusted according to the experimental results to ensure that important information in the image is retained.

[0069] Step 102, input the subgraph into the three-dimensional convolution module of the target neural network model, and perform convolution operation on the subgraph in the spatial dimension and the spectral dimension respectively to obtain a joint feature map, the joint feature map including spatial features and spectral features.

[0070] The target neural network model includes an input layer, a three-dimensional convolution module, a Transformer encoding module, a pyramid pooling module, a fully connected layer, and an output classification layer.

[0071] The three-dimensional convolution module convolves the hyperspectral data to obtain joint spatial-spectral features and preliminarily extracts feature representations of different bands and spatial positions in the image.

[0072] Specifically, unlike ordinary 2D convolution layers, the 3D convolution kernel can transform in both the "spectral" and "spatial" dimensions. When extracting spatial features, the 3D convolution kernel fixes the spectrum in one dimension, which can be understood as a plurality of 2D convolution kernels extracting features from corresponding multiple ordinary single-channel images and fusing these information.

[0073] This method effectively avoids the deficiency of traditional 2D convolution that cannot capture spectral information, so that the network can consider both the spatial structure and the spectral characteristics of the image.

[0074] As an optional embodiment, the loss function of the target neural network model during training is a cross-entropy loss function.

[0075] During the training of the target neural network model, a part of the hyperspectral mural image is divided into a training set and input into the model for supervised training. The cross-entropy loss function L used in model training is:

[0076]

[0077] where N represents the number of all pixel points in the hyperspectral image, represents the true label of the i-th pixel point, represents the probability that the model predicts the i-th pixel point as a defect.

[0078] ​As an optional embodiment, the sub-graphs in step 102 are respectively subjected to convolution operations in the spatial dimension and the spectral dimension to obtain a joint feature map, including:

[0079] In step 1021, a 3D convolution kernel is used to extract spatial features of different scales to obtain multi-scale spatial features.

[0080] In step 1022, the size of the 3D convolution kernel is changed to extract multi-scale spectral features from the multi-scale spatial features to obtain multi-scale spectral features.

[0081] In step 1023, the multi-scale spatial features and the multi-scale spectral features are fused to obtain a joint feature map.

[0082] In steps 1021-1023, the hyperspectral image is a three-dimensional data cube (height x width x band). A common 2D convolution kernel (size [k, k, 1]) can only slide in the spatial dimension (height and width) and cannot perceive the context information in the spectral dimension at the same time. It regards the same spatial position in different bands as completely independent features, thereby losing the chemical and physical information contained in the spectral curve.

[0083] A 3D convolution kernel (size [k, k, d]) is different, which is a real three-dimensional cube. A local spatial neighborhood and a local spectral neighborhood are covered at the same time in one convolution operation, so that joint spatial-spectral features can be directly extracted.

[0084] First, a 3D convolution kernel with a size of [k, k, 1] is used for spatial convolution. The size of the 3D convolution kernel in the spectral dimension is 1. This convolution kernel does not slide in the spectral dimension, but is fixed in a certain spectral dimension, and performs spatial convolution independently in each band.

[0085] For each band slice, this convolution kernel slides and convolves on it like a common 2D convolution kernel to extract spatial features such as edge features and texture features in the band.

[0086] Then, a 3D convolution kernel with a size of [1, 1, d] is used. The size of this convolution kernel in the spatial dimension is 1x1. This 1x1 convolution kernel does not slide in the spatial dimension, but is fixed in a certain spatial position (x, y coordinates). It slides and convolves along the spectral dimension to fuse information of adjacent multiple bands and extract spectral features.

[0087] In specific implementation, a 3D convolution kernel is used to extract spatial features of different scales to obtain multi-scale spatial features. Then, the size of the convolution kernel is changed to extract multi-scale spectral features from the multi-scale spatial features, and spatial-spectral feature fusion is performed.

[0088] The spatial features and spectral features are extracted by the following 3D convolution calculation formula:

[0089]

[0090] wherein X is the input data feature, Y is the output feature, W is the weight parameter of the 3D convolution kernel, M, N and L are the sizes of different dimensions of the convolution kernel, m, n and l represent different dimension indexes of the convolution kernel, that is, the offset when the convolution kernel moves, and the summation in the formula represents the traversal process of the convolution kernel.

[0091] The module adopts a residual connection mechanism, directly connects between every two convolution layers by a skip connection method, ensures stable gradient back propagation, and effectively alleviates the deep network training difficulty problem.

[0092] Step 1021 ensures that the model can see the spatial features in each band, and step 1022 ensures that the model can see the spectral features at each spatial position. Finally, the output joint feature map, each value in the map is calculated by the information in a piece of spatial region and a spectral interval around it, which is the joint spatial-spectral feature.

[0093] Step 103, modeling the global self-attention of the joint feature map through the Transformer encoding module of the model, obtaining the self-attention weight of each position, and weighting the spatial features and spectral features of each position according to the self-attention weight, obtaining a weighted feature map.

[0094] The Transformer encoding module is used to model the global self-attention based on the preservation of the spatial structure, further capture the long-range dependency relationship between different bands, and enhance the understanding ability of the model to complex spectral and spatial information.

[0095] This step uses the self-attention weight as the mixing coefficient to dynamically and selectively aggregate the information of all positions in the joint feature map into each output position, thereby generating a new enhanced feature map that each point contains global context information.

[0096] According to the self-attention weight, the spatial features and spectral features of each position are weighted and processed to obtain a weighted feature map.

[0097] Subsequently, by evaluating and processing the weighted features of normal samples and abnormal samples, the probability that the pixel point is an abnormal point can be obtained, and by integrating these probabilities globally, the abnormal region distribution of the whole image can be obtained.

[0098] As an optional embodiment, the global self-attention of the joint feature map is modeled by the Transformer encoding module of the model in step 103, and the self-attention weight of each position is obtained, including:

[0099] In step 1031, the absolute position encoding is generated by using the sine-cosine encoding method for each position in the joint feature map, and an encoded feature map is obtained.

[0100] On the basis of 3D convolutional layer feature extraction, the Transformer model is further used for global modeling of the features. Since the Transformer itself does not have position sensing capability, position encoding needs to be performed on the input features. According to the characteristics of the image, we choose a fixed encoding method to encode each part of the image cube using sine and cosine, so as to embed spatial information into the features and ensure that the model can recognize the spatial position.

[0101] In step 1032, the feature to be analyzed of the current pixel region in the encoded feature map is obtained, and the global feature of other regions except the current pixel region is obtained, and a global region feature is obtained. The feature to be analyzed is taken as a query feature, and the global region feature is taken as a key feature.

[0102] In step 1033, similarity calculation is performed on the query feature and the key feature, and a similarity value is obtained.

[0103] In step 1034, linear mapping is performed on the feature to be analyzed of the current pixel region, and a linear mapping value is obtained.

[0104] In step 1035, the self-attention weight of the current pixel region is calculated according to the similarity value, the linear mapping value, the feature dimension of the query feature and the key feature.

[0105] In steps 1032-1035, through the multi-head self-attention mechanism, the model can capture the long-range dependence relationship between different wavebands and different spatial positions. The formula of the self-attention mechanism is as follows:

[0106]

[0107] Where Q, K and V are query, key and value respectively, d k is the dimension of the key.

[0108] Specifically, Q represents the feature to be analyzed of the current region of the image. If the region is an abnormal region, the query feature of the region will be significantly different from the Q distribution of the normal sample.

[0109] K represents the characteristics of other global regions other than the current region, as a reference, similarity calculation is performed with the query characteristics of the current region, so as to measure the abnormal degree of the to-be-tested region;

[0110] V is a value obtained by linearly mapping the to-be-analyzed characteristics of the current pixel region.

[0111] The values of Q, K, V, d k of the current pixel region are substituted into the above formula to obtain the self-attention weight of the current pixel region.

[0112] The above steps compare the characteristics of each pixel point with the global information, thereby enhancing the understanding ability of the model to the global structure of the image and enhancing the understanding ability of the model to the complex spectral-spatial information.

[0113] As an optional embodiment, after obtaining the weighted feature map, the method further comprises:

[0114] Step 301, mapping the low-dimensional characteristics in the weighted feature map to high dimensions through linear transformation of the linear layer of the feedforward network to obtain a high-dimensional feature map;

[0115] Step 302, filtering the non-key characteristics of the high-dimensional feature map through the nonlinear activation layer of the feedforward network.

[0116] In steps 301-302, after the self-attention module, the feedforward network is used for nonlinear mapping of the characteristics, specifically, the low-dimensional characteristics are mapped to high dimensions through linear transformation of the linear layer, the representation ability of the characteristics is improved, and then a nonlinear activation layer is used to filter out irrelevant characteristics, so that the key characteristics are highlighted.

[0117] Step 104, performing standardization transformation on the weighted feature map in the output feature dimension through the layer normalization method to obtain a standard weighted feature map.

[0118] After the self-attention mechanism performs weighted summation, the obtained “weighted feature map” may have problems of internal covariate shift and different feature scales, and layer normalization can solve these problems.

[0119] For the feature vector of any spatial position (i, j) on the weighted feature map, the layer normalization separately calculates the mean and variance of the feature vector in each channel dimension, and uses them to standardize the vector.

[0120] Through the layer normalization method, the output feature dimension of the image is standardized and transformed, which can stabilize the gradient and accelerate the convergence of the model.

[0121] Step 105, performing a pooling operation of multiple scales on the standard weighted feature map through a pyramid pooling module of the model to obtain multi-scale features; and splicing the multi-scale features to obtain a spliced feature map.

[0122] After feature extraction through 3D convolution and Transformer, a Spatial Pyramid Pooling (SPP) module is used to process defects of different sizes in the image.

[0123] For defect detection tasks, the morphology and size of defects vary greatly. Large defects may require a larger receptive field and more global context to confirm. Small defects require fine, local features to capture their subtle traces. If the network only uses a single scale of pooling, then: for large defects, this pooling scale may be too small to effectively integrate the information of the entire defect region. For small defects, this pooling scale may be too large, mixing in a lot of background noise and diluting the features of the defects themselves.

[0124] The Spatial Pyramid Pooling module generates a spatial pyramid, divides the input features into different scales according to the specific dimensions of the input features (starting from 1x1, not exceeding the maximum size of the features), performs a pooling operation, and fuses the pooled features, thereby effectively processing defect regions of different sizes. Through multi-scale feature extraction, the network can accurately respond to defects of different sizes and morphologies.

[0125] As an optional embodiment, step 105 includes:

[0126] Step 1051, dividing the input features into multiple scales according to the dimensions of the input features in the standard weighted feature map through the pyramid pooling module of the model;

[0127] Step 1052, for the features of the multiple scales, respectively performing a pooling operation using different scale pooling windows to obtain pooled features of different scales;

[0128] Step 1053, performing weighted splicing on the pooled features of different scales to obtain a spliced feature map.

[0129] In steps 1051-1053, the pyramid pooling module performs pooling through different scales (specifically 1x1, 2x2, and 4x4) to extract global, local, and detailed information, and then splices the pooling results of each scale before sending them to the fully connected layer. This multi-scale processing mechanism enables the model to have good detection capability for defects of different scales such as cracks and shedding.

[0130] Step 106, input the spliced feature map into the full connection layer of the model to generate the abnormal probability of each pixel point in the subgraph, and convert the subgraph into a binary image according to the abnormal probability and a preset probability threshold; splice each binary image according to the position to obtain a defect detection result map of the hyperspectral mural image.

[0131] The spliced feature map is mapped into a two-dimensional defect probability map via a full connection network, representing the probability of each pixel point in the input image being abnormal, and the output shape is 1xHxW. In the present application, the output layer activation function adopts a Softmax function to realize the classification and prediction of multiple defects.

[0132] A suitable probability threshold is set in advance, for example, set to 0.85.

[0133] The subgraph is converted into a binary image according to the abnormal probability and the preset probability threshold. Specifically, the pixel points with an abnormal probability greater than or equal to the probability threshold can be determined as defect points and displayed as white, and the pixel points with an abnormal probability less than the probability threshold are displayed as black, so that the white pixel points are defect pixels.

[0134] Finally, each binary image is aligned and spliced according to the position, and multiple wave bands or multiple image regions are merged into a complete image to obtain a defect detection result map consistent with the size of the hyperspectral mural image.

[0135] Figure 3 A defect detection result map provided by the embodiment of the present application.

[0136] As shown in Figure 3 , the white part of the figure is a defect pixel point, and the black part is a normal pixel point.

[0137] In summary, the hyperspectral image defect detection method provided by the embodiment of the application cuts a hyperspectral mural image to be detected into a plurality of subgraphs according to a preset size, records the position of each subgraph in the image; the image includes spatial dimension information and spectral dimension information; the subgraph is input into a three-dimensional convolution module of a target neural network model, convolution operations are performed on the subgraph in the spatial dimension and the spectral dimension respectively, a joint feature map is obtained, and the joint feature map includes spatial features and spectral features; global self-attention modeling is performed on the joint feature map by a Transformer coding module of the model, self-attention weights of each position are obtained, the spatial features and the spectral features of each position are weighted and processed according to the self-attention weights, a weighted feature map is obtained; the weighted feature map is standardized and transformed in the output feature dimension by a layer normalization method, a standard weighted feature map is obtained; a plurality of scale pooling operations are performed on the standard weighted feature map by a pyramid pooling module of the model, a multi-scale feature is obtained; the multi-scale feature is spliced to obtain a spliced feature map; the spliced feature map is input into a fully connected layer of the model, an abnormal probability of each pixel point in the subgraph is generated, and the subgraph is converted into a binary image according to the abnormal probability and a preset probability threshold; the binary images are spliced according to the positions, and a defect detection result image of the hyperspectral mural image is obtained. The scheme combines the advantages of a 3D convolutional neural network in spatial-spectral joint feature extraction and the global modeling capability of a Transformer model, can accurately detect various defects in murals through multi-scale feature extraction, and has high accuracy and robustness in different scale defect detection tasks, especially in hyperspectral mural images with complex spectral dimensions and various defects. The method is not only suitable for mural restoration in cultural heritage protection, but also has strong generalizability and can be applied to other hyperspectral image processing fields.

[0138] Figure 4 is a structural block diagram of a cross-scale spatial target detection device provided by the embodiment of the application. As shown in the figure, Figure 4 The device 400 includes:

[0139] The cutting module 401 is configured to cut a hyperspectral mural image to be detected into a plurality of subgraphs according to a preset size, and record the position of each subgraph in the image; the image includes spatial dimension information and spectral dimension information.

[0140] The convolution module 402 is configured to input the subgraph into a three-dimensional convolution module of a target neural network model, perform convolution operations on the subgraph in the spatial dimension and the spectral dimension respectively, obtain a joint feature map, and the joint feature map includes spatial features and spectral features.

[0141] The self-attention module 403 is configured to perform global self-attention modeling on the joint feature map by a Transformer encoding module of the model to obtain self-attention weights of each position, and perform weighted processing on the spatial features and the spectral features of each position according to the self-attention weights to obtain a weighted feature map.

[0142] The standard transformation module 404 is configured to perform standardization transformation on the weighted feature map in the output feature dimension by a layer normalization method to obtain a standard weighted feature map.

[0143] The pooling module 405 is configured to perform a plurality of scale pooling operations on the standard weighted feature map by a pyramid pooling module of the model to obtain multi-scale features, and splice the multi-scale features to obtain a spliced feature map.

[0144] The detection module 406 is configured to input the spliced feature map into a fully connected layer of the model to generate an anomaly probability of each pixel point in the subgraph, convert the subgraph into a binary image according to the anomaly probability and a preset probability threshold, and splice each binary image according to the position to obtain a defect detection result map of the hyperspectral mural image.

[0145] As to the device in the above-mentioned embodiments, the specific manners in which various modules perform operations have been described in detail in the embodiments of the method, and will not be described in detail here.

[0146] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, storage, databases, or other media in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in many forms such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0147] Other embodiments of the disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the features disclosed herein. This application is intended to cover any variations, uses, or adaptations of the disclosure that are deemed to fall within the general principles of the disclosure and include commonly known or customary practice in the art. The specification and examples are to be considered exemplary only, with the true scope and spirit of the disclosure being indicated by the following claims.

[0148] It should be understood that the present disclosure is not limited to the precise structures herein described and illustrated in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the claims that follow.

Claims

1. A method for detecting defects in hyperspectral images, characterized in that, The method includes: The hyperspectral mural image to be detected is cut into multiple sub-images according to a preset size, and the position of each sub-image in the image is recorded; the image includes spatial dimension information and spectral dimension information. The subgraph is input into the three-dimensional convolution module of the target neural network model, and convolution operations are performed on the subgraph in the spatial dimension and the spectral dimension to obtain a joint feature map, which includes spatial features and spectral features. The joint feature map is modeled globally using the Transformer encoding module of the model to obtain the self-attention weight at each position. The spatial features and spectral features at each position are then weighted according to the self-attention weight to obtain a weighted feature map. By using layer normalization, the weighted feature map is standardized in the output feature dimension to obtain a standard weighted feature map. The standard weighted feature map is pooled at multiple scales using the pyramid pooling module of the model to obtain multi-scale features; the multi-scale features are then concatenated to obtain a concatenated feature map. The stitched feature map is input into the fully connected layer of the model to generate the anomaly probability of each pixel in the sub-image. The sub-image is converted into a binary image based on the anomaly probability and a preset probability threshold. The binary images are stitched together according to their positions to obtain the defect detection result map of the hyperspectral mural image. The step of performing convolution operations on the sub-graphs in both the spatial and spectral dimensions to obtain a joint feature map includes: 3D convolution kernels are used to extract spatial features at different scales to obtain multi-scale spatial features; By changing the size of the 3D convolution kernel, multi-scale spectral feature extraction is performed on the multi-scale spatial features to obtain multi-scale spectral features; The multi-scale spatial features and the multi-scale spectral features are fused to obtain a joint feature map; The step of performing global self-attention modeling on the joint feature map using the Transformer encoding module of the model to obtain the self-attention weights at each location includes: For each position in the joint feature map, an absolute position code is generated using sine and cosine coding to obtain the coded feature map; Obtain the features to be analyzed in the current pixel region of the encoded feature map, obtain the global features of other regions outside the current pixel region, and obtain the global region features; use the features to be analyzed as query features and the global region features as key features; The similarity between the query feature and the key feature is calculated to obtain a similarity value; The features to be analyzed in the current pixel region are linearly mapped to obtain linear mapping values; The self-attention weight of the current pixel region is calculated based on the similarity value, the linear mapping value, the query feature, and the feature dimension of the key feature; After obtaining the weighted feature map, the following steps are also included: The low-dimensional features in the weighted feature map are mapped to a high-dimensional feature map by the linear transformation of the linear layer of the feedforward network. Non-critical features of the high-dimensional feature map are filtered through the non-linear activation layer of the feedforward network.

2. The method according to claim 1, characterized in that, The process involves performing multi-scale pooling operations on the standard weighted feature map using the pyramid pooling module of the model to obtain multi-scale features; then concatenating the multi-scale features to obtain a concatenated feature map, including: The pyramid pooling module of the model divides the input features into multiple scales according to the dimension of each input feature in the standard weighted feature map. For the features at the multiple scales, pooling operations are performed using pooling windows of different scales to obtain pooled features at different scales. The pooling features at different scales are weighted and concatenated to obtain a concatenated feature map.

3. The method according to claim 1, characterized in that, Before inputting the subgraph into the 3D convolutional module of the target neural network model, the method further includes: Median filtering is used to remove random noise from the sub-image to obtain a denoised image; Normalize each band in the denoised image to obtain a normalized image; Principal component analysis was used to reduce the dimensionality of the normalized image.

4. The method according to claim 1, characterized in that, The loss function of the target neural network model during training is the cross-entropy loss function.

5. A hyperspectral image defect detection device, characterized in that, The device includes: The cutting module is used to cut the hyperspectral mural image to be detected into multiple sub-images according to a preset size, and record the position of each sub-image in the image; the image includes spatial dimension information and spectral dimension information. A convolution module is used to input the subgraph into the three-dimensional convolution module of the target neural network model, and to perform convolution operations on the subgraph in the spatial dimension and the spectral dimension to obtain a joint feature map, wherein the joint feature map includes spatial features and spectral features. The self-attention module is used to perform global self-attention modeling on the joint feature map through the Transformer encoding module of the model, obtain the self-attention weight at each position, and perform weighted processing on the spatial features and spectral features at each position according to the self-attention weight to obtain a weighted feature map. The standard transformation module is used to perform a standard transformation on the weighted feature map in the output feature dimension by using layer normalization to obtain a standard weighted feature map. The pooling module is used to perform multi-scale pooling operations on the standard weighted feature map through the pyramid pooling module of the model to obtain multi-scale features; and to concatenate the multi-scale features to obtain a concatenated feature map. The detection module is used to input the stitched feature map into the fully connected layer of the model, generate the abnormal probability of each pixel in the sub-image, convert the sub-image into a binary image according to the abnormal probability and a preset probability threshold, and stitch the various binary images according to their positions to obtain the defect detection result map of the hyperspectral mural image. The convolution module is specifically used for: 3D convolution kernels are used to extract spatial features at different scales to obtain multi-scale spatial features; By changing the size of the 3D convolution kernel, multi-scale spectral feature extraction is performed on the multi-scale spatial features to obtain multi-scale spectral features; The multi-scale spatial features and the multi-scale spectral features are fused to obtain a joint feature map; The self-attention module is specifically used for: For each position in the joint feature map, an absolute position code is generated using sine and cosine coding to obtain the coded feature map; Obtain the features to be analyzed in the current pixel region of the encoded feature map, obtain the global features of other regions outside the current pixel region, and obtain the global region features; use the features to be analyzed as query features and the global region features as key features; The similarity between the query feature and the key feature is calculated to obtain a similarity value; The features to be analyzed in the current pixel region are linearly mapped to obtain linear mapping values; The self-attention weight of the current pixel region is calculated based on the similarity value, the linear mapping value, the query feature, and the feature dimension of the key feature; The device is also used for: The low-dimensional features in the weighted feature map are mapped to a high-dimensional feature map by the linear transformation of the linear layer of the feedforward network. Non-critical features of the high-dimensional feature map are filtered through the non-linear activation layer of the feedforward network.

6. An electronic device, characterized in that, The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set, or instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the hyperspectral image defect detection method as described in any one of claims 1-4.

7. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, the at least one program, the code set, or instruction set is loaded and executed by a processor to implement the hyperspectral image defect detection method as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Hyperspectral image classification method based on multi-scale feature attention

    CN118247588A

  • Hyper-spectral feature extraction method and apparatus, and storage medium

    WO2023173884A1