A frequency domain enhanced in-situ hyperspectral feature extraction and classification method

By combining fractional Fourier transform with a local-global spectral attention mechanism, the problems of high computational complexity and poor noise robustness in medical hyperspectral image classification are solved, achieving efficient and accurate pathological tissue classification, which is suitable for real-time analysis in clinical settings.

CN122265719APending Publication Date: 2026-06-23BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING INST OF TECH
Filing Date
2026-03-24
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

Existing medical hyperspectral image classification methods struggle to balance local details with global context when processing non-stationary pathological signals, resulting in high computational complexity, poor noise robustness, and an inability to meet the needs of real-time clinical diagnosis.

Method used

By combining fractional Fourier transform with a local-global spectral attention mechanism, and through adaptive spatial-frequency feature extraction and multi-scale feature fusion, deep feature extraction and efficient classification of medical hyperspectral images are achieved.

Benefits of technology

It significantly improves the robustness of processing non-stationary pathological signals, reduces computational complexity, meets the needs of real-time clinical diagnosis, improves classification accuracy and generalization ability, and solves the class imbalance problem.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122265719A_ABST
    Figure CN122265719A_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of medical image processing and artificial intelligence, and particularly relates to a frequency domain enhanced in-situ biological hyperspectral feature extraction and classification method. By fusing the adaptive spatial-frequency feature extraction capability of fractional Fourier transform and the local-global spectral attention mechanism, the medical hyperspectral image can be deeply featured and efficiently classified, and the rapid and accurate diagnosis of early pathological lesions can be realized. The method is particularly suitable for real-time analysis of complex non-stationary pathological signals and multi-scale spectral-spatial information in a clinical environment, and provides accurate lesion classification and diagnosis assistance for doctors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of medical image processing and artificial intelligence technology, and in particular relates to a frequency domain enhanced method for in-situ hyperspectral feature extraction and classification of biological features. Background Technology

[0002] Rapid pathological classification of brain tumors is crucial for successful surgery and patient prognosis. Accurate intraoperative determination of the nature and boundaries of a brain tumor helps surgeons precisely remove diseased tissue while preserving as much normal tissue as possible, reducing the risk of postoperative complications. Rapid pathological classification not only improves real-time decision-making during surgery but also effectively shortens surgical time. However, traditional methods are limited by time, accuracy, and subjective judgment, making them unsuitable for complex cases. Therefore, more efficient and reliable technologies are urgently needed to assist intraoperative decision-making.

[0003] Medical hyperspectral imaging (MHSI) can simultaneously capture spatial structural information and spectral data across dozens to hundreds of consecutive bands. Compared to traditional RGB images, it covers a wider electromagnetic spectrum, making it invaluable in clinical applications such as the detection of occult lesions, tissue biochemical identification, and differentiation of pathological patterns. It is particularly crucial for the diagnosis of early-stage lesions—for example, in gastric cancer, the five-year survival rate for early-stage patients exceeds 90%, while the rate for late-stage patients, even with surgery, is less than 30%. The differences in spectral characteristics between pathological and healthy tissues provide a core basis for early identification. MHSI demonstrates significantly higher sensitivity in detecting these differences than traditional visual and histological analysis, becoming a powerful tool for advancing precision diagnosis.

[0004] However, medical hyperspectral image classification still faces many technical challenges: on the one hand, pathological signals have non-stationary characteristics, with high-frequency pathological details and low-frequency global structures intertwined, making them difficult to separate effectively; on the other hand, hyperspectral data has high dimensionality and strong correlation between bands, which can easily lead to the "curse of dimensionality," and the computational complexity of traditional methods increases twice when processing high-dimensional data, making it difficult to meet the needs of real-time clinical diagnosis.

[0005] Currently, medical hyperspectral image classification methods are mainly divided into traditional machine learning methods and deep learning methods. Traditional machine learning methods, such as Support Vector Machines (SVM) and Random Forests (RF), rely on manual dimensionality reduction and feature extraction using Principal Component Analysis (PCA), ignoring spatial information. They perform poorly in classifying heterogeneous tissue regions and are limited by subjective experience, resulting in a high risk of error. The emergence of deep learning methods has significantly improved classification capabilities: Convolutional Neural Networks (CNNs) extract local spectral spatial features through three-dimensional convolutional kernels, but their fixed receptive field limits the modeling of non-local spectral correlations and easily introduces grid artifacts and parameter redundancy; the Transformer architecture models long-range dependencies through a self-attention mechanism, but it suffers from two major drawbacks—insufficient robustness to sensor noise and non-stationary pathological signals, leading to increased errors in identifying local spectral variations; and the computational complexity of self-attention increases quadratically with sequence length, resulting in extremely high computational costs for high-dimensional MHSI.

[0006] Frequency domain methods have attracted attention due to their excellent noise robustness. Traditional Fourier transform (FT) achieves noise suppression and feature extraction through signal frequency domain transformation. However, fixed basis functions lack adaptability to non-stationary pathological signals, easily leading to global over-smoothing and loss of local details, and failing to dynamically balance spatial and frequency domain features. Although existing methods have made many attempts in spectral-spatial feature modeling, they have not yet effectively solved the problem of balancing non-stationary signal processing, computational efficiency, and classification accuracy. A more efficient and robust technical solution is urgently needed.

[0007] In summary, accurate and efficient classification of medical hyperspectral images plays a crucial role in early lesion diagnosis and patient prognosis. While traditional methods and existing deep learning models can assist in diagnosis to some extent, they still suffer from limitations in accuracy, computational complexity, poor noise robustness, and difficulty in balancing local details with global context. The fractional Fourier transform (FrFT), with its adjustable fractional-order parameters, achieves a continuous transition in the spatial frequency domain, offering new possibilities for non-stationary signal processing. Combined with a local-global spectral attention mechanism, it can effectively capture multi-scale features. The fusion of these two approaches holds promise for overcoming current technological bottlenecks and achieving efficient and accurate classification of medical hyperspectral images. Summary of the Invention

[0008] Based on the above analysis, this application provides a frequency-domain enhanced in-situ hyperspectral feature extraction and classification method for biological organisms. By fusing the adaptive spatial-frequency feature extraction capability of fractional Fourier transform with a local-global spectral attention mechanism, it can perform deep feature extraction and efficient classification of medical hyperspectral images, enabling rapid and accurate diagnosis of early pathological lesions. This method is particularly suitable for real-time analysis of complex non-stationary pathological signals and multi-scale spectral-spatial information in clinical settings, providing doctors with accurate lesion classification and diagnostic support for decision-making.

[0009] To achieve the above objectives, the first technical solution of this application discloses a frequency-domain enhanced in-situ hyperspectral feature extraction and classification model for biological organisms, which includes the following along the data processing direction:

[0010] Multi-scale patch embedding module: used to receive target medical hyperspectral images and generate multi-scale feature sequences;

[0011] Feature extraction module: connected to the multi-scale patch embedding module, it consists of alternating fractional domain feature extraction module FrTrans and local-global spectral attention module L-GSA, used for deep feature extraction of the multi-scale feature sequence;

[0012] Patch merging and cross-scale fusion module: connected to the feature extraction module, used for multi-scale fusion and spatial downsampling of deep features; and

[0013] Feature fusion and classification module: connected to the patch merging and cross-scale fusion module, used to generate pathological tissue classification results based on fused features;

[0014] The FrTrans module includes an adaptive fractional Fourier transform unit configured with a learnable fractional-order parameter α, used to achieve a continuous transition between the spatial domain and the frequency domain through a rotation angle φ = α·π / 2.

[0015] The random frequency pattern selection unit is used to randomly select K frequency components from N frequency components of the input feature to reduce computational complexity.

[0016] A frequency domain attention unit is used to calculate attention weights for the selected K frequency components and output frequency domain enhancement features;

[0017] The L-GSA module includes a Local Spectral Attention (LSA) submodule and a Global Spectral Attention (GSA) submodule. The LSA submodule uses a shift-and-merge sliding window mechanism to capture local correlations of spectral channels, while the GSA submodule uses a dual-channel clustering mechanism to capture long-range dependencies of spectral channels.

[0018] Furthermore, the adaptive fractional Fourier transform unit also includes a regularization constraint unit, used to pass the objective function. The fractional-order parameter α is constrained to fluctuate around 0.5; the range of the learnable fractional-order parameter α is [0.05, 0.95].

[0019] Furthermore, the LSA submodule includes: a window partitioning unit for partitioning the spectral channel into windows of size P; a shift window unit for shifting the window along the channel dimension by P / 2 units; and a feature fusion unit for fusing the attention outputs of the original window and the shift window.

[0020] The GSA submodule includes: a dual-channel clustering unit for clustering spectral channels into Z cluster centers to generate cluster weights; and a global attention calculation unit for calculating long-range dependencies based on the cluster centers.

[0021] Furthermore, the feature fusion and classification module includes:

[0022] Pathology Focused Attention Unit: Used to generate attention weights through a linear layer and a Sigmoid activation function, and then element-wise summed with the input features to enhance local discriminative details;

[0023] Global pooling unit: used for adaptive average pooling of weighted features;

[0024] Classification Header Unit: Used to map the pooled global feature vector to pathological tissue categories.

[0025] The second technical solution of this application discloses a frequency-domain enhanced in-situ hyperspectral feature extraction and classification method for biological organisms, which adopts the above-mentioned classification model and includes the following steps:

[0026] The target medical hyperspectral image to be classified is input into the multi-scale patch embedding module to generate a multi-scale feature sequence.

[0027] The multi-scale feature sequence is input into the feature extraction module and processed alternately by the FrTrans module and the L-GSA module to obtain a deep feature representation. In the FrTrans module, an adaptive fractional Fourier transform is performed on the features based on a learnable fractional-order parameter α to achieve a continuous transition in the spatial frequency domain. A random frequency mode selection mechanism is used to randomly select K frequency components from N frequency components for attention calculation, reducing computational complexity from... Down to In the L-GSA module, local spectral correlations are captured through a shift-merging sliding window mechanism, and long-range spectral dependence is captured through a dual-channel clustering mechanism.

[0028] The deep feature representation is input into the patch merging and cross-scale fusion module, and multi-scale feature fusion is performed by calculating weights based on feature entropy; and

[0029] The fused features are input into the feature fusion and classification module, and pathological tissue classification results are generated after the pathological focus attention enhancement is used to identify details.

[0030] Furthermore, the classification model is obtained through the following training steps: inputting the medical hyperspectral image training dataset into the classification model, and optimizing the network parameters, including the learnable fractional-order parameter α, through the backpropagation algorithm; wherein, a weighted combination of label smoothing cross-entropy loss and focus loss is used as the total loss function.

[0031] Furthermore, the weights of the label smoothing cross-entropy loss and the focus loss in the total loss function are 0.7 and 0.3, respectively; the learnable fractional-order parameter α is iteratively updated through the backpropagation algorithm and through L... α The regularization constraint fluctuates around 0.5.

[0032] Furthermore, in the K frequency components selected by the random frequency pattern selection mechanism, the value of K ranges from 32 to 128; the activation function used by the frequency domain attention unit is the tanh function.

[0033] as well as,

[0034] A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, it implements the frequency-domain enhanced biological in-situ hyperspectral feature extraction and classification method, or constructs the frequency-domain enhanced biological in-situ hyperspectral feature extraction and classification model.

[0035] A frequency-domain enhanced in-situ hyperspectral feature extraction and classification device for biological organisms, comprising:

[0036] The memory is used to store computer programs; the processor is used to control the frequency-domain enhanced in-situ hyperspectral feature extraction and classification model to perform classification operations when executing the computer program.

[0037] Compared with the prior art, the present invention has the following significant advantages:

[0038] (1) The adaptive spatial frequency decomposition capability significantly improves the robustness of non-stationary pathological signal processing.

[0039] This invention achieves a continuous transition and adaptive rotation between the spatial and frequency domains through a learnable fractional-order parameter α, enabling dynamic separation of high-frequency pathological details from low-frequency global trends. Compared to traditional Fourier transforms with fixed basis functions, this invention effectively overcomes the problems of global oversmoothing and loss of local details, significantly enhancing the model's ability to perceive non-stationary pathological signals (such as subtle spectral differences between early-stage cancerous tissue and normal tissue). Regularization constraints cause α to fluctuate around 0.5, ensuring an optimal balance between spatial and frequency representations, exhibiting excellent noise robustness in complex pathological scenarios such as cholangiocarcinoma and precancerous lesions of the stomach.

[0040] (2) The computational complexity is significantly reduced, meeting the needs of real-time clinical diagnosis.

[0041] By employing a synergistic design of a random frequency pattern selection mechanism and a dual-channel clustering mechanism, this invention reduces the computational complexity of traditional self-attention from... to This reduces the complexity of global spectral attention from Down to The model requires only 5.791 GFLOPs of computation and 8.792M parameters, which is the lowest among all mainstream methods. It supports real-time inference and rapid intraoperative decision-making in clinical settings, effectively solving the technical bottleneck of the traditional Transformer architecture, which has extremely high computational costs and difficulty in meeting real-time requirements when processing high-dimensional medical hyperspectral data.

[0042] (3) Multi-scale spectral-spatial feature collaborative modeling improves classification accuracy and generalization ability

[0043] The Local-Global Spectral Attention (L-GSA) module captures fine-grained local spectral correlations through a shift-and-merge sliding window mechanism, while simultaneously modeling long-range global dependencies through a dual-channel clustering mechanism. These two mechanisms work together to achieve in-depth mining of multi-scale spectral-spatial information. Combining cross-scale fusion with a pathological-focused attention mechanism, the model can adaptively enhance discriminative pathological features and suppress background noise. Experimental results show that the proposed method achieves a classification accuracy of over 93% and an AUC exceeding 98% on various pathological datasets, significantly outperforming traditional CNNs and standard Transformer methods. Furthermore, it demonstrates excellent generalization ability through five-fold cross-validation, providing reliable technical support for the accurate identification of early lesions.

[0044] (4) Effectively solves the problems of category imbalance and scarce annotations.

[0045] To address the common problem of an imbalance in the number of normal tissue and lesion samples in medical hyperspectral images, this invention employs a weighted combination of label smoothing cross-entropy loss and focal loss as the training objective function. This effectively alleviates the model's bias towards the majority class and improves the sensitivity to rare lesion types. Simultaneously, through data augmentation strategies and adaptive optimization of learnable parameters, the model maintains stable feature extraction capabilities even with limited labeled data, reducing reliance on large-scale labeled medical data and better aligning with real-world clinical applications.

[0046] In summary, this invention achieves high-precision and robust classification of medical hyperspectral images by deeply fusing fractional Fourier transform with local-global spectral attention, while maintaining extremely low computational overhead. This provides an effective technical solution for the rapid and accurate identification of intraoperative pathological tissues and the automated diagnosis of early lesions. Attached Figure Description

[0047] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0048] Appendix Figure 1 This is a diagram of the overall model framework;

[0049] Appendix Figure 2 Structure diagram of the feature extraction module

[0050] Appendix Figure 3 This is a structural diagram of the local spectral attention module;

[0051] Appendix Figure 4 This is a structural diagram of the global spectral attention module. Detailed Implementation

[0052] To make the technical problems solved, the technical solutions, and the beneficial effects of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0053] To achieve the above objectives, the first embodiment of this application discloses a frequency-domain enhanced in-situ hyperspectral feature extraction and classification model for biological organisms, as shown in the attached figure. Figure 1 As shown, it includes the following along the data processing direction:

[0054] 1. Multi-scale patch embedding module: used to receive target medical hyperspectral images and generate multi-scale feature sequences.

[0055] In this process, based on the purpose of deep learning processing, the target medical hyperspectral image first needs to be preprocessed into a medical hyperspectral image tensor. (Where, batch size B, spatial dimensions H×W, and spectral channels D), then multi-scale features are extracted through dual-path convolution:

[0056] Where Conv(4×4) and Conv(8×8) are 4×4 and 8×8 convolutions with a stride of 4, Interp is the bilinear interpolation alignment space size, and Concat is concatenated along the channel dimension to obtain... Subsequently, it unfolds into a sequence in space. (Number of tokens).

[0057] 2. Feature extraction module: Connected to the multi-scale patch embedding module, the fractional domain feature extraction module FrTrans and the local-global spectral attention module L-GSA are used to perform deep feature extraction on the multi-scale feature sequence.

[0058] The FrTrans module includes an adaptive fractional Fourier transform unit configured with a learnable fractional-order parameter α, used to transform the Fourier transform by a rotation angle. A continuous transition between the spatial and frequency domains can be achieved. Among them, a learnable fractional-order parameter α (with a value range of [0.05, 0.95]) can be dynamically adjusted to achieve a continuous transition between the spatial and frequency domains and dynamically separate high-frequency pathological details from low-frequency global trends.

[0059] The random frequency pattern selection unit is used to randomly select K frequency components (K is preferably 64) from the N frequency components of the input features, thereby reducing the self-attention complexity from... Down to To balance computational efficiency and feature integrity in order to reduce computational complexity;

[0060] The frequency domain attention unit is used to calculate the attention weights for the selected K frequency components and output the frequency domain enhancement features; the tanh activation function is used to preserve the positive and negative correlation of spectral features, selectively amplify key pathological features, and overcome the compression limitation of the SoftMax function on bipolar changes.

[0061] The process is as follows Figure 2 As shown:

[0062] Adaptive Fractional Fourier Transform (FrTrans module): A linear projection is performed on the input feature sequence X_seq to obtain the query Q, key K, and value V tensors. A one-dimensional FrFT is then applied along the feature sequence dimension N. The FrFT kernel function is defined as:

[0063]

[0064]

[0065] in For a one-dimensional spatial signal, a fractional-order parameter α∈[0.05,0.95] can be learned, corresponding to a rotation angle φ=α・π / 2, constrained by a regularized objective function:

[0066]

[0067] Where L CLS The sum of the label smoothing cross-entropy and the focus loss is λ=0.1;

[0068] Random frequency mode selection: K=64 frequencies are randomly selected from N frequency features to obtain Q. f K f , ;

[0069] Frequency domain attention: Q f K f V f Remodeling Attention weights are calculated by scaling the dot product:

[0070]

[0071] The result was reshaped back. The original spatial sequence is transformed back by inverse FrFT, connected with the input residual, and then fed into the feedforward network (FFN).

[0072] The L-GSA module includes a Local Spectral Attention (LSA) submodule and a Global Spectral Attention (GSA) submodule.

[0073] As attached Figure 3 As shown, the LSA submodule uses a shift-merge sliding window mechanism to capture local correlations in spectral channels. Specifically, the spectral channels are divided into multiple windows of size P=12. Local attention is calculated through linear transformation. The sliding window is shifted by P / 2 along the channel dimension to alleviate boundary discontinuities. Finally, the original and shifted window features are fused.

[0074] LSA submodule: Apply LayerNorm normalization to the input feature tensor.

[0075]

[0076] Divide the channel dimension D into M = D / P windows of size P = 12, and compute local attention for each window i:

[0077]

[0078] The sliding window is shifted P / 2 along the channel dimension, and the shift attention Attn_shift(j) is calculated (calculated in the same way as above). After fusion, local features are obtained.

[0079] As attached Figure 3 As shown, the submodule employs a dual-channel clustering mechanism to capture long-range dependencies of spectral channels. Through this mechanism, the spectral channels are clustered into Z=32 clusters, generating cluster weights and cluster centers. Long-range spectral dependencies are modeled using scaled dot product attention, reducing computational complexity from... Down to Specifically, this involves reshaping the input features into... After LayerNorm normalization, clustering weights are generated using the current projection. With cluster center The cluster center projection is After calculating the global attention, it is mapped back to the original channel space and connected to the input residual.

[0080] 3. Patch merging and cross-scale fusion module: connected to the feature extraction module, used for multi-scale fusion and spatial downsampling of deep features.

[0081] Patch merging: Use 2×2 convolution to downsample the spatial dimension and expand the channels, and then perform feature transformation with BatchNorm and a linear layer;

[0082] Cross-Scale Fusion (CSF): Adjusts feature channels through 1×1 convolution, aligns spatial scales using bilinear interpolation, and calculates weights based on feature entropy.

[0083]

[0084]

[0085]

[0086] Weighted fusion multi-stage features are obtained Emphasis is placed on information-rich components.

[0087] 4. Feature Fusion and Classification Module: Connected to the patch merging and cross-scale fusion module, it is used to generate pathological tissue classification results based on fused features;

[0088] Pathological Focused Attention: Attention weights are generated through linear layers and Sigmoid activation, and then compared with X. fused Element-level addition enhances local detail discrimination and suppresses background noise:

[0089]

[0090] Global pooling: for X weighted Adaptive average pooling is used to obtain the global feature vector. ;

[0091] The classification head consists of stacked linear layers and GELU activation, mapping the pooled global feature vector to the output class: X. pooled The data is fed into a stacked linear layer with GELU activation, and the predicted probabilities for each class are output: .

[0092] In a further implementation, the above model is obtained through the following training method:

[0093] The batch size was set (B=4 for the MDC dataset), the training epochs were 100, and a warm-up strategy was adopted for the learning rate (linearly increasing to the base rate for the first 20 epochs, then decreasing to 0.01 times after cosine annealing). The optimizer was Adam, and the loss function was a weighted combination of label smoothed cross-entropy loss and focus loss (weights of 0.7 and 0.3, respectively), to address the class imbalance problem. The preprocessed medical hyperspectral image training dataset was input into the model, and the network parameters, including the learnable fractional-order parameter α, were optimized through backpropagation. The total loss function was a weighted combination of label smoothed cross-entropy loss and focus loss, with weights of 0.7 and 0.3, respectively. The learnable fractional-order parameter α was iteratively updated through the backpropagation algorithm and fluctuated around 0.5 by Lα regularization constraints. The random frequency pattern selection mechanism selected K frequency components, with K ranging from 32 to 128. The activation function used for the frequency domain attention unit was the tanh function. Performance was monitored using a validation set, and early stopping was implemented. After training, save the optimized network parameters and weights file, and use the test dataset to verify and evaluate the classification performance.

[0094] Furthermore, during training, the model performance is monitored in real time using a validation dataset, and an early stopping strategy is employed to avoid overfitting. After training is complete, the optimized network parameters and weights are saved for subsequent deployment.

[0095] The training set is processed as follows:

[0096] S1. High-quality hyperspectral images of pathological tissues (including normal tissues, precancerous lesions, cancerous tissues, etc.) are acquired through a medical hyperspectral standard acquisition platform, with a spectral range covering 450-1000nm. At the same time, the corresponding tissue specimens are subjected to pathological "gold standard" diagnosis as the image annotation results to ensure the accuracy of the training dataset annotation.

[0097] S2. Use the pathological "gold standard" diagnostic results to classify and label the acquired medical hyperspectral images, and clarify the classification labels of different pathological tissues (such as normal bile duct tissue, bile duct cancer tissue, normal gastric tissue, intestinal metaplasia tissue, gastric intraepithelial neoplasia tissue, etc.).

[0098] S3. To verify the effectiveness of the algorithm and the generalization ability of the model, the standard dataset is scientifically divided, and a five-fold cross-validation or random seed random distribution strategy is adopted: 80% of the hyperspectral data is used as training data, the remaining 10% of the data is used as validation data, and 10% of the data is used as test data. The experiment is repeated many times to verify the generalization ability of the model.

[0099] S4. Preprocess the hyperspectral image data, including data augmentation operations such as random flipping, 90-degree rotation, brightness and contrast adjustment, and channel noise addition. Align the data format through standard procedures to form a standard training dataset for medical hyperspectral image classification.

[0100] After training, the frequency-domain enhanced biological in-situ hyperspectral feature extraction and classification model described in this application is obtained.

[0101] The second embodiment of this application discloses a frequency-domain enhanced in-situ hyperspectral feature extraction and classification method for biological applications. Employing the aforementioned classification model, and deploying this model on a standard medical hyperspectral acquisition platform, it enables real-time classification of medical hyperspectral images in clinical settings. The method includes the following steps:

[0102] The target medical hyperspectral image to be classified is input into the multi-scale patch embedding module to generate a multi-scale feature sequence.

[0103] The multi-scale feature sequence is input into the feature extraction module and processed alternately by the FrTrans module and the L-GSA module to obtain a deep feature representation. In the FrTrans module, an adaptive fractional Fourier transform is performed on the features based on a learnable fractional-order parameter α to achieve a continuous transition in the spatial frequency domain. A random frequency mode selection mechanism is used to randomly select K frequency components from N frequency components for attention calculation, reducing the computational complexity from O(N²) to O(KN). In the L-GSA module, a shift-and-merge sliding window mechanism is used to capture local spectral correlations, and a dual-channel clustering mechanism is used to capture long-range spectral dependencies.

[0104] The deep feature representation is input into the patch merging and cross-scale fusion module, and multi-scale feature fusion is performed by calculating weights based on feature entropy; and

[0105] The fused features are input into the feature fusion and classification module, and pathological tissue classification results are generated after the pathological focus attention enhancement is used to identify details.

[0106] To demonstrate the strong feature extraction capabilities of our model (FracTrans) from medical hyperspectral data, we evaluated FracTrans on a medical hyperspectral dataset classification task. First, we introduce the experimental setup, including the dataset, evaluation metrics, and implementation details. Then, we present a comprehensive comparative experiment.

[0107] Dataset: Experiments were conducted on the publicly available medical hyperspectral dataset (Multidimensional Bile Duct Dataset, MDC), which is suitable for classifying complex medical tissues. This dataset contains 2460 hyperspectral images with dimensions of 256×256×60 pixels, covering the spectral range of 550-1000 nm, including 1360 normal samples and 1100 abnormal samples (containing cancerous and precancerous tissues). The dataset was randomly split into training and test sets in a 4:1 ratio.

[0108] Evaluation metrics: The model performance is comprehensively evaluated using seven standard metrics: accuracy (ACC), specificity, precision, recall, Kappa coefficient, F1 score, and area under the ROC curve (AUC).

[0109] Comparative experiments: Comparison with eight state-of-the-art models, covering different architectures: ResNet50 based on CNN, FITNet and Swin Transformer based on visual Transformer, H2Former hybrid Transformer-CNN, FCP and SF-Unet frequency domain models, HSLabeling sparse labeling, and UFPF general feature perception model of microscopic MHSI.

[0110] Implementation details: Data augmentation is applied during training, including random flipping, 90-degree rotation, brightness and contrast adjustment, and channel noise addition. The MDC batch size is [value missing]. The learning rate employs a warm-up strategy: linearly increasing to the base rate for the first 20 rounds, then decreasing to 0.01 times through cosine annealing. The loss function combines label-smoothed cross-entropy and focus loss to handle class imbalance.

[0111] Classification Performance: Table 1 shows a comprehensive performance comparison across the two datasets. Methods integrating frequency domain information consistently outperform purely spatial methods on all evaluation metrics, demonstrating superior spectral-spatial modeling capabilities. The significant performance gap between FracTrans and purely spatial feature extraction methods like SwinTransformer highlights the fundamental importance of frequency domain analysis in capturing the inherent spectral characteristics of hyperspectral data. ResNet50, Swin Transformer, and FIT-Net exhibit low Kappa coefficients for accuracy, indicating limited classification consistency. ResNet50 performs moderately, suggesting that pure spatial convolution struggles to capture complex spectral patterns. Swin Transformer performs particularly poorly on MDC, possibly because its window attention mechanism is unsuitable for processing continuous spectral information. These observations suggest that integrating frequency domain information alone without an adaptive spectral modeling mechanism is insufficient to achieve optimal performance. FCP, SF-Unet, and HSLabeling show improvements in accuracy and recall, but have significant limitations. FCP is effective in handling frequency components but lacks adaptive frequency selection. SF-Unet exhibits large performance fluctuations across datasets and limited generalization. UFPF showed significant deficiencies in AUC and specificity for MDC, indicating that fixed-frequency processing may limit its adaptability to different spectral features. As shown in the table, FracTrans achieved the best or comparable levels across all evaluation metrics. In particular, the Kappa coefficient for binary classification was significantly improved, confirming the effectiveness of the adaptive spatial-frequency decomposition mechanism in capturing and differentiating complex pathological patterns.

[0112] Table 1 Comparison Results of MDC Datasets

[0113] .

[0114] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A biological in-situ hyperspectral feature extraction and classification model based on frequency domain enhancement, characterized in that, Along the data processing direction, it includes: Multi-scale patch embedding module: used to receive target medical hyperspectral images and generate multi-scale feature sequences; Feature extraction module: connected to the multi-scale patch embedding module, it consists of alternating fractional domain feature extraction module FrTrans and local-global spectral attention module L-GSA, used for deep feature extraction of the multi-scale feature sequence; Patch merging and cross-scale fusion module: connected to the feature extraction module, used for multi-scale fusion and spatial downsampling of deep features; Feature fusion and classification module: connected to the patch merging and cross-scale fusion module, used to generate pathological tissue classification results based on fused features; The FrTrans module includes an adaptive fractional Fourier transform unit configured with a learnable fractional-order parameter α, used to achieve a continuous transition between the spatial domain and the frequency domain through a rotation angle φ = α·π / 2. The random frequency pattern selection unit is used to randomly select K frequency components from N frequency components of the input feature to reduce computational complexity. A frequency domain attention unit is used to calculate attention weights for the selected K frequency components and output frequency domain enhancement features; The L-GSA module includes a Local Spectral Attention (LSA) submodule and a Global Spectral Attention (GSA) submodule. The LSA submodule uses a shift-and-merge sliding window mechanism to capture local correlations of spectral channels, while the GSA submodule uses a dual-channel clustering mechanism to capture long-range dependencies of spectral channels.

2. The model according to claim 1, characterized in that, The adaptive fractional Fourier transform unit further includes a regularization constraint unit, used to pass the objective function. The fractional-order parameter α is constrained to fluctuate around 0.5; the range of the learnable fractional-order parameter α is [0.05, 0.95].

3. The model according to claim 1, characterized in that, The LSA submodule includes: a window partitioning unit for dividing the spectral channel into windows of size P; a shift window unit for shifting the window along the channel dimension by P / 2 units; and a feature fusion unit for fusing the attention outputs of the original window and the shift window. The GSA submodule includes: a dual-channel clustering unit for clustering spectral channels into Z cluster centers to generate cluster weights; and a global attention calculation unit for calculating long-range dependencies based on the cluster centers.

4. The model according to claim 1, characterized in that, The feature fusion and classification module includes: Pathology Focused Attention Unit: Used to generate attention weights through a linear layer and a Sigmoid activation function, and then element-wise summed with the input features to enhance local discriminative details; Global pooling unit: used for adaptive average pooling of weighted features; Classification Header Unit: Used to map the pooled global feature vector to pathological tissue categories.

5. A frequency-domain enhanced method for in-situ hyperspectral feature extraction and classification of biological organisms, characterized in that, The classification model described in any one of claims 1-4 includes the following steps: The target medical hyperspectral image to be classified is input into the multi-scale patch embedding module to generate a multi-scale feature sequence. The multi-scale feature sequence is input into the feature extraction module and processed alternately by the FrTrans module and the L-GSA module to obtain a deep feature representation. In the FrTrans module, an adaptive fractional Fourier transform is performed on the features based on a learnable fractional-order parameter α to achieve a continuous transition in the spatial frequency domain. A random frequency mode selection mechanism is used to randomly select K frequency components from N frequency components for attention calculation, reducing computational complexity from... Down to In the L-GSA module, local spectral correlations are captured through a shift-merging sliding window mechanism, and long-range spectral dependence is captured through a dual-channel clustering mechanism. The deep feature representation is input into the patch merging and cross-scale fusion module, and multi-scale feature fusion is performed by calculating weights based on feature entropy; and The fused features are input into the feature fusion and classification module, and pathological tissue classification results are generated after the pathological focus attention enhancement is used to identify details.

6. The method according to claim 5, characterized in that, The classification model is obtained through the following training steps: inputting the medical hyperspectral image training dataset into the classification model, and optimizing the network parameters, including the learnable fractional-order parameter α, through the backpropagation algorithm; wherein, a weighted combination of label smoothing cross-entropy loss and focus loss is used as the total loss function.

7. The method according to claim 6, characterized in that, In the total loss function, the weights of the label smoothing cross-entropy loss and the focus loss are 0.7 and 0.3, respectively; the learnable fractional-order parameter α is iteratively updated through the backpropagation algorithm and through L... α The regularization constraint fluctuates around 0.

5.

8. The method according to claim 5, characterized in that, The random frequency pattern selection mechanism selects K frequency components, where K ranges from 32 to 128; the frequency domain attention unit uses the tanh function as its activation function.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the frequency domain enhanced biological in-situ hyperspectral feature extraction and classification method as described in any one of claims 5-8, or constructs a frequency domain enhanced biological in-situ hyperspectral feature extraction and classification model as described in any one of claims 1-4.

10. A frequency-domain enhanced in-situ hyperspectral feature extraction and classification device for biological organisms, characterized in that, include: Memory, used to store computer programs; The processor, when executing the computer program, controls the frequency-domain enhanced biological in-situ hyperspectral feature extraction and classification model as described in any one of claims 1-4 to perform classification operations.