Terahertz medical image classification method based on fractional Fourier transform and VGG16
Through the combination of fractional Fourier transform and VGG16, the multi-scale feature extraction and noise robustness problems in terahertz medical image classification are solved, and the effective classification of small lesions and minor changes is achieved, which improves the accuracy and robustness of image classification.
Patent Information
- Application Number
- CN202510621070.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-07-25
AI Technical Summary
The existing medical image classification technology is difficult to effectively extract multi-scale features when processing complex medical images, and is poorly robust to noise and edge blur, especially in terahertz medical images, which is difficult to distinguish between small lesions or minor changes.
The terahertz medical image classification method based on fractional Fourier transform and VGG16 is adopted. By acquiring terahertz medical images for preprocessing, the fractional Fourier convolution module is used for feature extraction, combined with the full connection module for image classification, and the weight is updated through backpropagation, to achieve flexible time-frequency domain representation and multi-scale information extraction.
It improves the multi-scale information extraction capability of terahertz medical image classification and the analysis ability of complex images, can better distinguish small lesions or minor changes, suppress noise at certain frequencies, and enhances the edge and detail feature extraction of images.
Smart Images

Figure CN120375091A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image analysis, and particularly relates to a terahertz medical image classification method based on fractional Fourier transform and VGG16. Background Art
[0002] Existing medical images mainly rely on CT and X-ray. Their image classification technology mainly depends on traditional image processing methods and deep learning models. When dealing with complex medical images, traditional image processing methods often fail to effectively extract multi-scale features. Existing deep learning models focus on local areas and rarely consider complex global information and frequency domain features, and have poor robustness to noise and edge blurring.
[0003] As a low-energy electromagnetic wave, terahertz has significant safety advantages (no ionizing radiation), has high resolution in the imaging of biological tissues, and has broad application prospects.
[0004] To address the above problems, the present invention proposes a terahertz medical image classification method based on fractional Fourier transform VGG16 (FrFT-VGG16). This method has a flexible time-frequency domain representation, can suppress noise at certain frequencies, has a stronger multi-scale information extraction ability, and the network's parsing ability for complex images enables the network to better distinguish small lesions or subtle changes in terahertz medical images. Summary of the Invention
[0005] The purpose of the present invention is to provide a terahertz medical image classification method based on fractional Fourier transform and VGG16, which has a flexible time-frequency domain representation, can suppress noise at certain frequencies, has a stronger multi-scale information extraction ability, and the network's parsing ability for complex images enables the network to better distinguish small lesions or subtle changes in terahertz medical images.
[0006] To solve the above technical problems, the technical solution adopted by the present invention is as follows:
[0007] A terahertz medical image classification method based on fractional Fourier transform and VGG16 includes the following steps:
[0008] S1: Obtain a terahertz medical image and preprocess the terahertz medical image;
[0009] S2: Apply a fractional Fourier convolution module to extract features from the preprocessed terahertz medical image to obtain network output features;
[0010] S3: Classify the terahertz medical image through a fully connected module and add classification labels;
[0011] S4: Match the classification tags with the network output features, and continuously update the weights through backpropagation to obtain the final image classification result.
[0012] Preferably, the specific process of preprocessing the terahertz medical image in step S1 is as follows:
[0013] Divide the data set into a training set, a validation set, and a test set according to a ratio of 7:2:1. Adjust terahertz medical images of different sizes to 224*224 pixels, and read their RGB images.
[0014] Preferably, the fractional Fourier convolutional module in step S2 sequentially includes a first convolutional layer, a first fractional Fourier convolutional layer, a first max-pooling layer, a second convolutional layer, a second fractional Fourier transform layer, a second max-pooling layer, a third convolutional layer, a third fractional Fourier transform layer, a third max-pooling layer, a fourth convolutional layer, a fourth fractional Fourier transform layer-1, a fourth fractional Fourier transform layer-2, a fourth max-pooling layer, a fifth convolutional layer, a fifth fractional Fourier transform layer-1, a fifth fractional Fourier transform layer-2, a fifth max-pooling layer, a first fully connected layer, a second fully connected layer, and an output layer.
[0015] Preferably, the first convolutional layer includes 64 convolution kernels with a kernel size of 3x3, a stride of 1, and a padding of 1. Perform a convolution operation on the input image. The output of the first convolutional layer is:
[0016]
[0017] where O i 1 (x, y) is the value of the output feature map of the i-th convolution kernel at position (x, y), and K i (j, k) is the value of the weight of the i-th convolution kernel at position (j, k);
[0018] After the convolution operation, use a preset activation function to introduce non-linear features. The mathematical expression is as follows:
[0019] f(x) = max(0, x);
[0020] where f(x) is the output of the activation function. When the input x is negative, the output is 0; when the input x is positive, the output is x;
[0021] After the convolution operation, obtain the output feature map O i 1 (x, y) of each convolution kernel, and apply the preset activation function to each feature map:
[0022] R i 1 (x, y) = f(Oi 1 (x, y)) = max(0, O i 1 (x, y))
[0023] Wherein, R i 1 (x, y) is the feature map processed by a preset activation function, that is, the output of the first convolutional layer after passing through the activation function.
[0024] Input shape: The preprocessed medical image, with a size of (224, 224, 3).
[0025] Output shape: 64 feature maps, with a size of (224, 224, 64).
[0026] Preferably, the first max pooling layer performs downsampling through a max pooling layer with a size of 2x2 and a stride of 2 to reduce the spatial size of the feature map and reduce the computational amount.
[0027] Input feature map R i 1 (x, y), after a 2x2 max pooling operation, the output feature map P i 1 (x, y) can be expressed by the following calculation formula:
[0028]
[0029] Wherein, P i 1 (x, y) is the output feature map after pooling, R i 1 (x, y) is the feature map processed by the activation function, and j and k represent the relative positions within the pooling window, taking values of 0 or 1 respectively.
[0030] Input shape: (224, 224, 64);
[0031] Output shape: (112, 112, 64), and the pooling operation halves the size of the feature map.
[0032] Preferably, the first fractional Fourier convolutional layer:
[0033] The input shape is (224, 224, 64)
[0034] The output shape is that the size of the image remains unchanged after the fractional Fourier convolutional layer transformation, still being (224, 224, 64).
[0035] Preferably, the specific process of applying the fractional Fourier convolutional module in step S2 to extract features from the preprocessed terahertz medical image to obtain the network output features is as follows:
[0036] S21: Channel splitting, dividing the input features into four equal parts along the channel dimension;
[0037] S22: Perform two-dimensional fractional Fourier transform;
[0038] S23: Convolution processing, separating the outputs after the four fractional Fourier transforms and performing convolutions of 3×3, 1×1, 1×1, and 1×1 respectively;
[0039] S24: Perform two-dimensional inverse fractional Fourier transform to convert the fractional-domain features back to the spatial domain;
[0040] S25: Channel merging, concatenating the four features along the channel dimension to keep the input-output shape of the FRFT layer unchanged.
[0041] Preferably, the specific expression for performing the two-dimensional fractional Fourier transform in step S22 is as follows:
[0042]
[0043] where the kernel:
[0044] K p1,p2 (h, k; u, v) = K p1 (h, u)K p2 (k, v);
[0045] K p1 (h, u) is defined as follows:
[0046]
[0047] where p1, p2 are the orders of the two dimensions, h, k are the spatial-domain coordinates of the input image, u, v are the coordinates in the transformed two-dimensional domain, δ in δ(h - u) is the delta function, and φ h = p1π / 2 is its rotation angle.
[0048] In addition, K p1 (h, u) and K p2 (k, v) have the same form. When both are set to the same value p1 = p2 = p, where p is the order of the 2D-DFrFT.
[0049] Preferably, the specific expression for performing the two-dimensional inverse fractional Fourier transform in step S24 to convert the fractional-domain features back to the spatial domain is as follows:
[0050]
[0051] Among them, p1 and p2 are the orders of two dimensions, h and k are the spatial domain coordinates of the input image, and u and v are the coordinates in the transformed two-dimensional domain. is the kernel function of the IFRFT, corresponding to the kernel function K p1,p2 (h, k; u, v) of the forward transform. The inverse kernel function can be expressed as:
[0052]
[0053] Among them, and are the kernel functions of the IFRFT in each dimension respectively.
[0054] The beneficial effects of the present invention include:
[0055] The terahertz medical image classification method based on the fractional Fourier transform and VGG16 provided by the present invention obtains a terahertz medical image and performs preprocessing, applies a fractional Fourier convolution module to extract features from the preprocessed terahertz medical image to obtain network output features, classifies the terahertz medical image through a fully connected module and adds classification labels, matches the classification labels with the network output features, and continuously updates the weights through backpropagation to obtain the final image classification result. It can have a flexible time-frequency domain representation in image classification, suppress noise at certain frequencies, improve the multi-scale information extraction ability and the network's parsing ability for complex images, enabling the network to better distinguish small lesions or minute changes in terahertz medical images. Description of the Drawings
[0056] Figure 1 is a schematic flow chart of the terahertz medical image classification method based on the fractional Fourier transform and VGG16 of the present invention.
[0057] Figure 2 is a schematic architecture diagram of the fractional Fourier transform convolutional neural network (FrFT-VGG16) model of the present invention.
[0058] Figure 3 is a schematic architecture diagram of the fractional Fourier convolution module of the present invention.
[0059] Figure 4 is the prediction result of the lung CT FRFT-VGG16 of the present invention. Detailed Embodiments
[0060] The following further describes the present invention in detail with reference to the attached Figures 1 to 4 drawings:
[0061] Example 1
[0062] See the attached Figure 1As shown, the terahertz medical image classification method based on fractional Fourier transform and VGG16 includes the following steps:
[0063] S1: Acquire a terahertz medical image and preprocess the terahertz medical image;
[0064] S2: Apply the fractional Fourier convolution module to extract the features of the preprocessed terahertz medical image to obtain the network output features;
[0065] S3: Classify terahertz medical images and add classification labels through a fully connected module;
[0066] S4: Use the classification label to match the network output feature, and continuously update the weight through back propagation to obtain the final image classification result.
[0067] In this embodiment, the specific process of preprocessing the terahertz medical image in step S1 is as follows:
[0068] The dataset was divided into training set, validation set, and test set in a ratio of 7:2:1. Terahertz medical images of different sizes were adjusted to 224*224 pixels, and their RGB images were read.
[0069] The main implementation steps of the fractional Fourier transform convolutional neural network (FrFT-VGG16) are: preprocessing the input image and label, converting the image into a matrix form that can be calculated, and adjusting the size; applying the designed fractional Fourier convolution (FrFT Layer) module and convolution, pooling and other modules to extract features of the data; using a 3-layer fully connected module to classify the image, matching the classification label with the network output features, and back propagating to continuously update the weights to obtain the final accurate terahertz medical image classification results. The above image classification process can have a flexible time-frequency domain representation in image classification, suppress noise on certain frequencies, improve the ability to extract multi-scale information and the network's ability to analyze complex images, so that the network can better distinguish small lesions or minor changes in terahertz medical images.
[0070] Example 2
[0071] On the basis of Embodiment 1, the present invention constructs a fractional Fourier transform convolutional module (FrFT Layer). Based on VGG16, some convolutional layers are replaced with fractional Fourier transform convolutional (FrFT Layer) modules to process images from multiple angles in the spatio-frequency plane, enhancing the model's multi-scale information extraction ability and the network's parsing ability for complex images. The feature extraction part has a total of five modules, each module contains a convolutional layer, a fractional Fourier transform convolutional module, a max pooling layer, etc. In the figure, the blue blocks are conventional convolutional blocks plus ReLU activation functions, the yellow blocks are fractional Fourier transform convolutional modules plus ReLU activation functions, and the red blocks are max pooling layers. The image classification part is mainly composed of three fully connected layers. In the figure, the green blocks are fully connected layers plus ReLU activation functions.
[0072] See Figure 2 As shown, the fractional Fourier transform convolutional module sequentially includes a first convolutional layer, a first fractional Fourier transform convolutional layer, a first max pooling layer, a second convolutional layer, a second fractional Fourier transform layer, a second max pooling layer, a third convolutional layer, a third fractional Fourier transform layer, a third max pooling layer, a fourth convolutional layer, a fourth fractional Fourier transform layer - 1, a fourth fractional Fourier transform layer - 2, a fourth max pooling layer, a fifth convolutional layer, a fifth fractional Fourier transform layer - 1, a fifth fractional Fourier transform layer - 2, a fifth max pooling layer, a first fully connected layer, a second fully connected layer, and an output layer.
[0073] The first convolutional layer includes 64 convolutional kernels with a kernel size of 3x3, a stride of 1, and a padding of 1, and performs a convolution operation on the input image. The output of the first convolutional layer is:
[0074]
[0075] Among them, O i 1 (x, y) is the value of the output feature map of the i-th convolutional kernel at the position (x, y), and K i (j, k) is the value of the weight of the i-th convolutional kernel at the position (j, k);
[0076] After the convolution operation, a preset activation function is used to introduce non-linear features. The mathematical expression is as follows:
[0077] f(x) = max(0, x);
[0078] Among them, f(x) is the output of the activation function. When the input x is negative, the output is 0; when the input x is positive, the output is x;
[0079] After the convolution operation, the output feature map O of each convolutional kernel is obtained i 1(x, y), apply a preset activation function to each feature map:
[0080] R i 1 (x, y) = f(O i 1 (x, y)) = max(0, O i 1 (x, y))
[0081] where R i 1 (x, y) is the feature map processed by the preset activation function, that is, the output of the first convolutional layer after passing through the activation function.
[0082] Input shape: The preprocessed medical image, with a size of (224, 224, 3).
[0083] Output shape: 64 feature maps, with a size of (224, 224, 64).
[0084] Fractional Fourier Transform Convolutional Layer (FrFT Layer 1):
[0085] Input shape: (224, 224, 64);
[0086] Output shape: The size of the image remains unchanged after the FRFT layer transformation, still (224, 224, 64).
[0087] The first max pooling layer performs downsampling through a max pooling layer with a size of 2x2 and a stride of 2 to reduce the spatial size of the feature map and the computational amount.
[0088] Input feature map R i 1 (x, y), after a 2x2 max pooling operation, the output feature map P i 1 (x, y) can be calculated by the formula:
[0089]
[0090] where P i 1 (x, y) is the output feature map after pooling, R i 1 (x, y) is the feature map processed by the activation function, and j and k represent the relative positions within the pooling window, taking values of 0 or 1 respectively.
[0091] Input shape: (224, 224, 64);
[0092] Output shape: (112, 112, 64). The pooling operation halves the size of the feature map.
[0093] Second Convolutional Layer (Conv2D Layer 2)
[0094] Apply 128 convolutional kernels of 3x3, stride 1, and padding 1 to perform a second convolution operation on the output of the previous layer and activate it with ReLU to further extract features. Its formula can be expressed as:
[0095] R i 2 (x, y) = f(O i 2 (x, y)) = max(0, O i 2 (x, y));
[0096] Similarly, R i 2 (x, y) is the feature map processed by the ReLU activation function, that is, the output of the second convolutional layer after passing through the activation function.
[0097] Input shape: (112, 112, 64).
[0098] Output shape: 128 feature maps, with a size of (112, 112, 128).
[0099] Fractional Fourier Transform Layer (FrFT Layer 2)
[0100] Input shape: (112, 112, 128)
[0101] Output shape: The size of the image remains unchanged after the FRFT layer transformation, still (112, 112, 128).
[0102] Second Max Pooling Layer (Max Pooling Layer 2)
[0103] Use 2x2 max pooling again to further downsample the feature map.
[0104]
[0105] Similarly, P i 2 (x, y) is the output feature map after pooling. j and k represent the relative positions within the pooling window, taking values 0 or 1 respectively.
[0106] Input shape: (112, 112, 128).
[0107] Output shape: (56, 56, 128).
[0108] Third Convolutional Layer (Conv2D Layer 3)
[0109] Apply 256 convolutional kernels of size 3x3, stride 1, and padding 1 to perform the third convolution operation on the output of the previous layer and activate it with ReLU to further extract features. Its formula can be expressed as:
[0110] R i 3 (x, y) = f(O i 3 (x, y)) = max(0, O i 3 (x, y));
[0111] Similarly, R i 3 (x, y) is the feature map processed by the ReLU activation function, that is, the output of the third convolutional layer after passing through the activation function.
[0112] Input shape: (56, 56, 128).
[0113] Output shape: 128 feature maps, size (56, 56, 256).
[0114] Third Fractional Fourier Transform Layer (FrFT Layer 3)
[0115] Input shape: (56, 56, 256).
[0116] Output shape: The size of the image remains unchanged after the FRFT layer transformation, still (56, 56, 256).
[0117] Third Max Pooling Layer (Max Pooling Layer 3)
[0118] Use 2x2 max pooling to finally downsample the feature map.
[0119]
[0120] Similarly, P i 3 (x, y) is the output feature map after pooling, and j and k represent the relative positions within the pooling window, taking values 0 or 1 respectively.
[0121] Input shape: (56, 56, 256).
[0122] Output shape: (28, 28, 256).
[0123] Fourth Convolutional Layer (Conv2D Layer 4)
[0124] Apply 512 convolution kernels of 3x3, stride 1, and padding 1 to perform the third convolution operation on the output of the upper layer and activate it with ReLU to further extract features. Its formula can be expressed as:
[0125] R i 4 R(x, y) = f(O i 4 (x, y)) = max(0, O i 4 (x, y));
[0126] R i 4 R(x, y) is the feature map processed by the ReLU activation function, that is, the output of the fourth convolution layer after passing through the activation function.
[0127] Input shape: (28, 28, 256).
[0128] Output shape: 512 feature maps, with a size of (28, 28, 512).
[0129] Fourth Fractional Fourier Transform Layer (FrFT Layer 4-1)
[0130] Input shape: (28, 28, 512).
[0131] Output shape: The size of the image remains unchanged after the FRFT layer transformation, still (28, 28, 512).
[0132] Fractional Fourier Transform Layer (FrFT Layer 4-2)
[0133] Input shape: (28, 28, 512).
[0134] Output shape: The size of the image remains unchanged after the FRFT layer transformation, still (28, 28, 512).
[0135] Fourth Max Pooling Layer (Max Pooling Layer 4)
[0136] Use 2x2 max pooling to finally downsample the feature map.
[0137]
[0138] P i 4 P(x, y) is the output feature map after pooling. j and k represent the relative positions within the pooling window, taking values of 0 or 1 respectively.
[0139] Input shape: (28, 28, 512).
[0140] Output shape: (14, 14, 512).
[0141] Fifth Convolutional Layer (Conv2D Layer 5)
[0142] Apply 512 convolutional kernels of 3x3, stride 1, and padding 1 to perform the third convolution operation on the output of the upper layer and activate it with ReLU to further extract features. Its formula can be expressed as:
[0143] R i 5 R(x, y) = f(O i 5 (x, y)) = max(0, O i 5 (x, y));
[0144] R i 5 R(x, y) is the feature map processed by the ReLU activation function, that is, the output of the fifth convolutional layer after the activation function.
[0145] Input shape: (14, 14, 512).
[0146] Output shape: 512 feature maps, with a size of (14, 14, 512).
[0147] Fifth Fractional Fourier Transform Layer (FrFT Layer 5-1)
[0148] Input shape: (14, 14, 512).
[0149] Output shape: The size of the image remains unchanged after the FRFT layer transformation, still (14, 14, 512).
[0150] Sixth Fractional Fourier Transform Layer (FrFT Layer 5-2)
[0151] Input shape: (14, 14, 512).
[0152] Output shape: The size of the image remains unchanged after the FRFT layer transformation, still (14, 14, 512).
[0153] Fifth Max Pooling Layer (Max Pooling Layer 5)
[0154] Use 2x2 max pooling to finally downsample the feature map.
[0155]
[0156] P i 5(x, y) is the output feature map after pooling. j and k represent the relative positions within the pooling window, taking values of 0 or 1 respectively.
[0157] Input shape: (14, 14, 512).
[0158] Output shape: (7, 7, 512).
[0159] First fully connected layer (Dense Layer 1):
[0160] Flatten the three-dimensional feature map into a one-dimensional vector for use in the fully connected layer.
[0161] Input shape: (7, 7, 512), 7 x 7 x 512 = 25088, flattened to (25088).
[0162] Output shape: 4096.
[0163] Second fully connected layer (Dense Layer 2):
[0164] Input the image into a fully connected layer with 4096 neurons and use the ReLU activation function.
[0165] Input shape: (4096).
[0166] Output shape: (4096).
[0167] Output layer
[0168] A fully connected layer with 2 output neurons (assuming 2 classes), using the softmax activation function for multi-class classification.
[0169] Input shape: (4096).
[0170] Output shape: (2) classes.
[0171] Example 3
[0172] Based on Example 1 or Example 2, refer to Figure 3 As shown, the specific process of applying the fractional Fourier convolution module in step S2 to extract features from the preprocessed terahertz medical image to obtain the network output features is as follows:
[0173] S21: Channel splitting, evenly divide the input features into four parts along the channel dimension;
[0174] S22: For the normalized image f′(h, k) of size M×N, perform a two-dimensional fractional Fourier transform. Compared with the convolutional two-dimensional discrete Fourier transform (DFT), the 2D-FrFT has higher adaptability and flexibility. As the rotation angle changes, the time-frequency domain characteristics of the transformed image also change.
[0175] S23: Convolution processing, separate the outputs after four fractional Fourier transforms, and perform convolutions of 3×3, 1×1, 1×1, and 1×1 respectively;
[0176] S24: Perform an inverse two-dimensional fractional Fourier transform to convert the fractional domain features back to the spatial domain;
[0177] S25: Channel merging, splice the four features along the channel dimension to keep the input and output shapes of the FRFT layer unchanged.
[0178] The specific expression for performing the two-dimensional fractional Fourier transform in step S22 is as follows:
[0179]
[0180] Among them, the kernel:
[0181] K p1,p2 (h, k; u, v) = K p1 (h, u)K p2 (k, v);
[0182] K p1 The definition of K
[0183]
[0184] where p1 and p2 are the orders of the two dimensions, h and k are the spatial domain coordinates of the input image, u and v are the coordinates in the transformed two-dimensional domain, δ in δ(h - u) is the delta function, and φ h = p1π / 2 is its rotation angle.
[0185] In addition, K p1 (h, u) and K p2 (k, v) have the same form. When both are set to the same value p1 = p2 = p, where p is the order of the 2D-DFrFT.
[0186] Based on the above equation, obviously the period of the transform kernel p is 4. Therefore, any real value within the range [0, 4) can be selected for p. Specifically, when p1 = p2 = π / 2, the FrFT is equivalent to the traditional Fourier transform. Due to the periodicity and symmetry of the fractional transform itself, we only need to study the values of the transform order within the range [0, 1].
[0187] The FrFT layer of the present invention is composed of four different order paths: spatial perspective, spatial-frequency perspective, frequency-spatial perspective, and frequency perspective. That is, the input features are transformed into four fractional domains of different orders through 2D-FrFT, and the fractional orders are 0, 0.25, 0.75, and 1 respectively.
[0188] In step S24, two-dimensional inverse fractional Fourier transform is performed to transform the fractional domain features back to the spatial domain, and the specific expression is as follows:
[0189]
[0190] where p1 and p2 are the orders of the two dimensions, h and k are the spatial domain coordinates of the input image, and u and v are the coordinates in the transformed two-dimensional domain. is the kernel function of IFRFT, corresponding to the kernel function K p1,p2 (h, k; u, v) of the forward transform, and the inverse kernel function can be expressed as:
[0191]
[0192] where and are the kernel functions of IFRFT in each dimension respectively.
[0193] In order to test the classification ability of the terahertz medical classification model FRFT-VGG16 for medical images, the present invention uses the images in the publicly available dataset SARS-COV-2 Ct-Scan Dataset to train, validate, and test the model. This dataset contains thousands of lung CT scan images of COVID-19 positive and negative patients. On it, the training loss of the FRFT-VGG16 model is 0.1538, and the precision on the test set is 94.88%, with good results. Some of its output images are shown in the appendix Figure 4 as shown. Flexible time-frequency domain representation: The fractional Fourier VGG16 can flexibly convert between the time domain and the frequency domain, retain the information in the time domain and the frequency domain to different degrees, and can enhance the edge and detail features of the image. Anti-noise ability: By transitioning between the time-frequency domains, the noise at certain frequencies is suppressed, and it has strong robustness. Rich multi-scale information extraction ability: The fractional Fourier VGG16 can extract information of different scales through different order parameters, and flexibly adapt to the multi-scale features of terahertz medical images.
[0194] In summary, the terahertz medical image classification method based on fractional Fourier transform and VGG16 provided by the present invention obtains terahertz medical images and performs preprocessing, applies a fractional Fourier convolution module to extract features from the preprocessed terahertz medical images to obtain network output features, classifies the terahertz medical images through a fully connected module and adds classification labels, matches the classification labels with the network output features, and continuously updates the weights through backpropagation to obtain the final image classification result. It can have a flexible time-frequency domain representation in image classification, suppress noise at certain frequencies, improve the ability to extract multi-scale information and the network's ability to analyze complex images, enabling the network to better distinguish small lesions or subtle changes in terahertz medical images. The fractional Fourier transform convolutional neural network VGG16 can flexibly convert between the time domain and the frequency domain, retain the information in the time domain and the frequency domain to varying degrees, and can enhance the edge and detail features of images. By transitioning between the time-frequency domain, it suppresses noise at certain frequencies and has strong robustness. The fractional Fourier VGG16 can extract information at different scales through different order parameters and flexibly adapt to the multi-scale features of terahertz medical images.
Claims
1. A terahertz medical image classification method based on fractional Fourier transform and VGG16, characterized in that It includes the following steps: S1: Obtain a terahertz medical image and preprocess the terahertz medical image; S2: Apply a fractional Fourier convolution module to extract features from the preprocessed terahertz medical image to obtain network output features; S3: Classify the terahertz medical image through a fully connected module and add classification labels; S4: Use the classification labels to match with the network output features, and continuously update the weights through backpropagation to obtain the final image classification result.
2. The terahertz medical image classification method based on fractional Fourier transform and VGG16 according to claim 1, wherein The specific process of preprocessing the terahertz medical image in step S1 is as follows: Divide the data set into a training set, a validation set, and a test set according to the ratio of 7:2:
1. Adjust terahertz medical images of different sizes to 224*224 pixels and read their RGB images.
3. The terahertz medical image classification method based on fractional Fourier transform and VGG16 according to claim 1, characterized in that The fractional Fourier convolution module in step S2 sequentially includes a first convolutional layer, a first fractional Fourier convolutional layer, a first max pooling layer, a second convolutional layer, a second fractional Fourier transform layer, a second max pooling layer, a third convolutional layer, a third fractional Fourier transform layer, a third max pooling layer, a fourth convolutional layer, a fourth fractional Fourier transform layer - 1, a fourth fractional Fourier transform layer - 2, a fourth max pooling layer, a fifth convolutional layer, a fifth fractional Fourier transform layer - 1, a fifth fractional Fourier transform layer - 2, a fifth max pooling layer, a first fully connected layer, a second fully connected layer, and an output layer.
4. The terahertz medical image classification method based on fractional Fourier transform and VGG16 according to claim 3, wherein The first convolutional layer includes 64 convolutional kernels with a kernel size of 3x3, a stride of 1, and a padding of 1, and performs a convolution operation on the input image. The output of the first convolutional layer is: Among them, is the value of the output feature map of the i-th convolutional kernel at the position (x, y), and K i (j, k) is the value of the weight of the i-th convolutional kernel at the position (j, k); After the convolution operation, a preset activation function is used to introduce non - linear features. The mathematical expression is as follows: f(x) = max(0, x); where f(x) is the output of the activation function. When the input x is negative, the output is 0; when the input x is positive, the output is x; After the convolution operation, the output feature map of each convolution kernel is obtained. Apply the preset activation function to each feature map: Among them, is the feature map processed by the preset activation function, that is, the output of the first convolutional layer after passing through the activation function. Input shape: The preprocessed medical image, with a size of (224, 224, 3). Output shape: 64 feature maps, with a size of (224, 224, 64).
5. The terahertz medical image classification method based on fractional Fourier transform and VGG16 according to claim 3, characterized in that The first max pooling layer performs downsampling through a max pooling layer with a size of 2x2 and a stride of 2 to reduce the spatial dimension of the feature map and reduce the amount of calculation. Input feature map After a 2x2 max pooling operation, the output feature map P i 1 (x,y)'s calculation formula can be expressed as: Among them, is the output feature map after pooling, is the feature map after being processed by the activation function. j and k represent the relative positions within the pooling window, taking values of 0 or 1 respectively. Input shape: (224, 224, 64); Output shape: (112, 112, 64), and the pooling operation halves the size of the feature map.
6. The terahertz medical image classification method based on fractional Fourier transform and VGG16 according to claim 3, characterized in that, The first fractional Fourier convolutional layer: The input shape is (224, 224, 64) The output shape is that the size of the image after the fractional Fourier convolutional layer transformation remains unchanged, still (224, 224, 64).
7. The terahertz medical image classification method based on fractional Fourier transform and VGG16 according to claim 1, wherein The specific process of applying the fractional Fourier convolution module in step S2 to extract features from the preprocessed terahertz medical image to obtain network output features is as follows: S21: Channel splitting, evenly divide the input features into four parts along the channel dimension; S22: Perform a two - dimensional fractional Fourier transform; S23: Convolution processing, separate the outputs after the four fractional Fourier transforms and perform convolutions of 3×3, 1×1, 1×1, and 1×1 respectively; S24: Perform two-dimensional inverse fractional Fourier transform to convert the fractional-domain features back to the spatial domain; S25: Channel merging, concatenate the four features along the channel dimension to keep the input and output shapes of the FRFT layer unchanged.
8. The terahertz medical image classification method based on fractional Fourier transform and VGG16 according to claim 7, characterized in that The specific expression for performing two-dimensional fractional Fourier transform in step S22 is as follows: where kernel: K p1,p2 (h, k; u, v) = K p1 (h, u)K p2 (k, v); K p1 (h, u) is defined as follows: where p1 and p2 are the orders of the two dimensions, h and k are the spatial-domain coordinates of the input image, u and v are the coordinates in the transformed two-dimensional domain, and δ in δ(h - u) is the delta function. φ h = p1π / 2 is its rotation angle. In addition, K p1 (h, u) and K p2 (k, v) have the same form. When both are set to the same value p1 = p2 = p, where p is the order of the 2D-DFrFT.
9. The terahertz medical image classification method based on fractional Fourier transform and VGG16 according to claim 7, wherein The specific expression for performing two-dimensional inverse fractional Fourier transform in step S24 to convert the fractional-domain features back to the spatial domain is as follows: where p1 and p2 are the orders of two dimensions, h and k are the spatial domain coordinates of the input image, and u and v are the coordinates in the transformed two-dimensional domain. is the kernel function of the IFRFT, which corresponds to the kernel function of the forward transform The inverse kernel function can be expressed as: Among them, and are the kernel functions of the IFRFT in each dimension respectively.