Pneumonia image identification and classification method based on improved Swin Transform model

By improving the Swin Transformer model and combining data augmentation and teacher-student guidance technology, the existing deep learning models are solved inadequate generalization capabilities and high computational complexity in large-scale data processing, and more efficient and accurate pneumonia image recognition classification is achieved.

CN119992176APending Publication Date: 2025-05-13GUANGZHOU HUAYI ELECTRONIC TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510051657.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing deep learning models have problems such as insufficient generalization ability, high computational complexity and insufficient extraction of local features of images when processing large-scale data, resulting in low efficiency and accuracy in pneumonia image recognition classification.

Method used

The improved Swin Transformer model is adopted to preprocess and enhance the original data, and use ResNet-34 as the teacher model to guide and train the Swin Transformer student model, and use MSGToken and shuffle operations to improve training efficiency and feature extraction capabilities.

Benefits of technology

It improves the efficiency and accuracy of pneumonia detection, reduces calculation costs, and enhances the generalization ability and feature extraction ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992176A_ABST
    Figure CN119992176A_ABST
Patent Text Reader

Abstract

The invention discloses a pneumonia image identification and classification method based on an improved Swin Transform model. The method comprises the following steps: S100, carrying out preprocessing of three aspects of content standardization, image colorization and format unification on original data; s200, performing data enhancement on the input training set image; s300, selecting ResNet-34 as a teacher model, and carrying out pre-training on the processed training set by using the teacher model to obtain a hard tag for guiding a student model; s400, carrying out training by using an improved Swin Transform student model, and improving the training efficiency of the student model by using a method including but not limited to MSG Token and shuffle; when the loss function is calculated, a hard tag output by the teacher model is used for guiding the student model; and S500, importing a chest radiograph image to be classified, and obtaining a classification result by using the trained student model. According to the method, the problems of insufficient generalization ability, high calculation complexity, insufficient image local feature extraction and the like when an existing deep learning model is used for processing large-scale data can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and in particular to a pneumonia image recognition and classification method based on an improved Swin Transformer model. Background Art

[0002] With the development of medical imaging technology, the amount of medical image data has increased dramatically. How to efficiently and accurately identify and classify these data has become an important research direction in medical imaging.

[0003] Traditional convolutional neural networks (CNNs) can automatically learn multi-level abstract features from raw images through structures such as convolutional layers, pooling layers, and fully connected layers, and perform well in image classification tasks. However, the local feature extraction characteristics of such methods may lead to misjudgment of recognition and classification that requires comprehensive consideration of global information. In addition, existing Transformer-based models usually have high computational complexity and long training time, making them difficult to effectively deploy in real-world environments with limited resources. Summary of the invention

[0004] The technical problem to be solved by the present invention is to provide a pneumonia image recognition and classification method based on a modified Swin Transformer model to solve the problems of insufficient generalization ability, high computational complexity and insufficient extraction of local image features in existing deep learning models when processing large-scale data.

[0005] In order to solve the above technical problems, the technical solutions adopted by the present invention are as follows.

[0006] A pneumonia image recognition and classification method based on a modified Swin Transformer model comprises the following steps:

[0007] S100. Preprocessing the original data in three aspects: content standardization, image colorization and format unification, to obtain an RGB three-channel image in jpg format with the to-be-identified classification area of ​​the chest X-ray image located in the middle;

[0008] S200. Perform data enhancement on the input training set images;

[0009] S300. Select ResNet-34 as the teacher model, and use the teacher model to perform pre-training on the processed training set, use multiple convolutional layers and residual blocks to extract image features, and obtain hard labels for guiding the student model;

[0010] S400. Use the improved Swin Transformer student model for training, and use methods including but not limited to MSGToken and shuffle to improve the training efficiency of the student model; and use the hard labels output by the teacher model to guide the student model when calculating the loss function;

[0011] S500. Import the chest X-ray image to be classified, use the trained student model to downsample the data, linearly map the image into a three-dimensional tensor, then extract the entire image information through several local multi-head self-attention operations and shuffle operations, and finally obtain the classification result based on the extracted information.

[0012] Preferably, the specific method of data enhancement in S200 includes rotation, horizontal flipping and brightness adjustment; assuming that the input original data set is The data processed by steps S100 and S200 are as follows:

[0013]

[0014] X″ i =γR ω (X′ i )+β

[0015] Among them, μ and σ are the mean and standard deviation of pixels respectively; f C is a mapping based on color wheel or histogram matching; R θ is the rotation operation, ω∈[-15°,15°]; γ controls the brightness intensity; β is the bias term.

[0016] Preferably, the step S300 specifically includes the following steps:

[0017] S310. Input the processed training set data into the first 7×7 convolutional layer, and then perform maximum pooling after passing through the batch normalization layer and activation function;

[0018] S320. Pass the result of the previous layer through several residual blocks in sequence, each residual block includes two convolutional layers, a BN layer and an activation function between the convolutional layers, and a skip connection;

[0019] S330. Perform global average pooling, downsampling and full connection on the output result of the previous step in turn to obtain the final output result; then calculate the loss function and update the parameters based on the prediction result of the teacher model to proceed to the next round of iteration.

[0020] Preferably, the step S400 specifically includes the following steps:

[0021] S410. Downsample the input image;

[0022] S420. Divide the image into multiple windows, add a MSG-Token to each window, and perform multi-head self-attention operation on the local window;

[0023] S430. Perform a shuffle operation on the MSG Token of each window, divide it into small blocks equal to the number of windows, and then exchange the small blocks to exchange information between MSG Tokens to achieve cross-window transmission of important feature information;

[0024] S440. Calculate the loss function based on the student model output and the teacher model output, and update the parameters using the Adam optimizer.

[0025] Preferably, the specific method of step S410 is: take a pixel at every other position to form 4 feature maps; then splice them in the channel dimension to obtain a tensor with a depth 4 times the original; then pass through the LayerNorm layer and the Linear layer to map the depth direction of each pixel, and finally obtain a three-dimensional tensor with a depth of twice the original and a height and width reduced to 1 / 2 of the original.

[0026] Preferably, in step S420, the image is divided into multiple windows, and a specific method of adding a MSG-Token in each window is as follows: the student model image feature extraction is a W-MSA operation that integrates MSG-Token, and for the processed image, it is divided into 8×8 windows, and each window is spliced ​​with a learnable MSG-Token and an exchange operation is performed:

[0027] Shuffle(m i )=m π(i)

[0028] The final output is:

[0029] Z final =Concat(Z1,Z2,...,Z n )W O

[0030] Among them, m i is the MSG Token in each window; n is the number of windows; π is the random permutation operation; W O is the learnable weight matrix of the output layer.

[0031] Preferably, the specific method of performing the multi-head self-attention operation on the local window in step S420 is: linearly transform the input pixel data, and then divide it into h heads, and perform a multi-head self-attention operation on each head. i Multiply by the trainable weight matrix Get Q, K for matching and calculating attention weights, and V for providing information, and substitute them into the formula to get each head i Attention:

[0032]

[0033] Among them, d k is the dimension of K; B is the position code;

[0034] Concatenate the results from each head:

[0035] MultiHead(Q,K,V)=Concat(head1,...,head h )W O

[0036] Get the final output.

[0037] Preferably, the step S430 is specifically expressed as:

[0038] T' MSG = reshape(T MSG ),

[0039] T' MSG = transpose(T' MSG )

[0040] T MSG = reshape(T' MSG ),

[0041] Among them, n×n is the number of windows, that is, the number of MSGTokens; d is the number of original channels.

[0042] Preferably, the step S440 is specifically as follows: the output of the previous layer is sequentially passed through the LinearNorm layer, the average pooling operation and the fully connected layer to convert the input features into the probability distribution of the category, thereby generating the classification prediction of the student model; then, by comparing the difference between the predicted probability distribution and the actual label, the classification loss is calculated; at the same time, in order to introduce the guidance information of the teacher model, the loss value between the predicted probability distribution of the student model and the predicted result of the teacher model is calculated to obtain the knowledge distillation loss, which is specifically calculated using the cross entropy loss function:

[0043]

[0044] Where N is the number of samples, C is the number of categories; y ic With y' icare the unique hot encodings of the true label and the teacher's predicted hard label, respectively, ic The student model predicts the probability that the i-th sample belongs to the c category; a is the weight coefficient.

[0045] Preferably, in step S440, the model parameters are finally updated using the Adam optimizer, specifically: the gradient g of the loss function with respect to each parameter θ is calculated by back propagation t , and update the first-order momentum and second-order momentum:

[0046] m t =β1×m t-1 +(1-β1)×g t

[0047]

[0048] Among them, m t-1 is the historical first-order momentum; v t-1 is the historical second-order momentum; β1 and β2 are exponential decay rates; in the Adam optimizer, by default m0=v0=0, β1=0.9, β2=0.999; then update the parameters based on the first-order momentum and the second-order momentum:

[0049]

[0050] Among them, α is the learning rate; ε is a very small number; λ is the weight decay coefficient.

[0051] Due to the adoption of the above technical scheme, the technical progress achieved by the present invention is as follows.

[0052] The present invention can solve the problems of insufficient generalization ability, high computational complexity, and insufficient extraction of local image features in existing deep learning models when processing large-scale data, thereby improving the efficiency and accuracy of pneumonia detection in chest X-rays. The advantages include the following:

[0053] 1) MSG Token is introduced to simplify the traditional SW-MSA calculation while maintaining effective information transmission between local windows, thereby reducing the computational cost and enhancing the model performance.

[0054] 2) Data augmentation techniques are used to enrich the diversity of the dataset and improve the generalization ability of the model, so that the model can understand the lung structure at different angles and adapt to different exposure conditions.

[0055] 3) By selecting the pre-trained ResNet-34 as the teacher model and Swin Transformer as the student model, hard labels are used to guide the learning process of the student model, reducing the overfitting phenomenon caused by insufficient data and improving the model's feature extraction and generalization capabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 is a flow chart of the present invention;

[0057] Figure 2 is a flow chart of step S300 of the present invention;

[0058] Figure 3 This is a flow chart of step S400 of the present invention. DETAILED DESCRIPTION

[0059] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0060] A pneumonia image recognition and classification method based on the improved Swin Transformer model, combined with Figure 1 As shown, the following steps are included:

[0061] S100. The original data is preprocessed in three aspects: content standardization, image colorization and format unification, to obtain an RGB three-channel image in jpg format with the to-be-identified classification area of ​​the chest X-ray image located in the middle.

[0062] Specifically, the original data were preprocessed in three aspects: content standardization, image colorization, and format unification, so that the lung window area of ​​all chest X-ray images was located in the middle, unnecessary noise artifacts were shielded through denoising, and the grayscale image in png format was converted into RGB three-channel image in jpg format through format conversion and other methods.

[0063] S200. Perform data augmentation on the input training set images.

[0064] The input training set images are enhanced by methods such as rotation, flipping, and brightness adjustment, so that the model has richer and more accurate data content for learning.

[0065] Specifically, data enhancement includes three methods: rotation, horizontal flipping, and brightness adjustment. The present invention only selects the above three transformation methods to imitate the changes in the real world to enhance the robustness of the model, because the difference in chest X-ray images in the real world does not show drastic color jitter, etc. Excessive and random transformations may reduce the performance of the model.

[0066] Assume that the original input data set is The data processed by steps S100 and S200 are as follows:

[0067]

[0068] X″ i =γR ω(X′ i )+β

[0069] Among them, μ and σ are the mean and standard deviation of pixels respectively; f C is a mapping based on color wheel or histogram matching; R θ is the rotation operation, ω∈[-15°,15°]; γ controls the brightness intensity; β is the bias term.

[0070] S300. Select ResNet-34 as the teacher model, and use the teacher model to perform pre-training on the processed training set, use multiple convolutional layers and residual blocks to extract image features, and obtain hard labels to guide the student model.

[0071] In order to improve the generalization ability of the model, the present invention uses a hard distillation network based on teacher-student guidance, using the experience already gained by the teacher model to enable the student model to acquire a certain ability to resist noise, thereby improving the performance of the student model.

[0072] like Figure 2 As shown, taking one iteration of the training process as an example, this step specifically includes the following steps:

[0073] S310. Input the processed training set data into the first 7×7 convolutional layer, and then perform maximum pooling after passing through the batch normalization layer and activation function. The details are as follows:

[0074] The input image size of the present invention is n×n=224×224, the convolution step size is s=2, the padding size is p=3, and the first layer convolution kernel size is f=7. Then the size after passing through the convolution layer is:

[0075]

[0076] The calculated value is o=112. Then the result is batch normalized (Batch Normalization, BN):

[0077]

[0078] Among them, x is the input data; μ B and are the mean and standard deviation of the current batch respectively; ε is a very small number; γ and β are trainable parameters; y is the output. Then y is input into the ReLu activation function:

[0079] f(y)=max(0,y)

[0080] All negative values ​​are replaced with 0. Finally, the data is max-pooled.

[0081] S320. The result of the previous layer is passed through several residual blocks in sequence. Each residual block includes two convolutional layers, a BN layer between the convolutional layers, and an activation function, in addition to a skip connection.

[0082] The jump connection adds a short-circuit connection between the input of the previous layer and the activation function of the next layer, so that the output of the residual block adds an identity mapping:

[0083] y l =h(x l )+F(x l ,W l )

[0084] x l+1 =f(y l )

[0085] Among them, x l is the input of the residual block; h(x l ) is an identity mapping, that is, h(x l )=x l ; F(x l ,W l ) is the mapping that needs to be learned; f(y l ) is the activation function; x l+1 is the input of the next residual block.

[0086] S330. Perform global average pooling, downsampling and full connection on the output result of the previous step in turn to obtain the final output result; then calculate the loss function and update the parameters based on the prediction result of the teacher model to proceed to the next round of iteration.

[0087] Since in the present invention, the loss function of the teacher model is the same as the optimizer update method, the specific formula will be described in step S440.

[0088] S400. Use the improved Swin Transformer student model for training, and use methods including but not limited to MSGToken and shuffle to improve the training efficiency of the student model; and when calculating the loss function, use the hard labels output by the teacher model to guide the student model.

[0089] Specifically, the present invention adopts an improved model based on the improved Swin Transformer (SwinT) as the student model, and on the basis of the original model, the present invention adds two modules, MSG Token and Distillation Token, to improve the operating efficiency and generalization ability of the model. Among them, MSG Token refers to the MSG Transformer proposed by Jiemin Fang et al. in 2022. This module replaces the SW-MSA in the traditional SwinT model and improves the operating efficiency by adding an additional Token for information exchange; and Distillation Token comes from the teacher model trained in S300, through which the student model is guided.

[0090] like Figure 3 As shown, taking one iteration in the training process as an example, this step specifically includes the following steps:

[0091] S410. Downsample the input image. The specific method is:

[0092] Take a pixel at every other position to form 4 feature maps; then splice them in the channel dimension to get a tensor with 4 times the original depth; then pass through the LayerNorm layer and the Linear layer to map the depth direction of each pixel, and finally get a three-dimensional tensor with twice the original depth and reduced height and width to 1 / 2 of the original.

[0093] S420. Divide the image into multiple windows, add a MSG-Token to each window, and perform multi-head self-attention operation on the local windows.

[0094] Specifically, the image is divided into multiple windows, and a MSG-Token is added to each window. The specific method is as follows: the image feature extraction of the student model is a W-MSA operation that integrates MSG-Token. For the processed image, it is divided into 8×8 windows, and each window is spliced ​​with a learnable MSG-Token and exchanged:

[0095] Shuffle(m i )=m π(i)

[0096] The final output is:

[0097] Z final =Concat(Z1,Z2,...,Z n )W O

[0098] Among them, m iis the MSG Token in each window; n is the number of windows; π is the random permutation operation; W O is the learnable weight matrix of the output layer.

[0099] Specifically, the specific method of performing the multi-head self-attention operation on the local window is as follows: within the local window, perform the multi-head self-attention operation proposed by Vaswani et al. in 2017. Perform a linear transformation on the input pixel data, and then divide it into h heads. i Multiply by the trainable weight matrix Get Q, K for matching and calculating attention weights, and V for providing information, and substitute them into the formula to get each head i Attention:

[0100]

[0101] Among them, d k is the dimension of K; B is the position code.

[0102] Concatenate the results from each head:

[0103] MultiHead(Q,K,V)=Concat(head1,...,head h )W O

[0104] Get the final output.

[0105] S430. Perform a shuffle operation on the MSG Token of each window, split it into small blocks equal to the number of windows, and then exchange the small blocks to exchange information between MSG Tokens to achieve cross-window transmission of important feature information. Specifically expressed as:

[0106] T' MSG = reshape(T MSG ),

[0107] T' MSG = transpose(T' MSG )

[0108] T MSG = reshape(T' MSG ),

[0109] Among them, n×n is the number of windows, that is, the number of MSG Tokens; d is the number of original channels.

[0110] S440. Calculate the loss function based on the student model output and the teacher model output, and update the parameters using the Adam optimizer.

[0111] Specifically, the output of the previous layer is passed through the LinearNorm layer, the average pooling operation, and the fully connected layer in turn to convert the input features into the probability distribution of the category, thereby generating the classification prediction of the student model; then, by comparing the difference between the predicted probability distribution and the actual label, the classification loss is calculated; at the same time, in order to introduce the guidance information of the teacher model, the loss value between the predicted probability distribution of the student model and the predicted result of the teacher model is calculated to obtain the knowledge distillation loss, which is specifically calculated using the cross entropy loss function:

[0112]

[0113] Where N is the number of samples, C is the number of categories; y ic With y' ic are the unique hot encodings of the true label and the teacher's predicted hard label, respectively, ic The student model predicts the probability that the i-th sample belongs to the c category; a is the weight coefficient.

[0114] Finally, the Adam optimizer is used to update the model parameters. Specifically, the gradient g of the loss function with respect to each parameter θ is calculated by back propagation. t , and update the first-order momentum and second-order momentum:

[0115] m t =β1×m t-1 +(1-β1)×g t

[0116]

[0117] Among them, m t-1 is the historical first-order momentum; v t-1 is the historical second-order momentum; β1 and β2 are exponential decay rates; in the Adam optimizer, by default m0=v0=0, β1=0.9, β2=0.999; then update the parameters based on the first-order momentum and the second-order momentum:

[0118]

[0119] Among them, α is the learning rate; ε is a very small number; λ is the weight decay coefficient.

[0120] S500. Import the chest X-ray image to be classified, use the trained student model to downsample the data, linearly map the image into a three-dimensional tensor, then extract the entire image information through several local multi-head self-attention operations and shuffle operations, and finally obtain the classification result based on the extracted information.

[0121] The parameters required in the classification process are all trained by the S400 student model.

[0122] When used, the present invention can be applied to pneumonia detection and can improve the efficiency and accuracy of pneumonia detection in chest X-rays. Compared with previous methods, the present invention has the following advantages:

[0123] 1) MSG Token is introduced to simplify the traditional SW-MSA calculation while maintaining effective information transmission between local windows, thereby reducing the computational cost and enhancing the model performance.

[0124] 2) Data augmentation techniques are used to enrich the diversity of the dataset and improve the generalization ability of the model, so that the model can understand the lung structure at different angles and adapt to different exposure conditions.

[0125] 3) By selecting the pre-trained ResNet-34 as the teacher model and Swin Transformer as the student model, hard labels are used to guide the learning process of the student model, reducing the overfitting phenomenon caused by insufficient data and improving the model's feature extraction and generalization capabilities.

Claims

1. A pneumonia image recognition and classification method based on a modified Swin Transformer model, characterized by: The following steps are involved: S100. Preprocessing the original data in three aspects: content standardization, image colorization and format unification, to obtain an RGB three-channel image in jpg format with the to-be-identified classification area of ​​the chest X-ray image located in the middle; S200. Perform data enhancement on the input training set images; S300. Select ResNet-34 as the teacher model, and use the teacher model to perform pre-training on the processed training set, use multiple convolutional layers and residual blocks to extract image features, and obtain hard labels for guiding the student model; S400. Use the improved Swin Transformer student model for training, and use methods including but not limited to MSGToken and shuffle to improve the training efficiency of the student model; and use the hard labels output by the teacher model to guide the student model when calculating the loss function; S500. Import the chest X-ray image to be classified, use the trained student model to downsample the data, linearly map the image into a three-dimensional tensor, then extract the entire image information through several local multi-head self-attention operations and shuffle operations, and finally obtain the classification result based on the extracted information.

2. According to claim 1, a pneumonia image recognition and classification method based on a modified SwinTransformer model is characterized in that: The specific method of data enhancement in S200 includes rotation, horizontal flipping and brightness adjustment; assuming that the input original data set is The data processed by steps S100 and S200 are as follows: X i ”=γR ω (X i ')+β Among them, μ and σ are the mean and standard deviation of pixels respectively; f C is a mapping based on color wheel or histogram matching; R θ is the rotation operation, ω∈[-15°,15°]; γ controls the brightness intensity; β is the bias term.

3. The pneumonia image recognition and classification method based on the improved SwinTransformer model according to claim 1, characterized in that: The step S300 specifically includes the following steps: S310. Input the processed training set data into the first 7×7 convolutional layer, and then perform maximum pooling after passing through the batch normalization layer and activation function; S320. Pass the result of the previous layer through several residual blocks in sequence, each residual block includes two convolutional layers, a BN layer and an activation function between the convolutional layers, and a skip connection; S330. Perform global average pooling, downsampling and full connection on the output result of the previous step in turn to obtain the final output result; then calculate the loss function and update the parameters based on the prediction result of the teacher model to proceed to the next round of iteration.

4. The pneumonia image recognition and classification method based on the improved SwinTransformer model according to claim 1, characterized in that: The step S400 specifically includes the following steps: S410. Downsample the input image; S420. Divide the image into multiple windows, add a MSG-Token to each window, and perform multi-head self-attention operation on the local window; S430. Perform a shuffle operation on the MSG Token of each window, divide it into small blocks equal to the number of windows, and then exchange the small blocks to exchange information between MSG Tokens to achieve cross-window transmission of important feature information; S440. Calculate the loss function based on the student model output and the teacher model output, and update the parameters using the Adam optimizer.

5. The pneumonia image recognition and classification method based on the improved SwinTransformer model according to claim 4 is characterized in that: The specific method of step S410 is: take a pixel at every other position to form 4 feature maps; then splice them in the channel dimension to obtain a tensor with a depth four times the original; then pass through the LayerNorm layer and the Linear layer to map the depth direction of each pixel, and finally obtain a three-dimensional tensor with a depth of twice the original and a height and width reduced to 1 / 2 of the original.

6. The pneumonia image recognition and classification method based on the improved SwinTransformer model according to claim 4 is characterized in that: In step S420, the image is divided into multiple windows, and a specific method of adding a MSG-Token in each window is as follows: the student model image feature extraction is a W-MSA operation that integrates MSG-Token. For the processed image, it is divided into 8×8 windows, and each window is spliced ​​with a learnable MSG-Token and exchanged: Shuffle(m i )=m π(i) The final output is: WITH final =Concat(Z1,Z2,...,Z n )IN O Among them, m i is the MSG Token in each window; n is the number of windows; π is the random permutation operation; W O is the learnable weight matrix of the output layer.

7. The pneumonia image recognition and classification method based on the improved SwinTransformer model according to claim 6 is characterized in that: The specific method of performing the multi-head self-attention operation on the local window in step S420 is: linearly transform the input pixel data, and then divide it into h heads, and perform a multi-head self-attention operation on each head. i Multiply by the trainable weight matrix Get Q, K for matching and calculating attention weights, and V for providing information, and substitute them into the formula to get each head i Attention: Among them, d k is the dimension of K; B is the position code; Concatenate the results from each head: MultiHead(Q,K,V)=Concat(head1,...,head h )W O Get the final output.

8. The pneumonia image recognition and classification method based on the improved SwinTransformer model according to claim 4 is characterized in that: The step S430 is specifically expressed as follows: T' MSG =transpose(T' MSG ) Among them, n×n is the number of windows, that is, the number of MSGTokens; d is the number of original channels.

9. The pneumonia image recognition and classification method based on the improved Swin Transformer model according to claim 4, characterized in that: The step S440 is specifically as follows: the output of the previous layer is sequentially passed through the LinearNorm layer, the average pooling operation and the fully connected layer to convert the input features into the probability distribution of the category, thereby generating the classification prediction of the student model; then, by comparing the difference between the predicted probability distribution and the actual label, the classification loss is calculated; at the same time, in order to introduce the guidance information of the teacher model, the loss value between the predicted probability distribution of the student model and the predicted result of the teacher model is calculated to obtain the knowledge distillation loss, which is specifically calculated using the cross entropy loss function: Where N is the number of samples, C is the number of categories; y ic With y' ic are the unique hot encodings of the true label and the teacher's predicted hard label, respectively, ic The student model predicts the probability that the i-th sample belongs to the c category; a is the weight coefficient.

10. The pneumonia image recognition and classification method based on the improved SwinTransformer model according to claim 9, characterized in that: In step S440, the model parameters are finally updated using the Adam optimizer, specifically: the gradient g of the loss function with respect to each parameter θ is calculated by back propagation t , and update the first-order momentum and second-order momentum: m t =β1×m t-1 +(1-β1)×g t Among them, m t-1 is the historical first-order momentum; v t-1 is the historical second-order momentum; β1 and β2 are exponential decay rates; in the Adam optimizer, by default m0=v0=0, β1=0.9, β2=0.999; then update the parameters based on the first-order momentum and the second-order momentum: Among them, α is the learning rate; ε is a very small number; λ is the weight decay coefficient.