Cervical cell classification method based on attention mechanism and Swin Transform

By adopting the CFA-Former network model based on attention mechanism and Swin Transformer in cervical cell classification, combining channel attention and spatial attention, and using RMSLayerNorm, the problems of insufficient characteristics and limited adaptability in traditional methods are solved, and higher classification accuracy and robustness are achieved.

CN119942214AActive Publication Date: 2025-05-06CHONGQING NORMAL UNIVERSITY +2
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510082217.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-06
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

Traditional cervical cell classification methods rely on the characteristics of manual design, are not robust enough, and have limited adaptability to images collected by different devices, making it difficult to effectively deal with the challenges of imaging conditions and staining differences.

Method used

The cervical cell classification method based on attention mechanism and Swin Transformer is adopted, and the CFA-Former network model is used to combine channel attention with spatial attention, enhance feature learning ability, and introduce RMSLayerNorm to replace the traditional LayerNorm layer to improve training efficiency.

Benefits of technology

It significantly improves the accuracy and robustness in the cervical cell classification task, reduces computational complexity, and shows better performance on multiple data sets, verifying its effectiveness in cervical cell classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942214A_ABST
    Figure CN119942214A_ABST
Patent Text Reader

Abstract

The invention discloses a cervical cell classification method based on an attention mechanism and a Swin Transform, and relates to the technical field of medical image processing. The invention provides a CFA-Former network model based on a CFA module, channel attention and space attention are combined, the limitation of a traditional model in capturing multi-scale features is solved, and the model can effectively improve the accuracy and robustness in a cervical cell classification task by intensifying the attention on important information and positions; in the model design, the CFA module adaptively focuses and inhibits important features through two learning paths and comprises a CDA sub-module and an SFA sub-module, and the CDA module can optimize information extraction of the network in the channel dimension through a lightweight channel attention mechanism; and the SFA module improves the feature expression ability of the model in the spatial dimension through an enhanced spatial attention mechanism, and especially shows significant advantages when coping with complex cell images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and in particular to a cervical cell classification method based on an attention mechanism and SwinTransformer. Background Art

[0002] Cervical cancer is the fourth most deadly cancer among women, but its incidence can be significantly reduced through effective cervical cancer screening. However, cervical cancer screening relies on accurate cell classification, but different imaging conditions and staining methods may lead to diverse cell morphology and sizes. Cells of the same category vary greatly, and cells of different categories may be very similar in appearance, which increases the difficulty of classification. Manual film reading requires a high degree of professional knowledge and rich clinical experience, and the process is cumbersome and time-consuming. Therefore, the study of automatic cervical cell screening has become a hot topic in academia.

[0003] In early studies, traditional machine learning algorithms were widely used for cervical cell classification, such as support vector machines (SVM) and AdaBoost. Wang et al. optimized the shape, texture, and Gabor features through feature selection algorithms and used SVM classification; Arya et al. used a variety of texture features including first-order histograms, gray-level co-occurrence matrices, local binary patterns, etc. for classification; Win et al. used random forests for feature selection and integrated multiple classifiers through bagging to generate the final result. However, these traditional methods are highly dependent on manually designed features, and manual feature design requires professional knowledge and is easily affected by subjectivity, resulting in insufficient robustness of the features. This makes the traditional methods have low classification accuracy and blurred boundaries when distinguishing cervical lesion cells, and the generalization ability is relatively limited when facing images collected by different devices.

[0004] With the continuous development of computer hardware and the increasing power of computing resources, deep learning has been widely used in the field of image processing. Deep learning technology can automatically extract complex high-level features from raw images and significantly improve the classification effect. For example, DeepCervix achieves excellent performance by integrating mixed deep features of multiple models; Deeppap directly classifies normal and abnormal single cervical cells based on image blocks with cell nuclei as the core. Lin et al. proposed a CNN model that combines cell image appearance and cell morphology for cervical cell classification; Zhao et al. designed the CCG-Taming Transformer structure, which generates images that are very similar to cervical cell images by introducing a novel convolutional data enhancement architecture and Token-to-Token visual transformer, thereby effectively improving the classification accuracy. Zhang et al. developed a new cervical cell dataset and designed a binary tree network with dual-path fusion attention features (BTTFA) to segment cell nuclei and distinguish different types of lesions; Hao et al. constructed an automatic cervical cell classification model using a feature fusion method based on VGG-16. Dong et al. combined Inception V3 with artificially designed features, fully integrating domain knowledge and achieved an accuracy of over 98% on the Herlev dataset; Basak et al. used Gray WolfOptimizer to optimize CNN dimensionality reduction features, thereby further improving classification performance.

[0005] In the traditional machine learning methods proposed above, cervical cell classification faces many problems. First, these methods are highly dependent on manually designed features. The extraction and design of features require rich professional knowledge and are easily affected by subjectivity, resulting in poor robustness of features. Secondly, the traditional methods have fuzzy boundaries and low accuracy during classification. In addition, due to the diversity of imaging devices and staining methods, these methods have poor adaptability to images collected by different devices, limited generalization ability, and it is difficult to effectively cope with the challenges of imaging conditions and staining differences.

[0006] Although deep learning methods have shown high classification accuracy, they also have some limitations. First, the training of deep learning models relies on a large amount of high-quality annotated data, while the datasets used in many current studies are mainly private datasets. The process of obtaining high-quality cervical cell datasets is expensive and time-consuming. Second, the traditional machine learning algorithms or CNN models used by Lin, Dong, Basak, and Zhang et al., although CNNs are powerful in image classification, may be biased in classification due to the complexity and diversity of cervical cell images. Although Zhao et al. used a Transformer structure that can capture global information, the Transformer's lack of attention to local details may lead to limited ability to accurately identify cancer cells, which is unacceptable in practical tasks. In addition, the Transformer structure is often accompanied by a large amount of computational complexity, and the traditional self-attention mechanism still has certain limitations in extracting local feature information and weighted processing of channel and spatial information, especially when processing images with rich details and local differences such as cervical cell images. Therefore, a more lightweight Transformer structure network that can reduce computational complexity and improve the classification performance of the model with only a small increase in parameters is needed, as well as a Transformer structure network that can better pay attention to channel and spatial information.

[0007] Therefore, a new solution to the above problems needs to be proposed. Summary of the invention

[0008] The purpose of the present invention is to provide a cervical cell classification method based on attention mechanism and Swin Transformer to solve the technical problems raised in the background technology.

[0009] To achieve the above object, the present invention provides the following technical solution: a cervical cell classification method based on attention mechanism and SwinTransformer, comprising at least the following steps:

[0010] S1: Build a cervical cell classification network framework based on the fusion of attention mechanism and Swin Transformer. The overall network structure is designed based on Swin Transformer and maintains the original structural characteristics of Swin Transformer, that is, image features are gradually extracted through four stages. In each stage, the number of CFA-Former blocks is consistent with the number of Swin Transformer blocks to ensure the effectiveness and consistency of the feature extraction process;

[0011] S2: Input an image into the cervical cell classification network framework. The input image is first divided into blocks by the PatchPartition module, and each 4×4 adjacent pixels is divided into a patch.

[0012] S3: The image is flattened in the channel direction. For cervical cell images, which are generally RGB three-channel images, the size of the image will change after flattening.

[0013] S4: In Stage 1, linear embedding is used for processing. Linear embedding is used to linearly transform the channel data of each pixel, transforming the image from the original dimension to the new dimension.

[0014] S5: In Stage 2 to Stage 4, the patch merging layer is used for downsampling to reduce the image size while retaining key information to further improve feature expression capabilities;

[0015] S6: When the image passes through the CFA-Former module, the RMSLayerNorm layer is normalized, and global average pooling is performed before finally reaching the classification head for classification. The RMSLayerNorm layer is used to replace the traditional LayerNorm layer to speed up the convergence of the network and reduce the consumption of computing resources.

[0016] Furthermore, the core components of the CFA-Former module include at least a CFA module and a W-MSA / SW-MSA module.

[0017] Furthermore, the CFA module is a parallel attention mechanism, which effectively combines channel attention with spatial attention, that is, an effective combination of CDA and SFA, thereby enhancing the feature learning ability of the model. This parallel structural design allows the model to simultaneously evaluate the criticality of the channel dimension and the spatial dimension, thereby avoiding the risk that a single attention mechanism may ignore certain important dimensions. Through weighted fusion, a more accurate feature representation can ultimately be obtained.

[0018] Furthermore, the CFA module generates a corresponding three-dimensional attention map for the input three-dimensional image to highlight important features;

[0019] The process of generating a 3D attention map is decomposed into two lightweight branches, each of which uses simplified components, effectively reducing the amount of computation and parameter consumption;

[0020] Each element of the feature map can be regarded as a feature detector, so the channel attention and spatial attention branches can clearly learn "which features to pay attention to" and "where to focus on", further optimizing the feature extraction process;

[0021] The CFA module information transmission process is as follows:

[0022] CFAM=σ(CDA+SFA)

[0023] CFA=(x*CFAM)+x

[0024] Among them, CFAM represents the attention feature map output after the fusion of channel attention and spatial attention, σ represents the Sigmoid function, CFA represents the output of the CFA module, and x represents the feature vector input to the CFA module.

[0025] Furthermore, the CDA adopts a lightweight and efficient channel attention mechanism to model the importance of each channel through adaptive pooling and adaptive convolution operations, without relying on complex fully connected layers, significantly reducing the computational complexity and the number of parameters that need to be learned, ensuring that the network can still maintain efficient computing performance when processing high-dimensional feature spaces;

[0026] The CDA has a convolution operation that can adaptively select the convolution kernel size according to the channel dimension of the input feature map;

[0027] Among them, k represents the amount of information transmission between the current channel and other k different channels. This value is dynamically determined by a nonlinear mapping related to the number of channels, so that the convolution operation can be flexibly adjusted to meet the information interaction requirements between different channels;

[0028] Since the input channels of cervical cell images are generally 3, they become 96 after the first stage of Liner Embedding. The number of channels doubles in each stage thereafter, so the number of channels is an even number. As the subject of the mapping;

[0029] Therefore, the mapping rule is given by:

[0030]

[0031] Where log(C) represents the natural logarithm of the number of channels C, b represents the offset in the general linear mapping, which is set to 1 in this paper. It means rounding down to the nearest odd number, and the convolution kernel size is generally set to an odd number;

[0032] The overall process within the CDA is as follows:

[0033]

[0034] Among them, AAP stands for adaptive average pooling, represents an adaptive k×k size convolution, σ represents the Sigmoid function, and expand() means expanding the output to the same shape as the input.

[0035] Furthermore, the SFA adopts the Bottleneck structure in ResNet, that is, SFA adopts a dilated convolution to expand the receptive field in a low-cost and high-efficiency manner, which can save both the number of parameters and the computational overhead.

[0036] Furthermore, the application of the SFA comprises at least the following steps:

[0037] First, a 1×1 convolution is used to reduce the dimension of the feature map C×H×W projection to The feature maps are integrated and compressed across the channel dimension, and the compression ratio reduction is set to 16 in this paper;

[0038] After dimensionality reduction, two dilated convolutions are applied to effectively capture and utilize contextual information;

[0039] Finally, the feature map is compressed to 1×H×W again using 1×1 convolution;

[0040] For the final output scale adjustment, expand() is used at the end to expand the feature map to the same size as the CDA branch so that the two branches can be fused;

[0041] The calculation process of the SFA is:

[0042]

[0043] Among them, f 1xl represents a convolution operation with a convolution kernel size of 1×1, represents a 3×3 atrous convolution, d is the dilation parameter of the atrous convolution, which is set to 4, and the padding is also set to 4.

[0044] Compared with the prior art, the present invention has the following beneficial effects:

[0045] 1. This paper proposes a CFA-Former network model based on the CFA (Cross Focus Attention) module. By combining channel attention and spatial attention, the limitations of traditional models in capturing multi-scale features are solved. In addition, the model can effectively improve the accuracy and robustness of cervical cell classification tasks by strengthening the focus on important information and locations.

[0046] 2. In the model design of the present invention, the CFA module adaptively focuses on and suppresses important features through two learning paths, including two sub-modules: CDA (Channel DimFocus Attention) and SFA (Spatial Focus Attention). The CDA module can optimize the information extraction of the network in the channel dimension through a lightweight channel attention mechanism; while the SFA module improves the feature expression ability of the model in the spatial dimension through an enhanced spatial attention mechanism, especially when dealing with complex cell images, showing significant advantages;

[0047] 3. The present invention also introduces RWSLayerNorm to replace the traditional LayerNorm layer, further improving the training efficiency of the model. This optimization not only speeds up the training process, but also effectively avoids common gradient vanishing and convergence problems, making the network more stable when processing large-scale data.

[0048] 4. The present invention has been comprehensively experimentally verified on multiple different data sets, including quantitative evaluation of labeled data sets and qualitative analysis of model results. The experimental results show that the present invention is significantly superior to existing methods in multiple evaluation indicators, verifying its effectiveness in cervical cell classification. Through reasonable architecture design and module innovation, the present invention successfully realizes the efficient extraction and fusion of global and local features of cervical cell images, overcoming the limitations of traditional methods and single deep learning models. In practical applications, the present invention can not only provide more accurate classification support for cervical cancer screening, but also provide valuable reference for medical image classification tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for describing the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0050] Figure 1 It is the overall network framework diagram of the present invention;

[0051] Figure 2 This is a CFA structure diagram of the present invention;

[0052] Figure 3 This is a comparison chart of the results of transplanting the module in the present invention to other models;

[0053] Figure 4 FLOPs comparison chart when the module in the present invention is transplanted to other models;

[0054] Figure 5 This is a comparison chart of the classification results of the present invention and other networks for specific categories. DETAILED DESCRIPTION

[0055] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0056] The present invention proposes a CFA (Cross Focus Attention) module and designs a CFA-Former network model based on this module. CFA, as the core component of CFA-Former, aims to enhance the representation ability of the network. Through two different learning paths, this module can efficiently identify and focus on the content and positions that need to be emphasized or suppressed, thereby effectively refining intermediate features. The CFA module includes two sub-modules: CDA (Channel DimFocus Attention) and SFA (Spatial Focus Attention). Among them, CDA is a lightweight channel attention module, while SFA is an enhanced spatial attention module. In addition, the model also uses RWSLayerNorm to replace the traditional LayerNorm layer, thereby significantly improving the training efficiency of the network.

[0057] A cervical cell classification method based on attention mechanism and Swin Transformer includes at least the following steps:

[0058] S1: Build a cervical cell classification network framework based on attention mechanism fusion and Swin Transformer (see Figure 1 ), the overall network structure is designed based on Swin Transformer and maintains the original structural characteristics of Swin Transformer, that is, image features are gradually extracted through four stages. In each stage, the number of CFA-Former blocks is consistent with the number of Swin Transformer blocks to ensure the effectiveness and consistency of the feature extraction process;

[0059] S2: Input an image into the cervical cell classification network framework. The input image is first divided into blocks by the PatchPartition module, and each 4×4 adjacent pixels is divided into a patch.

[0060] S3: The image is flattened in the channel direction. For cervical cell images, which are generally RGB three-channel images, the size of the image will change after flattening.

[0061] S4: In Stage 1, linear embedding is used for processing. Linear embedding is used to linearly transform the channel data of each pixel, transforming the image from the original dimension to the new dimension.

[0062] S5: In Stage 2 to Stage 4, the patch merging layer is used for downsampling to reduce the image size while retaining key information to further improve feature expression capabilities;

[0063] S6: When the image passes through the CFA-Former module, the RMSLayerNorm layer is normalized, and global average pooling is performed before finally reaching the classification head for classification. The RMSLayerNorm layer is used to replace the traditional LayerNorm layer to speed up the convergence of the network and reduce the consumption of computing resources.

[0064] Furthermore, the core components of the CFA-Former module include at least a CFA module and a W-MSA / SW-MSA module.

[0065] Further, the CFA module (see Figure 2 ) is a parallel attention mechanism. The CFA module effectively combines channel attention with spatial attention, that is, the effective combination of CDA and SFA, thereby enhancing the feature learning ability of the model. This parallel structural design allows the model to simultaneously evaluate the criticality of the channel dimension and the spatial dimension, thereby avoiding the risk that a single attention mechanism may ignore certain important dimensions. Through weighted fusion, a more accurate feature representation can be obtained in the end.

[0066] Furthermore, the CFA module generates a corresponding three-dimensional attention map for the input three-dimensional image to highlight important features;

[0067] The process of generating a 3D attention map is decomposed into two lightweight branches, each of which uses simplified components, effectively reducing the amount of computation and parameter consumption;

[0068] Each element of the feature map can be regarded as a feature detector, so the channel attention and spatial attention branches can clearly learn "which features to pay attention to" and "where to focus on", further optimizing the feature extraction process;

[0069] The CFA module information transmission process is as follows:

[0070] CFAM=σ(CDA+SFA)

[0071] CFA=(x*CFAM)+x

[0072] Among them, CFAM represents the attention feature map output after the fusion of channel attention and spatial attention, σ represents the Sigmoid function, CFA represents the output of the CFA module, and x represents the feature vector input to the CFA module.

[0073] Furthermore, the CDA adopts a lightweight and efficient channel attention mechanism to model the importance of each channel through adaptive pooling and adaptive convolution operations, without relying on complex fully connected layers, significantly reducing the computational complexity and the number of parameters that need to be learned, ensuring that the network can still maintain efficient computing performance when processing high-dimensional feature spaces;

[0074] The CDA has a convolution operation that can adaptively select the convolution kernel size according to the channel dimension of the input feature map;

[0075] Among them, k represents the amount of information transmission between the current channel and other k different channels. This value is dynamically determined by a nonlinear mapping related to the number of channels, so that the convolution operation can be flexibly adjusted to meet the information interaction requirements between different channels;

[0076] Since the input channels of cervical cell images are generally 3, they become 96 after the first stage of Liner Embedding. The number of channels doubles in each stage thereafter, so the number of channels is an even number. As the subject of the mapping;

[0077] Therefore, the mapping rule is given by:

[0078]

[0079] Among them, log(C) represents the natural logarithm of the number of channels C, b represents the offset in the general linear mapping, which is set to 1 in this paper, and the convolution kernel size is generally set to an odd number, so It means round up to the nearest odd number.

[0080] The overall process within the CDA is as follows:

[0081]

[0082] Among them, AAP stands for adaptive average pooling, represents an adaptive k×k size convolution, σ represents the Sigmoid function, and expand() means expanding the output to the same shape as the input.

[0083] Furthermore, the SFA adopts the Bottleneck structure in ResNet, that is, SFA adopts dilated convolution to expand the receptive field in a low-cost and high-efficiency manner, which can save both the number of parameters and the computational overhead;

[0084] In the SFA, a spatial attention map is finally generated to emphasize the features of different spatial positions or suppress unimportant spatial position features, that is, context information is crucial to identifying which spatial positions need attention. In order to efficiently utilize context information, a larger receptive field is required. However, using too large a convolution kernel will result in excessive computational overhead. Therefore, SFA uses dilated convolution to expand the receptive field in a low-cost, high-efficiency manner. Compared with standard convolution, dilated convolution can more easily construct an effective spatial mapping, because through the arrangement of dilated convolutions, the receptive field can be expanded exponentially, so that the SFA module can efficiently summarize and integrate context information.

[0085] Furthermore, the application of the SFA comprises at least the following steps:

[0086] First, a 1×1 convolution is used to reduce the dimension of the feature map C×H×W projection to The feature maps are integrated and compressed across the channel dimension, and the compression ratio reduction is set to 16 in this paper;

[0087] After dimensionality reduction, two dilated convolutions are applied to effectively capture and utilize contextual information;

[0088] Finally, the feature map is compressed to 1×H×W again using 1×1 convolution;

[0089] For the final output scale adjustment, expand() is used at the end to expand the feature map to the same size as the CDA branch so that the two branches can be fused;

[0090] The calculation process of the SFA is:

[0091]

[0092] Among them, f 1xl represents a convolution operation with a convolution kernel size of 1×1, represents a 3×3 atrous convolution, d is the dilation parameter of the atrous convolution, which is set to 4, and the padding is also set to 4.

[0093] The following technical assessment is proposed:

[0094] Quantitative evaluation

[0095] We evaluate the performance of the present invention through quantitative measurements. In order to effectively evaluate the proposed method, we use several evaluation indicators commonly used in the field of image classification as evaluation results, namely accuracy, precision, recall, F1-score, AUC value, sensitivity and specificity, where Sensitivity and Specificity are only used as evaluation indicators for binary classification. Accuracy is used to evaluate the probability of correct prediction in all samples, Precision is used to evaluate the proportion of samples predicted as positive samples that are actually positive, which can reflect the accuracy of model prediction, Recall evaluates the proportion of samples predicted as positive samples that are actually positive, which can reflect the comprehensiveness of model prediction, and F1 score is used to solve the balance between precision and recall, which is the harmonic mean of the two. The AUC value is the area under the ROC (Receiver Operating Characteristic) curve, which indicates the model's ability to distinguish between positive and negative samples. The closer the value is to 1, the stronger the model's classification ability is. The calculation formula for Sensitivity is the same as that for Recall, both of which reflect the comprehensiveness of the model's prediction. Specificity is used to measure the model's ability to correctly identify negative samples, that is, the proportion of samples that are actually negative that are correctly identified as negative by the model.

[0096] In addition to the basic classification indicators, the present invention also uses FLOPs, Throughout and Params to more comprehensively evaluate the performance of the model. Among them, FLOPs represents the number of floating-point operations required by the model during execution, which is used to measure the complexity of an algorithm or model calculation. The size is related to the complexity of the model and the set Batch Size, but the size is generally an integer multiple of Batch Size 1. Params represents the number of parameters in the model and is also used to evaluate the complexity of the algorithm or model. Throughput (images / s) represents the number of images processed by the model per second, which is used to measure the efficiency of the model. Under the same hardware conditions, the inference speed of the model can be understood.

[0097] Table 1 and Table 2 show the comparison of the indicators of the present invention with other methods on the LBC dataset and the Tianchi Cervical Cell Challenge dataset, respectively, where the LBC dataset is a four-category dataset and the Tianchi Cervical Cell Challenge dataset is a two-category dataset. The test results show that the present invention exhibits better performance and is superior to other methods in most indicators.

[0098] Since CFA-Former is designed based on Swin Transformer, Swin Transformer is selected as the baseline model of this experiment. According to the different structural designs of Swin Transformer, different structures of CFA-Former are designed accordingly. Table 1 shows the corresponding CFA-Former structural designs. Due to hardware limitations, the Tiny structure is used for experiments.

[0099] Table 1 Different structures of CFA-Former model

[0100]

[0101] Among them, Table 2 shows the results of quantitative measurements on the LBC dataset. The network of the Transformer architecture, which is more mainstream and advanced in recent years, is selected. Whether in terms of accuracy or F1 score, CFA-Former has reached the best comparison baseline model. The number of parameters has only increased by 0.16M, and FLOPs has hardly changed, but there is a significant improvement in performance, highlighting the superiority of CFA-Former.

[0102] Table 2 Quantitative measurement results on the Mendeley LBC dataset

[0103]

[0104] Table 3 records the comparison with the baseline model, proving that the model has a certain degree of generalization and the results on different data sets are better than the baseline model.

[0105] Table 3 Quantitative measurement results on the Tianchi Cervical Cell Challenge dataset

[0106]

[0107] Table 4 shows the effectiveness of the CFA module and RMSLayerNorm. First, experiments were conducted with or without the CDA module and the SFA module. The experimental results show that when the CDA module and the SFA module are removed, the accuracy decreases slightly, and when CDA and SFA are added at the same time, the accuracy increases significantly. Secondly, experiments were conducted on whether to replace the RMSLayerNorm layer. It can be clearly seen that the inference speed of the model has increased. Finally, CDA and SFA are added to the baseline model at the same time, that is, the CFA module is added, but the RMSLayerNorm layer is not replaced. It can be clearly seen that the accuracy is not much different, but the inference speed is significantly slower than CFA-Former.

[0108] Table 4 Quantitative measurement results on the Mendeley LBC dataset

[0109]

[0110]

[0111] 2) Qualitative evaluation

[0112] According to the characteristics of CFA module and RMSLayerNorm, considering whether it is portable, this paper merges the designed CFA module into some common Transformer structures, such as DeiT, T2T, PerViT, TNT, PVT, and also achieves good results. Figure 3 The following are the results of training these models from scratch on the LBC dataset and combining them with the CFA module proposed in this paper. Figure 4 These figures show that the CFA module is portable and reusable, and can reduce the complexity of the model and effectively improve the performance of the network.

[0113] Figure 4 and Figure 5 The accuracy comparison of CFA-Former, T2T, Focal, DeiT, and PerViT for each category (HSIL, LSIL, NILM, SCC) in the dataset is shown. As can be seen from the figure, the accuracy of each network for normal negative cells (NILM) is very high, showing excellent classification ability. However, in the two categories of HSIL and SCC, which are more difficult to classify, the accuracy of each network has decreased, but CFA-Former's performance is significantly better than other networks, with an accuracy of 99.38% in the HSIL class and 99.91% in the SCC class, both of which are the best performance. For the LSIL category, the accuracy of CFA-Former, PerViT, and Focal networks is not much different, showing relatively close performance, reaching 99.97% (CFA-Former), 99.87% (PerViT), and 99.88% (Focal) respectively.

[0114] In summary:

[0115] (1) The present invention innovatively designs a CFA module based on the fusion of attention mechanism, and constructs a CFA-Former network architecture based on the SwinTransformer structure, which is specifically used for the efficient classification of cervical cells. This architecture fully utilizes the advantages of Swin Transformer, effectively integrates it with the attention mechanism, and realizes the efficient combination of channel attention and spatial attention within the model. The CFA module contains two sub-modules: CDA (Channel DimFocusAttention) and SFA (Spatial Focus Attention). Among them, CDA is a lightweight channel attention module that focuses on improving classification performance by enhancing the features of important channels; while SFA is an enhanced spatial attention module that can focus on key spatial areas in the image. The synergy of the two sub-modules greatly enhances the feature extraction capability of the model, thereby significantly improving the overall classification performance.

[0116] (2) This invention successfully integrates the RMSLayerNorm (Root Mean Square Layer Normalization) technology into the Swin Transformer framework for the first time, aiming to improve the model performance in the cervical cell classification task. As a new normalization method, RMSLayerNorm can more effectively adapt to different input data distributions compared to the traditional LayerNorm, especially in deep learning, it has stronger generalization ability. Experimental results show that RMSLayerNorm can significantly improve the inference speed of the model while reducing overfitting in the application of Swin Transformer. In the specific task of cervical cell classification, the introduction of RMSLayerNorm not only optimizes the training process, but also further improves the accuracy and efficiency of the model in processing cervical cell images, providing a feasible and superior solution for this task.

[0117] (3) The present invention conducted a large number of experiments on two different cervical cell image datasets to verify the effectiveness and advantages of the CFA-Former architecture and its CFA module. The experimental results show that the CFA module can not only significantly improve the classification performance of cervical cells, but also has strong portability and scalability, and can be used in combination with other mainstream deep learning network frameworks (such as ResNet, VGG, etc.). By embedding the CFA module into the existing network architecture, the network model can still achieve a large performance improvement with a small parameter increment, showing significant flexibility. This feature enables the CFA module to be widely used in various classification tasks, especially in fields such as medical image analysis that require high-precision classification.

[0118] It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other specific forms without departing from the spirit or essential features of the invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations falling within the meaning and scope of the equivalent elements of the claims be included in the invention. Any reference numeral in a claim should not be considered as limiting the claim to which it relates.

Claims

1. A cervical cell classification method based on attention mechanism and Swin Transformer, characterized by: At least the following steps are included: S1: Build a cervical cell classification network framework based on the fusion of attention mechanism and Swin Transformer. The overall network structure is designed based on Swin Transformer and maintains the original structural characteristics of Swin Transformer, that is, image features are gradually extracted through four stages. In each stage, the number of CFA-Former blocks is consistent with the number of Swin Transformer blocks to ensure the effectiveness and consistency of the feature extraction process; S2: Input an image into the cervical cell classification network framework. The input image is first divided into blocks by the patch segmentation module. Each 4×4 adjacent pixels is divided into a patch. S3: The image is flattened in the channel direction. For cervical cell images, which are generally RGB three-channel images, the size of the image will change after flattening. S4: In Stage 1, linear embedding is used for processing, and linear embedding is used to linearly transform the channel data of each pixel, transforming the image from the original dimension to the new dimension; S5: In Stage 2 to Stage 4, the patch merging layer is used for downsampling to reduce the image size while retaining key information to further improve the feature expression capability; S6: When the image passes through the CFA-Former module, the RMSLayerNorm layer is normalized, and global average pooling is performed before finally reaching the classification head for classification. The RMSLayerNorm layer is used to replace the traditional LayerNorm layer to speed up the convergence of the network and reduce the consumption of computing resources.

2. According to claim 1, a cervical cell classification method based on attention mechanism and Swin Transformer is characterized in that: The core components of the CFA-Former module include at least a CFA module and a W-MSA / SW-MSA module.

3. According to claim 2, a cervical cell classification method based on attention mechanism and Swin Transformer is characterized in that: The CFA module is a parallel attention mechanism. The CFA module effectively combines channel attention with spatial attention, that is, an effective combination of CDA and SFA, thereby enhancing the feature learning ability of the model. This parallel structural design allows the model to simultaneously evaluate the criticality of the channel dimension and the spatial dimension, thereby avoiding the risk that a single attention mechanism may ignore certain important dimensions. Through weighted fusion, a more accurate feature representation can ultimately be obtained.

4. According to claim 3, a cervical cell classification method based on attention mechanism and Swin Transformer is characterized in that: The CFA module generates a corresponding 3D attention map for the input 3D image, highlighting important features; The process of generating a 3D attention map is decomposed into two lightweight branches, each of which uses simplified components, effectively reducing the amount of computation and parameter consumption; Each element of the feature map can be regarded as a feature detector, so the two branches of channel attention and spatial attention can clearly learn "which features to pay attention to" and "where to focus on", further optimizing the feature extraction process; The CFA module information transmission process is as follows: CFAM=σ(CDA+SFA) CFA=(x*CFAM)+x Among them, CFAM represents the attention feature map output after the fusion of channel attention and spatial attention, σ represents the Sigmoid function, CFA represents the output of the CFA module, and x represents the feature vector input to the CFA module.

5. According to claim 4, a cervical cell classification method based on attention mechanism and Swin Transformer is characterized in that: The CDA adopts a lightweight and efficient channel attention mechanism, which models the importance of each channel through adaptive pooling and adaptive convolution operations, without relying on complex fully connected layers, significantly reducing the computational complexity and the number of parameters that need to be learned, ensuring that the network can still maintain efficient computing performance when processing high-dimensional feature spaces; The CDA has a convolution operation that can adaptively select the convolution kernel size according to the channel dimension of the input feature map; Among them, k represents the amount of information transmission between the current channel and other k different channels. This value is dynamically determined by a nonlinear mapping related to the number of channels, so that the convolution operation can be flexibly adjusted to meet the information interaction requirements between different channels; Since the input channels of cervical cell images are generally 3, they become 96 after the first stage of Liner Embedding. The number of channels doubles in each stage thereafter, so the number of channels is an even number. As the subject of the mapping; Therefore, the mapping rule is given by: Where log(C) represents the natural logarithm of the number of channels C, b represents the offset in the general linear mapping, which is set to 1 in this paper. It means rounding down to the nearest odd number, and the convolution kernel size is generally set to an odd number; The overall process within the CDA is as follows: Among them, AAP stands for adaptive average pooling, represents an adaptive k×k convolution, σ represents the Sigmoid function, and expand() represents expanding the output to the same shape as the input.

6. The cervical cell classification method based on attention mechanism and Swin Transformer according to claim 5, characterized in that: The SFA adopts the Bottleneck structure in ResNet, that is, SFA adopts dilated convolution to expand the receptive field in a low-cost and high-efficiency manner, which can save both the number of parameters and the computational overhead.

7. The cervical cell classification method based on attention mechanism and Swin Transformer according to claim 6, characterized in that: The application of the SFA comprises at least the following steps: First, a 1×1 convolution is used to reduce the dimension of the feature map C×H×W projection to The feature maps are integrated and compressed across the channel dimension, and the compression ratio reduction is set to 16 in this paper; After dimensionality reduction, two dilated convolutions are applied to effectively capture and utilize contextual information; Finally, the feature map is compressed to 1×H×W again using 1×1 convolution; For the final output scale adjustment, expand() is used at the tail end to expand the feature map to the same size as the CDA branch so that the two branches can be fused; The calculation process of the SFA is: Among them, f 1xl represents a convolution operation with a convolution kernel size of 1×1, represents a 3×3 atrous convolution, d is the dilation parameter of the atrous convolution, which is set to 4, and the padding is also set to 4.

Citation Information

Patent Citations

  • Abnormal cell identification method and system based on cervical cytology image

    CN112215117A

  • Cervical cell image classification method based on fine granularity

    CN112990118A

  • Fast image segmentation method based on multi-scale Transform

    CN117173205A