Cervical abnormal cell detection method based on multi-scale feature fusion
By adopting multi-scale feature fusion strategy, cross-scale pooling model (CSPM) and multi-scale fusion attention module (MSFA) in the detection of cervical abnormal cells, the shortcomings of existing methods in small-objective detection, multi-scale adaptability and feature interaction are solved, and higher detection accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510015048.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-30
AI Technical Summary
The existing cervical abnormal cell detection methods have shortcomings in small-object detection, multi-scale adaptability, feature interaction and context modeling, resulting in limited detection accuracy and robustness.
A multi-scale feature fusion method is proposed to detect cervical abnormal cells. By constructing a multi-scale feature extraction module and feature fusion strategy, local details, global background and intercellular relationships are integrated, and the cross-scale pooling model (CSPM) and multi-scale fusion attention module (MSFA) are used to improve the performance of the detection model.
The performance of the detection model in different scales and complex scenarios is improved, the characteristic interaction between cells and the utilization of global context information is enhanced, and the detection accuracy and robustness are significantly improved.
Smart Images

Figure CN120071337A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of medical image processing and deep learning, and particularly to a method for detecting abnormal cervical cells with multi-scale feature fusion. Background Art
[0002] Cervical cancer is the fourth most common cancer among women globally, with approximately 604,000 confirmed cases and 342,000 deaths in 2020. Notably, the development of cervical cancer is relatively slow - it typically takes about 10 years from high-risk HPV infection to precancerous lesions and then to invasive cancer. However, if detected in the early stage, the disease is highly curable. This relatively long progression period provides a valuable time window for effective screening and intervention, and timely screening has been proven to reduce the incidence by at least 60%. Detecting abnormal cervical cells is a crucial step in the screening process; however, manually examining cell slides is laborious, time-consuming, and highly subjective, which increases the likelihood of diagnostic errors. In addition, due to the fact that abnormal cells only account for a very small proportion in the sample images, this inefficiency also leads to a waste of a large amount of medical resources.
[0003] With the development of image processing technology and deep learning, significant progress has been made in cervical cell detection technology. However, cervical cell images pose unique challenges, including a high degree of similarity between cells, subtle differences between different cell categories, significant variability within categories, and the diversity of cell sizes. In addition, the complexity of the microscope imaging environment further increases the difficulty of detection. Traditional methods rely on manually extracting features in the segmentation and classification steps, and are vulnerable to the influence of segmentation accuracy, resulting in a decline in detection accuracy.
[0004] Currently, end-to-end object detection methods provide new ideas for cervical cell detection, but existing methods still have deficiencies in dealing with the specific characteristics of cervical cells. For example, the detection sensitivity for small targets is relatively low, the adaptability to scale changes is limited, and there are defects in modeling the feature interaction and context information between cells. These problems limit the performance of the detection model in the task of identifying abnormal cervical cells.
[0005] Existing methods for detecting abnormal cervical cells have deficiencies in small target detection, multi-scale adaptability, and feature interaction and context modeling. Specifically, the sensitivity to small targets is relatively low, there is a lack of effective processing for targets with significantly different sizes, and the associated features and global context information between cells are not fully utilized, resulting in limited detection accuracy and robustness;
[0006] Therefore, a new solution to the above problems is needed. Summary of the Invention
[0007] The purpose of the present invention is to provide a method for detecting cervical abnormal cells with multi-scale feature fusion. It mainly focuses on the problems of large inter-class gap and small intra-class gap among cervical cells, as well as large variations in cell size, and focuses on multi-scale feature processing and fusion. A multi-scale feature fusion detection network (MSFF-Net) that can integrate local details, global background, and inter-cell relationships is proposed to improve the performance of the detection model in different scales and complex scenarios and solve the limitations of existing methods.
[0008] To achieve the above object, the present invention provides the following technical solutions: A method for detecting cervical abnormal cells with multi-scale feature fusion, which at least includes the following steps:
[0009] S1: Divide the cervical cell dataset into a training set and a test set. The main task in the model training process is to train the cervical pathological image cell detection model;
[0010] S2: Construct a cervical cell multi-scale feature extraction module. By using a multi-scale feature extraction strategy, capture feature information of different scales from cervical cell images. The feature information through the different scales of feature information at least includes local features and global features. The local features capture cell textures, and the global features extract cell structures and context information. Through a pyramid pooling structure or a similar multi-scale mechanism, achieve a comprehensive perception of cell size and morphological changes;
[0011] S3: Multi-scale feature fusion of cervical cell images. On the basis of feature extraction, adopt a feature fusion strategy to fuse shallow local features with deep global semantic features. This can not only enhance the feature interaction between cells but also effectively alleviate the problems of large differences among similar cells and high similarity among different cells. The fused features are more representative, improving the detection accuracy and robustness of the model. The local features capture cell textures, and the global features extract cell structures and context information;
[0012] S4: Detection of cervical abnormal cells. Based on the fused features, detect abnormal cells. The detection process combines a classification model to identify whether there are abnormal cells and mark the abnormal regions. The output results of the cervical pathological image cell detection model are fed back to the training process to improve the detection accuracy by further optimizing the model;
[0013] S5: Cervical cell prediction process. Input the cervical cell image to be detected, and use the multi-scale feature extraction and fusion network optimized in the training stage to predict abnormal cells. The model can efficiently detect abnormal cells of different sizes and shapes and provide more accurate results at the same time.
[0014] Further, the cervical cell multi-scale feature extraction module in S2 at least includes a backbone network, a neck, and a detection head;
[0015] The application of the cervical cell multi-scale feature extraction module comprises at least the following steps:
[0016] First, the input image is preprocessed, and the preprocessing includes at least a Mosaic high-order data enhancement strategy;
[0017] Features are extracted from images through a novel backbone network. The feature extraction process includes convolution, deconvolution, a connected bottleneck structure, and cross-scale pooling. The cross-scale pooling is CSPM. The CSPM is used to mine the connections between cells. The CSPM integrates multiple convolutional layers and multi-level pooling layers to more effectively capture the local and global features of cells in the feature extraction stage, so as to improve the detection accuracy and reliability of diseased cervical cells, effectively extract the local and global information of cervical cell images, and generate multi-level feature maps suitable for different target scales.
[0018] By fusing shallow feature maps with high resolution but less semantic information and deep feature maps with low resolution but rich semantic information, a feature map with both high resolution and rich semantic information is generated, thereby improving the performance of object detection.
[0019] The CSPM module combines multiple grouped convolutional blocks and three serial maximum pooling layers with different receptive fields, which enhances the model's expressiveness while reducing the number of parameters. The specific structure is as follows: Figure 2 As shown in a;
[0020] The neck part adopts a structure combining FPN and PAN to fuse deep features with shallow features, and selectively focuses on more important features through the MSFA module;
[0021] By using convolution kernels of different sizes, multi-scale feature information can be more efficiently fused in the feature fusion stage, making the fused features richer;
[0022] Finally, the cervical diseased cells are detected through the detection head part.
[0023] Furthermore, the cross-scale pooling (CSPM) designed in S2 aims to address the limitations of the YOLOv8 feature extraction stage in complex feature representation and long-distance context information capture. Especially when dealing with long-distance dependencies, it is difficult to meet the needs of pathologists for comprehensive judgment of surrounding cells when reading films. CSPM more efficiently integrates local and global features by deeply exploring the correlation between cells.
[0024] The application process of the cross-scale pooling at least includes:
[0025] Split the feature map in the channel dimension, divide the input feature map into two parts along the channel direction and perform different operations on each part, and then reassemble the two parts;
[0026] Extract the feature of a part of the implementation basis, and perform pooling operations of different scales on the other part after multiple convolutions, so as to obtain features from receptive fields of different scales and improve the detection level of the model for multi-scale targets;
[0027] This multi-scale pooling strategy adopts the way of serially connecting pooling layers of different sizes, enabling the feature map to gradually accumulate context information of different scales during the pooling process layer by layer. By using three pooling layers of different sizes to be responsible for capturing detailed features, medium-range features, and long-range context information respectively, and splicing the detailed features, medium-range features, and long-range context information together, a feature map rich in scale information can be generated, which contains both fine local details and complete global information and context content;
[0028] At the same time, combined with convolutional layers of various different specifications, further refine and strengthen the presentation of complex features in the image;
[0029] Re-integrate the feature maps of the two branches, so as to be able to shape a more complete feature representation and achieve the effect of enhancing object detection.
[0030] In the cervical cell dataset, the CSPM module can more accurately extract local and global features by fusing convolution and multi-level pooling, thereby improving the accuracy and reliability of detecting diseased cells.
[0031] Furthermore, the multi-scale feature fusion of cervical cell images in S3 is based on MSFA in the MSFF-Net architecture. The MSFF-Net has three detection heads, which are respectively used to detect targets of three different sizes: 8×8, 16×16, and 32×32. Therefore, in the MSFA before the three detection heads, the kernel size of each in the multi-branch depth convolution is also different, adjusted according to the size of the detection target of the corresponding detection head. In each branch, two depthwise strip convolutions are used to approximate the standard depth convolution. On the one hand, it can greatly reduce the amount of calculation and the number of parameters, thereby improving the training and inference speed of the model. On the other hand, since cervical diseased cells often show irregular shapes and sizes, the strip depth convolution, due to its directional kernel design, can better capture the edges and directional features of these cells, thus helping to distinguish normal cells from abnormal cells.
[0032] Furthermore, the application of the MSFA at least includes the following steps:
[0033] First, Dconv 5×5 uses depthwise convolution to reduce the number of parameters and aggregate local information at the same time;
[0034] Next, multi-branch depth convolution kernels of different sizes are used to capture feature information in different directions (horizontal and vertical);
[0035] Since the three detection heads mainly target three different sizes of objects, the MSFA used before each detection head is also a multi-branch depth convolution of different sizes;
[0036] Then, Conv 1×1 is used to model the relationship between different channels and use the obtained attention weights, and multiply each channel of the original feature map to obtain the channel feature map after attention weighting, which will emphasize the channels helpful for the current task and suppress irrelevant channels;
[0037] In addition, referring to the spatial attention, it can strengthen the accuracy of abnormal cell localization;
[0038] Avgpool and Maxpool are used to extract local features and maximum features respectively to generate features of different context scales;
[0039] The obtained features are concatenated along the channel dimension to obtain a feature map with different scale context information, and the features are further integrated through Conv 7×7 to generate attention weights;
[0040] Finally, the obtained spatial attention weights are applied to the original feature map to weight the features at each spatial position, which can highlight important image regions and reduce the influence of unimportant regions. The process description formula is as follows:
[0041]
[0042]
[0043]
[0044] M s (F″) = σ(Conv 7×7 ([Avgpool(F′); Maxpool(F′)])) (4)
[0045]
[0046] where F represents the input feature, F′ represents Figure 2 the output result after the element-wise matrix multiplication operation in b, F″ is the final output, is the element-wise matrix multiplication operation, σ is processed by the sigmoid activation function. DConv represents depth convolution, Scale i , i ∈ {1, 2, 3} represents the i-th branch of the multi-branch depth convolution stage used. Mc (F) represents Figure 2 The output after multi-branch convolution and 1×1 convolution processing in b, M s (F″) represents Figure 2 The output result obtained by processing through the spatial attention branch in the second half of b.
[0047] Furthermore, the loss function adopted by the optimization model in S4 is the Inner-CIoU loss function, and the Inner-CIoU loss function is formed by further optimizing the CIoU loss function.
[0048] Furthermore, the application process of the Inner-CIoU loss function at least includes the following steps:
[0049] The size of the auxiliary bounding box introduced in Inner-CIoU is controlled by the scale factor ratio. ratio is the scale scaling factor. When ratio is 1, the size of the auxiliary bounding box is equal to the size of the actual bounding box, and ratio is set to 0.7;
[0050] Its calculation method is as follows:
[0051]
[0052]
[0053]
[0054]
[0055]
[0056] union=(w gt ×h gt )×(ratio) 2 +(w×h)×(ratio) 2 -inter (11)
[0057]
[0058] where and represent the center coordinates of the ground truth box (GT), x c and y c represent the center coordinates of the predicted box (anchor), w gt and h gt represent the width and height of the ground truth box (GT), and w and h represent the width and height of the predicted box (anchor). ratio represents the scaling ratio factor of the internal auxiliary box. b t 、bb , b l , b r respectively represent the upper and lower boundaries and the left and right boundaries of the inner box. inter represents the intersection area between the inner auxiliary boxes, union represents the union area between the inner auxiliary boxes, and IoU inner represents the intersection over union of the inner auxiliary boxes.
[0059] Compared with the ordinary IoU loss, when ratio is less than 1, the size of the auxiliary bounding box is smaller than the actual bounding box, and its effective regression range is smaller than that of the IoU loss. However, the absolute value of its gradient is greater than the gradient obtained from the IoU loss, which can accelerate the convergence of high-IoU samples;
[0060] When ratio is greater than 1, the larger-scale auxiliary bounding box expands the effective regression range, which is beneficial to the regression of low-IoU;
[0061] The mathematical expression of the Inner-CIoU loss function is as follows:
[0062]
[0063]
[0064]
[0065] where α is a weight function, used to measure the aspect ratio, ρ 2 (b, b gt ) represents the Euclidean distance between the center point b of the predicted bounding box and the center point b of the ground-truth bounding box gt ;
[0066] Inner-CIoU not only introduces auxiliary bounding boxes, but also comprehensively considers the distance between the center points of the bounding boxes and the aspect ratio of the bounding boxes.
[0067] Furthermore, three detection heads for different scales are used in S4 to detect cervical lesion cells of different scales, where:
[0068] The small-object detection head inputs a high-resolution feature map and is specifically used to detect small lesion cells in pathological images, such as single abnormal cells or smaller lesion areas;
[0069] The medium-object detection head inputs a medium-resolution feature map and is used to process medium-sized lesion areas, such as cell clusters composed of several abnormal cells;
[0070] The large-object detection head inputs a low-resolution feature map and is responsible for detecting large-area lesion areas, such as larger lesions formed by the aggregation of multiple lesion cells.
[0071] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0072] First, the present invention uses a cervical lesion cell detection method based on multi-scale feature fusion. By proposing a cross-scale pooling model (CSPM), it realizes the efficient capture of multi-scale features from local details to overall structures in cervical lesion cell images. CSPM can dynamically extract features between different scales, making full use of context information, thereby compensating for the deficiency in the expression ability of single-scale features and effectively dealing with the problems of large cell size variations and complex backgrounds (such as cell adhesion and occlusion). Secondly, the designed multi-scale fusion attention module (MSFA) further enhances the feature fusion ability. MSFA can adaptively adjust the weights of different-scale features according to the importance of cell features, fully integrating local and global information, thus improving the network's perception and recognition ability of cervical abnormal cells, especially showing higher robustness when dealing with minor differences between cells and significant variability within categories.
[0073] Furthermore, through the synergistic effect of CSPM and MSFA, the present invention not only improves the adaptability of the detection model in the scenario of cell size changes but also significantly improves the separation ability of lesion cells in complex scenarios, ultimately achieving the simultaneous improvement of the detection accuracy and efficiency of cervical lesion cells. Other advantages and objectives of the present invention will be further elaborated in the subsequent description. These improvements will provide new solutions for medical image detection and have broad application value. Description of the Drawings
[0074] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0075] Figure 1 It is the flowchart of the example of the present invention;
[0076] Figure 2 It is the model diagram of the detection network (MSFF-Net) based on multi-scale feature fusion constructed in the example of the present invention;
[0077] Figure 3 It is the size statistical histogram and cell comparison diagram of the cervical cell image dataset used;
[0078] Figure 4 It is the feature comparison diagram obtained by the multi-scale feature extraction detection network of the example of the present invention;
[0079] Figure 5The following is the rendering effect of MSFF-Net and YOLOv8 in the examples of the present invention. Detailed implementation mode
[0080] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0081] As Figure 3 , the statistical results of the cervical cell dataset images used in the present invention. From the figure, we can observe that in addition, the cell sizes in cervical cell images vary greatly, and some cells may be very small (for example Figure 3 b), and some cells are larger in size (for example Figure 3 c). And cervical cells have the characteristics of small inter-class differences and large intra-class differences. As Figure 3 (d) shows cells of two different lesion categories, but they look similar, Figure 3 (e) shows two cells of the same lesion category, but they have large appearance differences. Therefore, relying solely on local reasoning is often insufficient. Clinically, when examining slides, cytopathologists usually use the surrounding cells as a reference to compare with the target cells to determine whether they are normal or abnormal. In order to imitate the behavior of pathologists when examining slides and enhance the feature interaction between cells, the present invention constructs a cervical lesion cell detection method based on multi-scale feature fusion, and its specific network model structure is as Figure 2 shown.
[0082] Based on the above, the following is proposed a method for detecting abnormal cervical cells with multi-scale feature fusion;
[0083] Please refer to Figure 1 , a method for detecting abnormal cervical cells with multi-scale feature fusion, at least including the following steps:
[0084] S1: Divide the cervical cell dataset into a training set and a test set. The main task in the model training process is to train the cervical pathological image cell detection model;
[0085] S2: Construct a cervical cell multi-scale feature extraction module. By using a multi-scale feature extraction strategy, capture feature information of different scales from cervical cell images. The feature information includes at least local features and global features through different scales of feature information;
[0086] S3: Multi-scale feature fusion of cervical cell images. Based on feature extraction, a feature fusion strategy is adopted to fuse shallow local features and deep global semantic features, which can not only enhance the feature interaction between cells, but also effectively alleviate the problems of large differences among similar cells and high similarity among different types of cells. The fused features are more representative, improving the detection accuracy and robustness of the model;
[0087] S4: Detection of cervical abnormal cells. Based on the fused features, the detection of abnormal cells is carried out. The detection process combines a classification model to identify whether there are abnormal cells and mark the abnormal areas. The output results of the cervical pathological image cell detection model are fed back to the training process to improve the detection accuracy by further optimizing the model;
[0088] S5: Cervical cell prediction process. Input the cervical cell image to be detected, and use the optimized multi-scale feature extraction and fusion network in the training stage to predict abnormal cells. The model can efficiently detect abnormal cells of different sizes and shapes and provide more accurate results at the same time.
[0089] Figure (2) is a specific detection network model based on multi-scale feature fusion. It includes the following specific steps:
[0090] This model mainly covers three parts: the backbone network, the neck, and the detection head. First, preprocess the input image with Mosaic high-order data augmentation strategy, etc. The backbone network is responsible for extracting features from the image, including a series of convolutions, deconvolutions, bottleneck structures with connections, and cross-scale pooling (CSPM). By fusing shallow feature maps with high resolution but less semantic information and deep feature maps with low resolution but rich semantic information, feature maps with both high resolution and rich semantic information are generated, thus improving the performance of object detection. The CSPM module combines multiple grouped convolution blocks (including 1×1 and 3×3 convolutions) and three serial maximum pooling layers with different receptive fields, reducing the number of parameters while enhancing the model's expressive ability. The neck part adopts a structure combining FPN and PAN to fuse deep features and shallow features, and selectively focuses on more important features through the MSFA module. By using convolution kernels of different sizes, multi-scale feature information can be more efficiently fused in the feature fusion stage, making the fused features more abundant. Finally, the detection head part detects cervical lesion cells.
[0091] Among them, the cross-scale pooling constructed in S2 is mainly to solve the problem of insufficient performance of the original SPPF (Spatial Pyramid Pooling-Fast) in YOLOv8 in tasks that require complex feature representation and context information capture. In particular, there are certain limitations in the ability to capture long-distance context information. The CSPM constructed in the present invention can better explore cross-cell relationships and fully capture local and global features.
[0092] The structure of the CSPM module is shown in Figure (2a). First, the feature map is segmented in the channel dimension, which can maintain the integrity of information while reducing the amount of computation and the number of parameters. It splits the feature map in the channel dimension, divides the input feature map into two parts along the channel direction and performs different operations separately, and then recombines the two. One part performs basic feature extraction work, and the other part uses pooling operations of different scales after passing through multiple convolutions to obtain features from receptive fields of different scales, thereby improving the model's detection level for multi-scale targets. This multi-scale pooling strategy uses a serial connection of pooling layers of different sizes, enabling the feature map to gradually accumulate context information of different scales during the pooling process layer by layer. Three pooling layers of different sizes are responsible for capturing detailed features, medium-range features, and long-distance context information respectively. By splicing the features extracted by these multi-scale pooling layers, a feature map rich in scale information can be generated, which contains both fine local details and complete global information and context content. At the same time, the module combines convolutional layers of various different specifications to further refine and strengthen the presentation of complex features in the image. By re-integrating the feature maps of the two branches, the CSPM module can form a richer feature representation and improve the performance of object detection. In the application of the cervical cell dataset, the CSPM module can more effectively capture the local and global features of cells during the feature extraction stage by integrating multiple convolutional layers and multi-level pooling layers, so as to improve the detection accuracy and reliability of diseased cervical cells.
[0093] In step S3, the multi-scale fusion attention mechanism constructed in the present invention is shown in Figure (2b). The MSFA mainly includes the following parts:
[0094] The MSFA in the multi-scale feature fusion of cervical cell images based on the MSFF-Net architecture. The MSFF-Net has three detection heads, which are respectively used to detect targets of three different sizes, 8×8, 16×16, and 32×32. Therefore, in the MSFA before the three detection heads, the kernel size of each in the multi-branch depth convolution is also different, adjusted according to the size of the detection target of the corresponding detection head. In each branch, two depthwise strip convolutions are used to approximate the standard depth convolution. On the one hand, it can greatly reduce the amount of calculation and the number of parameters, thus improving the training and inference speed of the model. On the other hand, since cervical lesion cells often show irregular shapes and sizes, the strip depth convolution, due to its directional kernel design, can better capture the edges and directional features of these cells, thus helping to distinguish normal cells from abnormal cells.
[0095] The application of MSFA includes at least the following steps:
[0096] First, Dconv 5×5 uses depthwise convolution to reduce the number of parameters while aggregating local information;
[0097] Then, multi-branch depth convolution kernels of different sizes are used to capture feature information in different directions (horizontal and vertical);
[0098] Since the three detection heads are mainly for detecting targets of three different sizes, the MSFA used before each detection head is also a multi-branch depth convolution of different sizes;
[0099] Then Conv 1×1 is used to model the relationship between different channels and use the obtained attention weights, and multiply each channel of the original feature map to obtain the attention-weighted channel feature map, which will emphasize the channels helpful for the current task and suppress irrelevant channels;
[0100] In addition, referring to the spatial attention, it can strengthen the accuracy of abnormal cell localization;
[0101] Avgpool and Maxpool are used respectively to extract local features and maximum features for generating features of different context scales;
[0102] The obtained features are concatenated along the channel dimension to obtain a feature map with different-scale context information, and the features are further integrated by Conv 7×7 to generate attention weights;
[0103] Finally, the obtained spatial attention weights are applied to the original feature map to weight the features at each spatial position, which can highlight important image regions and reduce the influence of unimportant regions. The process description formula is as follows:
[0104]
[0105]
[0106]
[0107] M s = σ(Conv 7×7 ([Avgpool(F′); Maxpool(F′)])) (4)
[0108]
[0109] where F represents the input feature, and F′ represents Figure 2 the output result after element-wise matrix multiplication in b, F″ is the final output, is the element-wise matrix multiplication, σ is processed by the sigmoid activation function. DConv represents depth convolution, Scale i , i ∈ {1, 2, 3} represents the i-th branch of the multi-branch depth convolution stage used. M c (F) represents Figure 2 the output after multi-branch convolution and 1×1 convolution processing of the result in b, M s (F″) represents Figure 2 the output result obtained by processing through the spatial attention branch in the second half of b.
[0110] The loss function adopted by the optimized model in S4 is the Inner-CIoU loss function, and the Inner-CIoU loss function is formed by further optimizing the CIoU loss function.
[0111] Furthermore, the application process of the Inner-CIoU loss function at least includes the following steps:
[0112] The size of the auxiliary bounding box introduced in Inner-CIoU is controlled by the scale factor ratio. ratio is the scale scaling factor. When ratio is 1, the size of the auxiliary bounding box is equal to the actual bounding box size, and ratio is set to 0.7;
[0113] Its calculation method is as follows:
[0114]
[0115]
[0116]
[0117]
[0118]
[0119] union = (w gt ×h gt ) × (ratio) 2 + (w × h) × (ratio) 2 - inter(11)
[0120]
[0121] where and represent the center coordinates of the ground truth (GT), x c and y c represent the center coordinates of the predicted box (anchor), w gt and h gt represent the width and height of the ground truth (GT), and w and h represent the width and height of the predicted box (anchor). ratio represents the scaling factor of the internal auxiliary box. b t 、b b 、b l 、b r represent the upper and lower boundaries and the left and right boundaries of the internal box respectively. inter represents the intersection area between the internal auxiliary boxes, union represents the union area between the internal auxiliary boxes, and IoU inner represents the intersection over union of the internal auxiliary boxes.
[0122] Compared with the ordinary IoU loss, when ratio is less than 1, the size of the auxiliary bounding box is smaller than the actual bounding box, and its effective regression range is smaller than the IoU loss. However, the absolute value of its gradient is greater than the gradient obtained from the IoU loss, which can accelerate the convergence of high-IoU samples;
[0123] When ratio is greater than 1, the larger-scale auxiliary bounding box expands the effective regression range and is beneficial to the regression of low IoU;
[0124] The mathematical expression of the Inner-CIoU loss function is as follows:
[0125]
[0126]
[0127]
[0128] where α is the weight function, used to measure the aspect ratio, ρ 2 (b, b gt ) represents the Euclidean distance between the center point b of the predicted box and the center point b of the ground truth box; gt
[0129] Inner-CIoU not only introduces auxiliary bounding boxes, but also comprehensively considers the distance between the center points of the bounding boxes and the aspect ratios of the bounding boxes.
[0130] Furthermore, in S4, three detection heads for different scales are used to detect cervical lesion cells of different scales, where:
[0131] The small-object detection head inputs a high-resolution feature map and is specifically used to detect small lesion cells in pathological images, such as single abnormal cells or smaller lesion areas;
[0132] The medium-object detection head inputs a medium-resolution feature map and is used to process medium-sized lesion areas, such as cell clusters composed of several abnormal cells;
[0133] The large-object detection head inputs a low-resolution feature map and is responsible for detecting large-area lesion areas, such as larger lesion cell masses formed by the aggregation of multiple lesion cells.
[0134] Furthermore, the following experimental proofs are proposed:
[0135] In Table 1 below, the present invention designs a set of ablation experiments to prove the effectiveness of the proposed methods;
[0136] Table 1 Ablation Experiments
[0137]
[0138] The present invention evaluates the contributions of several important elements to the method of the present invention, including CSPM, MSFA, and Inner-CIoU. The ablation experiments gradually integrate these three methods into the baseline YOLOv8 to enhance the feature extraction and utilization capabilities of the network. Table 1 shows the effects of sequentially embedding CSPM, MSFA, and Inner-CIoU on the detection performance of the cervical lesion cell detection model.
[0139] For the convenience of performance comparison, this experiment mainly uses mAP@0.5(%) as the main evaluation index. The experimental results in Table 1 show that when applied to the baseline model, each strategy proposed by the present invention improves the performance of detecting cervical lesion cells to varying degrees. The parallel calculation of CSPM and MSFA proves that the combined embedding has better detection performance compared with a single module, with mAP@0.5(%) increased by 2.2% and mAP@0.5:0.95(%) increased by 2.0%. The best performance is achieved when CSPM, MSFA, and Inner-CIoU are combined and embedded. This indicates that the three methods proposed in the embodiments of the present invention are all effective for the detection of cervical abnormal cells.
[0140] In addition, to more intuitively demonstrate the effectiveness of the examples of the present invention, heatmaps are also used to visualize some outputs of YOLOv8 and MSFF-Net. The heatmap visualizations of two instances in the dataset are as Figure 5 shown. The attention of YOLOv8 is relatively dispersed, and there is a situation where it also focuses on the background. While MSFF-Net only shows good focus on all abnormal cells. Respectively as Figure 5 a and Figure 5 b shown. This indicates that when using MSFF-Net to detect cervical lesion cells, the model can better focus on the characteristics of cervical lesion cells.
[0141] In the present invention, the Mosaic data augmentation strategy and multi-scale training method are adopted in the preprocessing stage to improve the generalization ability of the model. During the training process, the preprocessed training set images are input into the MSFF-Net network model for training, and the model structure is as Figure 2 shown. The initial weights of MSFF-Net are based on the pre-trained model of YOLOv8 on the COCO dataset. After training, the best model weights are saved and the experimental results are recorded. In the prediction stage of cervical cell images, the trained optimal MSFF-Net model is used to detect cervical abnormal cells and generate detection result images. The detection effects generated by the experiment are as Figure 4 shown, where Ground Truth represents the original annotation map of cervical cells, which is used to compare and evaluate the accuracy of the detection results.
[0142] In Table 2, the present invention designs a set of comparative experiments to compare the detection effects of the original YOLOv8 and the MSFF-Net model constructed by the examples of the present invention for cervical abnormal cells.
[0143] Table 2 Comparison of the effects of the baseline YOLOv8 network model and the MSFF-Net network model
[0144]
[0145] In addition, to verify the effectiveness of this method for detecting various abnormal cells, a set of comparative experiments on the detection effects of various lesion cells are also designed, as shown in Table 3.
[0146] Table 3 Comparison of the detection effects of the baseline YOLOv8 network model and the MSFF-Net network model for each lesion category
[0147]
[0148] The detection effects for LSIL, CAND, and FLORA in Table 3 have been improved to a large extent. Among them, the sizes of lsil cells vary greatly, including both large cell clusters and small single cells. The improvement in its detection effect further proves the effectiveness of the MSFA attention mechanism we proposed in dealing with different scale problems. Both CAND and FLORA have the problem of large intra-class differences, and the number of these two types of samples in the cervical cell dataset we used is small. The improvement in the detection accuracy of these two lesion categories also well illustrates the effectiveness of the MSFF-Net method we proposed, which can better fuse context information, obtain more effective features, and thus improve the detection performance. The original YOLOv8 had worse detection effects for ASCH and HSIL, mainly because most of these cells are small in size, and the surrounding environment in cervical cell images is complex, making them vulnerable to occlusion and other influences. It can be seen from the results that the Inner-IOU we adopted has improved these problems to a certain extent. However, the detection effects for the two categories of ACTIN and HERPS have decreased, mainly because the number of samples in the dataset is too small, and the model may have difficulty generalizing well for this category. The model may have more advantages in learning other categories.
[0149] The overall results show that the cervical lesion cell detection method based on multi-scale feature fusion proposed in the embodiments of the present invention can enhance the feature interaction of cervical cell images, effectively solve the problems of small inter-class gap and large intra-class gap of cervical cells, and at the same time can solve the problem of large variation in the size of cervical cells.
[0150] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claimed rights.
Claims
1. A method for detecting abnormal cervical cells by multi-scale feature fusion, characterized in that: At least the following steps are included: S1: The cervical cell dataset is divided into a training set and a test set. The main task of the model training process is to train the cervical pathology image cell detection model; S2: construct a cervical cell multi-scale feature extraction module, and capture feature information of different scales from the cervical cell image by using a multi-scale feature extraction strategy, and the feature information includes at least local features and global features through the feature information of different scales; S3: Multi-scale feature fusion of cervical cell images. Based on feature extraction, a feature fusion strategy is adopted to fuse shallow local features with deep global semantic features. This not only enhances the feature interaction between cells, but also effectively alleviates the problem of large differences between cells of the same type and high similarity between cells of different types. The fused features are more representative, improving the detection accuracy and robustness of the model. S4: Detection of abnormal cervical cells. Detection of abnormal cells is performed based on the fused features. The detection process is combined with the classification model to identify whether the cells are abnormal and mark the abnormal areas. The output results of the cervical pathology image cell detection model are fed back to the training process to improve the detection accuracy by further optimizing the model. S5: Cervical cell prediction process: input the cervical cell image to be detected, and use the multi-scale feature extraction and fusion network optimized in the training stage to predict abnormal cells. The model can efficiently detect abnormal cells of different sizes and shapes, and provide more accurate results.
2. The method for detecting abnormal cervical cells by multi-scale feature fusion according to claim 1, characterized in that: The cervical cell multi-scale feature extraction module in S2 at least includes a backbone network, a neck and a detection head; The application of the cervical cell multi-scale feature extraction module comprises at least the following steps: First, the input image is preprocessed, and the preprocessing includes at least a Mosaic high-order data enhancement strategy; Extracting features from images through a novel backbone network. The feature extraction process includes convolution, deconvolution, a connected bottleneck structure, and cross-scale pooling. The cross-scale pooling is CSPM. The CSPM is used to mine the connections between cells. The CSPM integrates multiple convolutional layers and multi-level pooling layers to more effectively capture the local and global features of cells in the feature extraction stage, so as to improve the detection accuracy and reliability of diseased cervical cells. By fusing shallow feature maps with high resolution but less semantic information and deep feature maps with low resolution but rich semantic information, a feature map with both high resolution and rich semantic information is generated, thereby improving the performance of object detection. The CSPM module combines multiple grouped convolutional blocks and three serial maximum pooling layers with different receptive fields, which enhances the model's expressiveness while reducing the number of parameters. The specific structure is shown in Figure 2a. The neck part adopts a structure combining FPN and PAN to fuse deep features with shallow features, and selectively focuses on more important features through the MSFA module; By using convolution kernels of different sizes, multi-scale feature information can be more efficiently fused in the feature fusion stage, making the fused features richer; Finally, the cervical diseased cells are detected through the detection head part.
3. The method for detecting abnormal cervical cells by multi-scale feature fusion according to claim 2, characterized in that: The application process of the cross-scale pooling at least includes: Split the feature map in the channel dimension, divide the input feature map into two parts along the channel direction and perform different operations on each part, and then reassemble the two parts; One part is used for basic feature extraction, and the other part is used for pooling operations of different scales after multiple convolutions to obtain features from receptive fields of different scales, thereby improving the model's detection level for multi-scale targets. This multi-scale pooling strategy uses a serial connection of pooling layers of different sizes, so that the feature map gradually accumulates context information of different scales in the layer-by-layer pooling process. Three pooling layers of different sizes are responsible for capturing detail features, medium-range features, and long-distance context information respectively. By splicing detail features, medium-range features, and long-distance context information together, a feature map rich in scale information can be generated, which contains both fine local details and complete global information and context content. At the same time, a variety of convolutional layers of different specifications are combined to further refine and enhance the complex features in the image; The feature maps of the two branches are reintegrated to create a more complete feature expression, thereby improving the effect of target detection.
4. The method for detecting abnormal cervical cells by multi-scale feature fusion according to claim 3, characterized in that: The multi-scale feature fusion of cervical cell images in S3 is based on MSFA in the MSFF-Net architecture. The MSFF-Net has three detection heads, which are used to detect targets of three different sizes: 8×8, 16×16 and 32×32. Therefore, in the MSFA before the three detection heads, the kernel size of each multi-branch deep convolution is also different, which is adjusted according to the detection target size of the corresponding detection head. In each branch, two deep strip convolutions are used to approximate the standard deep convolution. On the one hand, it can greatly reduce the amount of calculation and the number of parameters, thereby improving the training and inference speed of the model. On the other hand, since cervical lesion cells often show irregular shapes and sizes, the strip deep convolution can better capture the edge and directional characteristics of these cells due to its directional kernel design, thereby helping to distinguish normal cells from abnormal cells.
5. The method for detecting abnormal cervical cells by multi-scale feature fusion according to claim 4, characterized in that: The MSFA application is at least The following steps are involved: First, Dconv 5×5 uses depthwise convolution to reduce the number of parameters while aggregating local information; Next, multi-branch deep convolution kernels of different sizes are used to capture feature information in different directions; Since the three detection heads are mainly used to detect targets of three different sizes, the MSFA used before each detection head also uses multi-branch depth convolutions of different sizes; Then use Conv 1×1 to model the relationship between different channels and use the obtained attention weights and multiply them with each channel of the original feature map to obtain the attention-weighted channel feature map, which will emphasize the channels that are helpful for the current task and suppress irrelevant channels; In addition, referring to spatial attention, it is possible to enhance the accuracy of abnormal cell localization; Avgpool and Maxpool are used to extract local features and maximum features respectively, which are used to generate features of different context scales; The obtained features are concatenated along the channel dimension to obtain a feature map with context information of different scales, and the features are further integrated through Conv 7×7 to generate attention weights; Finally, the obtained spatial attention weights are applied to the original feature map to weight the features of each spatial position, which can highlight important image areas and reduce the influence of unimportant areas. The process description formula is as follows: M s (F ″ )=σ(Conv 7×7 ([Avgpool(F ′ );Maxpool(F ′ )])) (4) Where F represents the input feature, F ′ represents the output result after the element-by-element matrix multiplication operation in Figure 2b, F ″ is the final output, is an element-by-element matrix multiplication operation, σ is processed by the sigmoid activation function. DConv represents deep convolution, Scale i i∈{1,2,3} represents the i-th branch of the multi-branch depthwise convolution stage used; M c (F) represents the output after multiple branch convolutions and 1×1 convolutions in Figure 2b, M s (F ″ ) represents the output result obtained after the second half of the spatial attention branch processing in Figure 2b.
6. The method for detecting abnormal cervical cells by multi-scale feature fusion according to claim 5, characterized in that: The loss function used by the optimization model in S4 is the Inner-CIoU loss function, and the Inner-CIoU loss function is formed by further optimizing the CIoU loss function.
7. The method for detecting abnormal cervical cells by multi-scale feature fusion according to claim 6, characterized in that: The application process of the Inner-CIoU loss function includes at least the following steps: The size of the auxiliary border introduced in Inner-CIoU is controlled by the scale factor ratio. The ratio is the scale scaling factor. When the ratio is 1, the size of the auxiliary border is equal to the actual border size. The ratio is set to 0.
7. The calculation method is as follows: union=(w gt ×h gt )×(ratio) 2 +(w×h)×(ratio) 2 -inter (11) in and represents the center coordinate of the real box (GT), x c and c Represents the center coordinate w of the prediction box (anchor) gt and h gt Represents the width and height of the real box (GT), w and h represent the width and height of the predicted box (anchor); Ratio represents the scaling factor of the internal auxiliary frame; b t , b b , b l , b r They represent the upper and lower boundaries and the left and right boundaries of the internal box respectively; inter represents the intersection area between the internal auxiliary boxes, and union represents the union area between the internal auxiliary boxes. IoU inner Represents the intersection-combination ratio of the internal auxiliary frame; Compared with the ordinary IoU loss, when the ratio is less than 1, the auxiliary border size is smaller than the actual border, and its effective range of regression is smaller than the IoU loss, but its absolute value of the gradient is larger than the gradient obtained by the IoU loss, which can accelerate the convergence of high IoU samples; When the ratio is greater than 1, the larger auxiliary bounding box expands the effective range of regression and improves the regression of low Iou. The mathematical expression of the Inner-CIoU loss function is as follows: Where α is the weight function, θ is used to measure the aspect ratio, and ρ 2 (b,b gt ) represents the center point b of the predicted box and the center point b of the real box gt The Euclidean distance between Inner-CIoU not only introduces auxiliary bounding boxes, but also comprehensively considers the center point distance of the bounding box and the aspect ratio of the bounding box.
8. The method for detecting abnormal cervical cells by multi-scale feature fusion according to claim 1, characterized in that: The S4 uses three detection heads of different sizes to detect cervical lesion cells of different sizes, wherein: The small target detection head inputs a high-resolution feature map and is specifically designed to detect small diseased cells in pathological images, such as a single abnormal cell or a smaller lesion area. The medium-resolution feature map is input to the medium-target detection head, which is used to process medium-sized lesion areas, such as cell clusters composed of several abnormal cells. The large object detection head inputs a low-resolution feature map and is responsible for detecting large lesion areas, such as large lesions formed by the aggregation of multiple diseased cells.
Citation Information
Cited By
Medical image detection method based on cross-channel and cell density adaptive mechanism
CN121414749A
Medical image detection method based on cross-channel and cell density adaptive mechanism
CN121414749B