Cervical cancer cell detection method based on the Yolov5l model
By adding DMBST and CBS modules, feature map fusion and detection heads to the Yolov5l model, the problem of insufficient attention to the lack of end-to-end detection model and fine-grained feature in the prior art is solved, and the accuracy and reliability of cervical cancer cell detection is improved.
Patent Information
- Application Number
- CN202311311752.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-11
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-10-11
AI Technical Summary
The prior art cervical cancer cell detection methods lack end-to-end detection models and need to pay attention to fine-grained characteristics, resulting in poor detection results.
The cervical cancer cell detection method based on the Yolov5l model, by adding randomly zeroed multi-branch sliding window self-attention mechanism module DMBST and convolutional layer CBS in the backbone of the network, adding feature map fusion in the network neck Neck, and adding detection head Detect in the detection layer head to improve feature extraction capabilities.
Effectively extracting tiny detailed features in the image improves the accuracy and reliability of cervical cancer cell detection.
Smart Images

Figure CN117274220B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a cervical cancer cell detection method in the field of biomedical image processing, and in particular to a cervical cancer cell detection method based on a Yolov51 model. Background Art
[0002] The existing cervical cancer cell detection method needs to segment the cell nucleus before detection or classification, or only detect or classify. In order to detect or classify accurately, it is sometimes necessary to manually construct features and use these features as the basis for detection or classification, such as: the ratio of the cell nucleus to the cytoplasm, the diameter of the cell nucleus, the staining depth of the cell nucleus, etc. It can be seen that the technical effect of the existing cervical cancer cell detection method is unsatisfactory and lacks an end-to-end detection model. The so-called end-to-end detection model refers to a detection model that inputs a picture and obtains the output result through implicit feature extraction. The main reason for the above dilemma is that the differences in morphology, color and other features of cervical cells are much smaller than the differences between general detected objects, and cannot be well distinguished. More attention needs to be paid to fine-grained features to distinguish cervical cells. In addition, there are few sources of cervical cytology data sets and the annotation information is not unified, which increases the difficulty of cervical cell detection tasks to a considerable extent.
[0003] Yolov5 is one of the object detection algorithm model series developed by ultralytics. It is the fifth version based on the idea of "you only need to look once" (hereinafter referred to as the Yolov5 model). (The YOLO in the name is a combination of the initials of You Only LookOnce in English). The Yolov5 model can quickly and accurately detect multiple objects in an image and output their category labels and location coordinates. The Yolov5 model has five different parameter sizes: micro (n), small (s), medium (m), large (l) and extra large (x). The Yolov5 model can run on a variety of computer devices and operating systems, including desktop computers, servers, embedded systems, and cloud platforms. It needs to support deep learning frameworks (such as PyTorch) and related software libraries. The Yolov5 model has high speed and accuracy in object detection tasks.
[0004] Obviously, existing cervical cancer cell detection methods have problems such as the need to pay more attention to fine-grained features and the lack of an end-to-end detection model. Summary of the invention
[0005] In order to solve the problems existing in the prior art cervical cancer cell detection methods, such as the need to pay more attention to fine-grained features and the lack of an end-to-end detection model, the present invention proposes a cervical cancer cell detection method based on the Yolov5l model.
[0006] The cervical cancer cell detection method based on the Yolov5l model of the present invention uses the large Yolov5 model as the main architecture of the detection network, that is, uses Yolov5l as the main architecture of the detection network, including the network backbone Backbone, the network neck Neck, and the detection layer Head. It is characterized in that a Droped Multi-branch SwinTransformer (DMBST) module and a convolutional layer CBS are added after the sixth layer in the Backbone; where DMBST is the abbreviation of Droped Multi-branch SwinTransformer; multiple feature map fusions with a tensor size of (762, 20, 20) are added in the Neck; a detection head Detect with a tensor size of (762, 20, 20) is added in the Head; the processing mechanism of the Droped Multi-branch SwinTransformer (DMBST) module includes:
[0007] (1) Input a feature map with a tensor size of (c, h, w), and disassemble the channels into n equal parts as the input of each branch (c / n, h, w); where c is the number of image channels; h is the height of the image, in pixels; w is the width of the image, in pixels; n is the number of branches;
[0008] (2) Each branch uses the Swin Transformer of the sliding window self-attention mechanism to extract the global attention map;
[0009] (3) Use the LA network to compress the number of image channels; that is, first compress the input image to r channels with a convolution, and then compress it to one channel with another convolution, and finally pass through the sigmoid activation layer of the logistic regression function to form a single attention map FM i ; and there will be as many attention maps FM as there are branches i ; The LA network refers to a convolutional neural network with only 3 layers, including two convolutional layers and an activation function layer;
[0010] (4) Random Branch Dropout, that is, randomly zero out the attention map FM of a certain branch i as follows:
[0011]
[0012] In the formula, i is a random number between 0 and n, and n is the number of branches; that is, randomly select the feature map FM of a certain branch i , and then make this feature map matrix equal to a matrix of all zeros; where the widths of the feature map matrix and the all-zero matrix are both w, in pixels; the heights are both h, in pixels;
[0013] (5) Since each branch may have different attention regions, it is necessary to fuse these attention regions. Therefore, Channel Maximize is used to merge the attention regions of each channel into an attention map, that is, in all attention maps, the maximum value is found in the channel dimension. The formula for Channel Maximize is as follows:
[0014] FM merge(i,j) = Max(FM 0(i,j) , FM 1(i,j) , FM 2(i,j) ,..., FM n(i,j) ), i ∈ [0, h], j ∈ [0, w]
[0015] In the formula, FM merge(i,j) is the value of the i-th row and j-th column of the finally fused feature map matrix, which is equal to the maximum value of the values of the i-th row and j-th column on the feature map of each branch. The purpose of doing this is to find the regions that need to be noted in the model;
[0016] (6) Perform a dot product on the input feature map matrix, that is, directly multiply the numbers in the corresponding positions of the matrix, and finally output a feature map with a tensor size of (c, h, w), where c is the number of image channels, h is the height of the image in pixels; w is the width of the image in pixels.
[0017] Furthermore, adding a randomly zeroed multi-branch sliding window self-attention mechanism module DMBST and a convolutional layer CBS after the sixth layer in the Backbone includes stacking them in the way of setting one layer of DMBST and then one layer of CBS, then setting another layer of DMBST and then another layer of CBS, and the stacking times are at least two; and the tensor sizes of each stacking are different.
[0018] Furthermore, adding a randomly zeroed multi-branch sliding window self-attention mechanism module DMBST and a convolutional layer CBS after the sixth layer in the Backbone includes adding a DMBST with a tensor size of (762, 20, 20) after the sixth layer in the Backbone, and then setting a layer of CBS after it; then, adding a DMBST with a tensor size of (512, 40, 40) after the eighth layer, and then setting a layer of CBS after it.
[0019] Further, the Sliding Window Self-Attention Mechanism Swin Transformer refers to a computer vision detection model based on shifted windows developed by Microsoft Research. It includes two parts. The first part is: Layer Normalization - Window Self-Attention - Dropout of some neurons - Layer Normalization - Linear Layer; the second part is only different from the first part in that the window self-attention is changed to cross-window attention.
[0020] Further, adding multiple feature map fusions with a tensor size of (762, 20, 20) in the network Neck includes performing feature map fusions at the 14th, 15th, 16th, 29th, 30th, and 31st layers in the Neck, and adding the Deterministic Multibranch Sliding Window Self-Attention Mechanism Module (DMBST) at the 25th and 28th layers in the Neck respectively.
[0021] Further, adding the detection head Detect with a tensor size of (762, 20, 20) in the detection layer Head includes setting the detection head Detect with a tensor size of (762, 20, 20) in the detection layer Head, which together with three other Detect heads with different tensors forms the Head; and using the output of the 30th layer in the network Neck as the input to the Detect with a tensor size of (762, 20, 20).
[0022] The beneficial technical effect of the cervical cancer cell detection method based on the Yolov5l model of the present invention is that it can effectively extract the minute detail features in the image, thereby effectively improving the detection effect of small targets in the image, and further effectively improving the detection accuracy and reliability of cervical cancer cells. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Att Figure 1 is the main architecture diagram of the detection network of the cervical cancer cell detection method based on the Yolov5l model of the present invention;
[0024] Att Figure 2 is the network structure diagram of the Deterministic Multibranch Sliding Window Self-Attention Mechanism Module (DMBST) of the present invention;
[0025] Att Figure 3 is the network structure diagram of the Sliding Window Self-Attention Mechanism Swin Transformer;
[0026] Att Figure 4 is the LA network structure diagram;
[0027] Att Figure 5 is the schematic diagram of deterministic branch zeroing;
[0028] Att Figure 6It is a schematic diagram of channel maximization;
[0029] Appendix Figure 7 It is the C3STR network structure.
[0030] The following further describes the cervical cancer cell detection method based on the Yolov5l model of the present invention in conjunction with the attached drawings and specific embodiments. Specific Embodiments
[0031] Appendix Figure 1 It is the main architecture diagram of the detection network of the cervical cancer cell detection method based on the Yolov5l model of the present invention. Appendix Figure 2 It is the network structure diagram of the random zeroing multi-branch sliding window self-attention mechanism module DMBST module of the present invention. Appendix Figure 3 It is the network structure diagram of the sliding window self-attention mechanism Swin Transformer. Appendix Figure 4 It is the LA network structure diagram. Appendix Figure 5 It is the schematic diagram of random branch zeroing. Appendix Figure 6 It is the schematic diagram of channel maximization. Appendix Figure 7 It is the C3STR network structure. Text description in the figure:
[0032] Input: Input;
[0033] Backbone: Network backbone;
[0034] Neck: Neck network;
[0035] Head: Detection layer;
[0036] CBS: Convolutional layer, including convolution + batch normalization + silu activation;
[0037] C3: Convolutional network with 3 layers;
[0038] C3STR: Convolutional network with 3 layers of sliding window self-attention mechanism Swin Transformer (see Appendix Figure 7 );
[0039] DMBST: Random zeroing multi-branch sliding window self-attention mechanism;
[0040] SPPF: Feature pyramid network;
[0041] Upsample: Upsampling;
[0042] Concat: Merge;
[0043] Detect: Detection head;
[0044] C: Number of channels;
[0045] H: Image height, unit: pixel;
[0046] W: Image width, unit: pixel;
[0047] n: Number of branches;
[0048] LA_Net: LA network;
[0049] Random Branch Dropout: Random branch zeroing;
[0050] Channel Maximize: Channel maximization;
[0051] Output: Output;
[0052] LN: Layer normalization;
[0053] W-MSA: Window-based multi-head self-attention;
[0054] SW-MSA: Self-attention between sliding windows;
[0055] MLP: Multi-layer perceptron;
[0056] Conv1: Convolution layer 1;
[0057] Conv2: Convolution layer 2;
[0058] activte: Activation layer.
[0059] As can be seen from the figure, the cervical cancer cell detection method of the present invention is based on the Yolov5l model, with the large Yolov5 model as the main architecture of the detection network, that is, Yolov5l as the main architecture of the detection network (see attached Figure 1 ), including the network backbone Backbone, the network neck Neck and the detection layer Head. It is characterized in that a random zeroing multi-branch sliding window self-attention mechanism module DMBST and a convolution layer CBS are added after the sixth layer in the Backbone; where DMBST is the abbreviation of Droped Multi-branch Swin Transformer; multiple feature map fusions with a tensor size of (762, 20, 20) are added in the Neck; a detection head Detect with a tensor size of (762, 20, 20) is added in the Head; the processing mechanism of the random zeroing multi-branch sliding window self-attention mechanism module DMBST includes (see attached Figure 2 ):
[0060] (1) Input a feature map with a tensor size of (c, h, w), and split the channels into n equal parts as the input for each branch (c / n, h, w); where c is the number of image channels; h is the height of the image in pixels; w is the width of the image in pixels; n is the number of branches; the purpose of doing this is not only to reduce the computational amount, but also to make different branches of the model generate different feature maps during training;
[0061] (2) Each branch uses the sliding window self-attention mechanism Swin Transformer to extract the global attention map;
[0062] (3) Use the LA network to compress the number of image channels; that is, first compress the input image to r channels with a convolution, and then compress it to one channel with another convolution, and finally pass through the sigmoid activation layer of the logistic regression function to form a single attention map FM i ; moreover, there will be as many attention maps FM as there are branches i ; the LA network refers to a convolutional neural network with only 3 layers, including two convolutional layers and an activation function layer (see Appendix Figure 4 );
[0063] (4) Random Branch Dropout, that is, randomly zero out the attention map FM of a certain branch i (see Appendix Figure 5 ), and the formula is as follows:
[0064]
[0065] In the formula, i is a random number between 0 and n, and n is the number of branches; that is, randomly select the feature map FM of a certain branch i , and then make this feature map matrix equal to a matrix of all zeros; where the widths of the feature map matrix and the all-zero matrix are both w in pixels; the heights are both h in pixels; the purpose of doing this is to drive the network to automatically learn new features that are beneficial to improving image classification and improve the model robustness;
[0066] (5) Since each branch may have different attention regions, it is necessary to fuse these several attention regions. Therefore, Channel Maximize is used to merge the attention regions of each channel into a single attention map (see Appendix Figure 6 ), that is, in all the attention maps, find the maximum value in the dimension of the channel, and the formula for channel maximization is as follows:
[0067] FM merge(i,j) = Max(FM 0(i,j) , FM 1(i,j) , FM2(i,j) ,..., FM n(i,j) ), i ∈ [0, h], j ∈ [0, w]
[0068] In the formula, FM merge(i,j) is the value of the i-th row and j-th column of the finally fused feature map matrix, which is equal to the maximum value of the values of the i-th row and j-th column on the feature maps of each branch; the purpose of doing this is to find the areas that need to be noted in the model;
[0069] (6) Perform a dot product on the input feature map matrix, that is, directly multiply the numbers at the corresponding positions of the matrices, and finally output a feature map with a tensor size of (c, h, w), where c is the number of image channels, h is the height of the image in pixels; w is the width of the image in pixels.
[0070] As mentioned above, since the differences in the morphological and color characteristics of cervical cells are much smaller than the differences between general detected objects and cannot be well distinguished, more attention needs to be paid to fine-grained features to distinguish cervical cells. Obviously, if the depth of the network is not deep enough, that is, the number of layers of the network is not enough, it may affect the extraction of some tiny detail features, thus affecting the detection effect of small targets in the image, that is, affecting the detection effect of cervical cancer cells. The cervical cancer cell detection method based on the Yolov5l model of the present invention adds a DMBST module and a convolutional layer CBS after the 6th layer in the network backbone to improve the feature extraction effect of subsequent layers, and then adds a feature map fusion with a tensor size of (762, 20, 20) in the network neck, and a detection head with a tensor size of (762, 20, 20) is newly added in the detection layer. Thus, on the basis of giving full play to the ability of the Yolov5l model to quickly and accurately detect multiple objects in the image, the detection effect of small targets in the image is effectively improved, and further the detection accuracy and reliability of cervical cancer cells are effectively improved. It should be particularly noted that the tensor size of (762, 20, 20) is selected for the purpose of complementing the existing tensor sizes in the Yolov5l model to improve the detection effect of the detection network model on small and medium targets.
[0071] As one of the preferred solutions, adding a randomly zeroed multi-branch sliding window self-attention mechanism module DMBST and a convolutional layer CBS after the sixth layer in the Backbone, including stacking them in the way of setting one layer of DMBST, then one layer of CBS, then another layer of DMBST, and then another layer of CBS, and the stacking times are at least two; and the tensor sizes of each stacking are different. Increasing the stacking times can increase the network depth, which in turn affects the extraction of some subtle detail features. Thus, the image can be detected from more dimensions to improve the detection effect of the detection network model on small and medium-sized targets. As one of the embodiments, add a DMBST with a tensor size of (762, 20, 20) after the sixth layer in the Backbone, and then set a layer of CBS after it; then, add a DMBST with a tensor size of (512, 40, 40) after the eighth layer, and then set a layer of CBS after it.
[0072] As one of the preferred solutions, the sliding window self-attention mechanism Swin Transformer refers to a computer vision detection model based on shifted windows developed by Microsoft Research (see Appendix Figure 3 ), which includes two parts. The first part is: layer normalization - self-attention within the window - discarding a part of neurons - layer normalization - linear layer; the second part only differs from the first part in that the self-attention within the window is changed to the attention between sliding windows. Using the sliding window self-attention mechanism Swin Transformer, a computer vision detection model based on shifted windows developed by Microsoft Research, to extract the global attention map of the image for each branch can pay a great deal of attention to the subtle differences in the image and improve the accuracy of detection.
[0073] As one of the preferred solutions, adding multiple feature map fusions with a tensor size of (762, 20, 20) in the network Neck, including performing feature map fusions at the fourteenth, fifteenth, sixteenth, twenty-ninth, thirtieth, and thirty-first layers in the Neck, and adding randomly zeroed multi-branch sliding window self-attention mechanism modules DMBST at the twenty-fifth and twenty-eighth layers in the Neck respectively. Adding feature map fusions with a tensor size of (762, 20, 20) can better focus on each detail in the image. In order to perform better feature fusion on the feature maps stitched together in the Neck, two DMBST modules are specially added.
[0074] As one of the preferred solutions, adding a detection head Detect with a tensor size of (762, 20, 20) to the detection layer Head includes setting a detection head Detect with a tensor size of (762, 20, 20) in the detection layer Head, which together with three other Detect with different tensors constitutes Head; and using the output of the thirtieth layer in the network neck Neck as the input of Detect with a tensor size of (762, 20, 20).
[0075] Experimental detection results:
[0076] Table 1: Comparison of detection results
[0077]
[0078] *Data is from the following research: Liang Y, Tang Z, Yan M, et al. Comparison-Based Convolutional Neural Networks for Cervical Cell / Clumps Detection in the Limited Data Scenario:, 2018.
[0079] Table 2: Comparison of detection results of three models on the public dataset
[0080]
[0081] Table 3: Comparison of detection results of three models on the local dataset
[0082]
[0083] In order to compare the accuracy and reliability of the cervical cancer cell detection method based on the Yolov5l model of the present invention from different levels and in different ways. The experiment compares three models, namely the Yolov5l model, Yolo5l + C3STR, and YoGAnet, and uses two datasets, namely the public dataset and the local dataset, to train and detect different models. The descriptions of relevant models and index data are as follows:
[0084] Yolo5l: refers to the unmodified Yolov5l detection model obtained from https: / / github.com / ultralytics / yolov5;
[0085] YoGAnt: refers to the detection model obtained by modifying Yolov5l according to the method of the present invention;
[0086] Yolo5l+C3STR: Refers to the detection model that replaces all DMBST modules in YoGAnet with C3STR modules;
[0087] Public dataset: Derived from the academic research Liang Y, Tang Z, Yan M, et al. Comparison-Based Convolutional Neural Networks for Cervical Cell / Clumps Detection in the Limited Data Scenario:, 2018; The public dataset contains 7410 cervical microscopic images and 50687 object instance bounding boxes;
[0088] Local dataset: 12 cases from the Department of Pathology of the Army Medical Center of the Army Medical University, with a total of 1514 images and 4698 instances;
[0089] mAP@50: Refers to the detection result when the mAP value is at a confidence threshold of 0.5; mAP is an evaluation metric for the average accuracy of different classes of targets under different ratios of the intersection to the union of the ground truth boxes and the detection boxes; Generally, the larger the value of mAP@50, the more accurate the detection result;
[0090] mAP@50:95: Refers to the detection result when the mAP value is at a confidence threshold from 0.5 to 0.95;
[0091] Recall: The recall rate, which refers to the proportion of the number of samples correctly detected among all actual positive examples; Generally, the larger the Recall value, the more accurate the detection result;
[0092] Precision: The precision rate, which refers to the proportion of the number of samples that are actually positive among all samples detected as positive; Generally, the larger the Precision value, the more accurate the detection result;
[0093] F1: Refers to the harmonic mean defined as the precision rate and the recall rate, an evaluation metric for classification models;
[0094] ascus: Intraepithelial lesion cells;
[0095] asch: Cells with high-grade squamous intraepithelial lesion not excluded;
[0096] lsil: Low-grade squamous intraepithelial lesion cells;
[0097] hsil: High-grade squamous intraepithelial lesion cells;
[0098] scc: Squamous cancer cells;
[0099] AGC: Atypical glandular cells;
[0100] Trich: Trichomonas;
[0101] Candi: Candida;
[0102] Flora: Bacterial flora;
[0103] Herp: Herpes;
[0104] Actin: Actinomyces;
[0105] As can be seen from the data listed in Table 1, Table 2 and Table 3, from the two indicators of mAP@50 and Recall (Table 1), the YoGAnt model of the present invention has improved compared with the Yolov5l model, with increases of 6.2% and 3.3% respectively. Training and detection of each model were carried out using a comparison dataset, and the results showed (Table 2) that the YoGAnt model of the present invention has obvious advantages in detecting cells such as ascus, asch, lsil, hsil, scc and candida. The local dataset only has 5 types of squamous epithelial lesion cells such as ascus, asch, lsil, hsil and scc, and the data volume is relatively small. Training and detection of each model were carried out using the local dataset, and the results showed (Table 3) that the YoGAnt model of the present invention has obvious improvements compared with the Yolov5l model in terms of indicators such as mAP@50, Recall and F1, which are 3.5%, 20% and 4.7% respectively.
[0106] In summary, the beneficial technical effect of the cervical cancer cell detection method based on the Yolov5l model of the present invention is that it can effectively extract minute detail features in the image, thereby effectively improving the detection effect of small targets in the image, and further effectively improving the detection accuracy and reliability of cervical cancer cells.
Claims
1. A cervical cancer cell detection method based on the Yolov5l model, which uses the large Yolov5 model as the main architecture of the detection network, that is, uses Yolov5l as the main architecture of the detection network, including, A network backbone, a network neck, and a detection head. It is characterized in that a Droped Multi-branch Swin Transformer (DMBST) module and a convolutional layer CBS are added after the sixth layer in the backbone; where DMBST is the abbreviation of Droped Multi-branch Swin Transformer; multiple feature map fusions with a tensor size of (762, 20, 20) are added in the neck; a detection head Detect with a tensor size of (762, 20, 20) is added in the head; the processing mechanism of the Droped Multi-branch Swin Transformer (DMBST) module includes: (1) Input a feature map with a tensor size of (c, h, w), and split the channels into n equal parts as the input of each branch (c / n, h, w); where c is the number of image channels; h is the image height in pixels; w is the image width in pixels; n is the number of branches; (2) Each branch uses the Swin Transformer of the sliding window self-attention mechanism to extract the global attention map; (3) Compress the number of image channels using the LA network; that is, first compress the input image to r channels using convolution, then compress it to one channel again using convolution, and finally form a single attention map FM through the sigmoid activation layer of the logistic regression function i ; and there will be as many attention maps FM as there are branches i ; the LA network refers to a convolutional neural network with only 3 layers, including two convolutional layers and one activation function layer; (4) Random Branch Dropout, that is, randomly zero out the attention map FM of a certain branch i The formula is as follows: Wherein, i is a random number between 0 and n, and n is the number of branches; that is, randomly select the feature map FM of a certain branch i , and then make this feature map matrix equal to a matrix of all zeros; wherein, the widths of the feature map matrix and the all-zero matrix are both w, with the unit of pixel; the heights are both h, with the unit of pixel; (5) Since each branch may have different attention regions, it is necessary to fuse these several attention regions. Therefore, Channel Maximize is used to merge the attention regions of each channel into an attention map, that is, in all attention maps, find the maximum value in the channel dimension. The Channel Maximize formula is as follows: FM merge(i,j) = Max(FM 0(i,j) , FM 1(i,j) , FM 2(i,j) ,..., FM n(i,j) ), i ∈ [0, h], j ∈ [0, w] where, FM merge(i,j) is the value of the i-th row and j-th column of the finally fused feature map matrix, which is equal to the maximum value of the values of the i-th row and j-th column on the feature maps of each branch; the purpose of doing this is to find the areas that need to be noted in the model; (6) Perform a dot product on the input feature map matrix, that is, directly multiply the numbers at the corresponding positions of the matrix, and finally output a feature map with a tensor size of (c, h, w), where c is the number of image channels, h is the image height in pixels; w is the image width in pixels.
2. The cervical cancer cell detection method based on the Yolov5l model according to claim 1, wherein The addition of the Droped Multi-branch Swin Transformer (DMBST) module and the convolutional layer CBS after the sixth layer in the backbone includes stacking them in the way of setting one layer of DMBST and then one layer of CBS, and then setting one layer of DMBST and then one layer of CBS again, and the stacking times are at least two times; And the tensor sizes of each stack are different.
3. The cervical cancer cell detection method based on the Yolov5l model according to claim 2, wherein, The addition of the Droped Multi-branch Swin Transformer (DMBST) module and the convolutional layer CBS after the sixth layer in the backbone includes adding a DMBST with a tensor size of (762, 20, 20) in the sixth layer of the backbone, and then setting one layer of CBS behind it; then, adding a DMBST with a tensor size of (512, 40, 40) in the eighth layer, and then setting one layer of CBS behind it.
4. The cervical cancer cell detection method based on the Yolov5l model according to claim 1, wherein, The sliding window self-attention mechanism Swin Transformer refers to a computer vision detection model based on shifted windows developed by Microsoft Research. It consists of two parts. The first part is: layer normalization - self-attention within the window - dropping a part of neurons - layer normalization - linear layer; the second part is only different from the first part in that the self-attention within the window is changed to the attention between sliding windows.
5. The cervical cancer cell detection method based on the Yolov5l model according to claim 1, wherein, The addition of multiple feature map fusions with a tensor size of (762, 20, 20) in the network neck Neck includes performing feature map fusions at the fourteenth, fifteenth, sixteenth, twenty-ninth, thirtieth, and thirty-first layers in Neck, and adding a stochastic depth multi-branch sliding window self-attention mechanism module DMBST at the twenty-fifth and twenty-eighth layers in Neck respectively.
6. The cervical cancer cell detection method based on the Yolov5l model according to claim 1, characterized in that The addition of a detection head Detect with a tensor size of (762, 20, 20) in the detection layer Head includes setting a detection head Detect with a tensor size of (762, 20, 20) in the detection layer Head, which together with three other Detect with different tensors constitutes Head; and using the output of the thirtieth layer in the network neck Neck as the input of the Detect with a tensor size of (762, 20, 20).