Lightweight method for colorectal polyp detection based on improved YOLOv8 model
By improving the YOLOv8 model, replacing backbone with FastersNet, replacing the detection head with LThead and using DwConv, the problem of high deployment requirements and slow detection rate of colorectal polyp detection model is solved, and lightweight and efficient detection is achieved.
Patent Information
- Application Number
- CN202510337569.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-04
AI Technical Summary
The deployment of existing colorectal polyp detection models has high requirements for equipment, and it is difficult to break through the detection rate limit, and doctors are prone to missed diagnosis and misdiagnosis in colonoscopy.
Using the improved YOLOv8 model, the backbone is replaced with FastersNet, the detection head is replaced with LThead, and the convolution module of LThead is replaced with DwConv, reducing the calculation amount and parameter amount, and improving the lightweighting degree of the model and detection rate.
The colorectal polyp detection model is lightweight, which improves detection efficiency and speed, reduces the hardware requirements, and reduces the possibility of missed diagnosis and misdiagnosis.
Smart Images

Figure CN120259240A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of medical image recognition technology and deep learning technology, and specifically relates to a lightweight method for colorectal polyp detection based on an improved YOLOv8 model. Background Art
[0002] The prevalence and mortality of colorectal cancer rank third and second respectively worldwide. Reports show that in 2022, there were 517,100 new cases of colorectal cancer in China, accounting for 10.7% of all malignant tumor incidences. There were 240,000 deaths from colorectal cancer, accounting for 9.3% of all malignant tumor deaths. The incidence and mortality of colorectal cancer in the whole country were 36.63 / 100,000 and 17.00 / 100,000 respectively, showing an overall upward trend.
[0003] The best treatment plan for colorectal cancer is to perform surgical resection when colorectal cancer is in the form of early polyps. Colonoscopy is the most commonly used screening method for colorectal cancer. Doctors detect polyps by observing the surface of the colorectal wall. However, when doctors detect polyps during colonoscopy, there are problems such as the small size and large number of polyps, and the texture and other information of some polyps are similar to those of the surrounding tissues, making it difficult to directly identify them with the naked eye. This leads to misdiagnosis and missed diagnosis easily when doctors have insufficient experience or are visually fatigued. Some studies have shown that about 26.3% of patients have missed diagnosis after reexamination.
[0004] In addition, not all polyps need to be surgically resected. Hyperplastic polyps basically have no probability of canceration and do not cause pain to patients. Doctors usually do not recommend surgical resection for such polyps. However, adenomatous polyps have a relatively high incidence, especially villous polyps or villotubular polyps with a diameter greater than a certain size and a villous quantity greater than 25%, which have a relatively high risk of canceration. In this case, doctors will strongly recommend that patients undergo surgical resection.
[0005] Currently, the gold standard for examining polyps is to extract the tissue of the polyp, obtain a pathological report through pathological examination of the extracted polyp tissue, and then judge the type of the polyp. Some experienced doctors can also judge the type of the polyp by observing the characteristics of the polyp.
[0006] Facing the urgent need to improve the accuracy and detection efficiency in colonoscopy, the technology of object detection models based on deep learning has become crucial. We can use such technology to overcome the problem of missed diagnosis that easily occurs when doctors are fatigued. With the continuous breakthroughs of deep learning in the field of computer medical images, more and more researchers have used this technology to strengthen the ability of early disease detection and treatment, and have achieved good results in the judgment and positioning ability of polyps. However, there are still a series of problems in the polyp detection task based on deep learning, among which the deployment of the detection model has high requirements for equipment, and it is difficult to break through the limitation of the model detection rate.
[0007] Therefore, in view of the lightweight problem of the colorectal polyp detection model, this application proposes a colorectal polyp detection and classification method based on the YOLOv8 model. This method aims to improve the detection efficiency of the model that assists doctors in colorectal detection and build a more lightweight architecture for the model, so as to reduce the requirements of the model for hardware and meet the requirement of improving the model detection rate. Summary of the Invention
[0008] To achieve the above objectives, the present invention provides a lightweight method for colorectal polyp detection based on an improved YOLOv8 model, including the following steps:
[0009] S1. Obtain data from the colorectal polyp dataset locally;
[0010] S2. Perform data preprocessing and data augmentation operations on the data in the colorectal polyp dataset;
[0011] S3. Improve the YOLOv8 model by replacing the backbone of YOLOv8 with the FastersNet architecture to improve the model's utilization efficiency of channel information;
[0012] S4. Replace the detection head part of the model with the newly developed LThead to reduce the computational amount during model inference and improve the running efficiency of the model.
[0013] S5. Replace the Conv of the newly developed LThead with DwConv to reduce the computational amount and the number of parameters for extracting channel information during model convolution, and achieve the lightweight of the model.
[0014] S6. Input the dataset data obtained in step S2 into the improved YOLOv8 model obtained in step S6, adjust the parameters and perform training;
[0015] S7. Observe various indicators of the model during training to determine whether the model converges. If the model has converged, save the converged lightweight colorectal polyp detection model, use the lightweight colorectal detection model to detect and classify polyps, and calculate the inference time, number of parameters, and computational complexity of the model.
[0016] Preferably, the obtained colorectal polyp dataset contains 20,000 colorectal polyp images, and each image has a txt format label file with the same name. The label file records the category and location information of the polyp. The first digit being "0" indicates a hyperplastic polyp, and the first digit being "1" indicates an adenomatous polyp. The following four digits respectively represent the relative position of the colorectal polyp in the picture.
[0017] Preferably, the operations of data preprocessing and data augmentation on the colorectal polyp dataset data include:
[0018] S21. According to the storage method of the YOLO dataset, successively establish the images and labels folders to store colorectal polyp images and annotation information files, and divide the colorectal polyps into a training set, a validation set, and a test set in a ratio of 8:1:1.
[0019] S22. Perform operations such as randomly rotating, flipping, and adding various noises to the colorectal polyp images to achieve data augmentation. The formula for the rotation operation is as follows:
[0020]
[0021] In the formula: x, y are the coordinates before rotation, m, n are the coordinates of the rotation center, α is the rotation angle, and left, top are the coordinates of the upper left corner after rotation.
[0022] The formula for the image flipping operation is as follows:
[0023]
[0024] In the formula, f w represents the width of the image, x, y are the coordinates before flipping, and x0, y0 are the coordinates after flipping.
[0025] The noises include but are not limited to Gaussian noise, Poisson noise, and salt-and-pepper noise. The formula for Gaussian noise is as follows:
[0026]
[0027] Preferably, the improved YOLOv8 model includes: replacing the backbone of YOLOv8 with the FastersNet network. FastersNet uses the form of partial convolution to reduce the calculation of duplicate features, thereby improving the model's ability to extract channel information, reducing the computational load during model inference, and increasing the speed of model inference. The FastersNet includes:
[0028] S31. Partial Convolution module. The design idea of partial convolution comes from the fact that there is a lot of duplicate information in the feature channels obtained after image feature extraction. If convolutional operations are used to extract information from all channels, there will be a lot of useless calculations, that is, similar effects can also be obtained by performing feature extraction operations on partial feature maps. Thus, for a feature map F ∈ R H×W×C split F into and two parts. Among them, perform convolutional operations to obtain the feature information of F1.
[0029] S32. Small patch convolution module. It is calculated through the Frobenius norm. The important information of the picture is generally concentrated in the central part of the picture. For the remaining information, there is a lot of information highly similar to the channel information of F1. To simplify the inference speed and not lose the ability to extract information from the remaining channels, the central information of the F2 feature map is extracted in the way of small patch convolution, and finally the obtained information is concatenated with the previously obtained F1 feature map to form a new feature map.
[0030] Preferably, replace the detection head of YOLOv8 with LThead. The design idea of LThead is to combine the advantages of the decoupled head and the coupled head, and reduce the convolutional operations before loss to achieve a lightweight effect. The decoupled head divides the tasks of the detection head into two parts. The first part is to identify the position of the object and return the position information of the detected item. The second part of the task is to classify the objects identified in the first part. Since the decoupled head assigns different tasks to different convolutions, its performance in performing detection tasks is better than that of the coupled head. Such a design architecture also has certain defects. Compared with the coupled head, the design architecture of the decoupled head has almost twice as many convolutional modules. And in order to extract information of different scales, the number of detection heads generally ranges from three to five in object detection tasks. For this reason, the architecture of the decoupled head will increase the computational consumption of the model, so it lacks in terms of lightweight. The design concept of the coupled head is exactly the opposite of that of the decoupled head. The architecture of the coupled head combines the positioning and classification tasks in the same convolutional layer. The advantage of using the same convolutional module to perform both tasks is that it greatly reduces the computational resource consumption of the model. However, relatively speaking, such an architecture has a decline in positioning ability and classification ability. The design idea of LThead is to use one convolutional module in the coupled head to share the positioning task, that is, on the premise of keeping the same number of convolutional modules as the coupled head, the positioning task and the classification task are respectively handed over to two convolutional modules for processing, thereby achieving the effect of lightweight or improving the detection ability.
[0031] Based on LThead, we replace the convolutional modules of the backbone with DwConv to reduce the inference computational amount. DwConv replaces the convolutional process with two processes: Depthwise and Pointwise. In the Dwpthwise operation, the operation of using C out filters to extract feature maps in ordinary convolution is abandoned. First, each channel uses one filter to extract its respective features once, obtaining the number of input feature channels of feature maps. In this way, the feature information carried by each feature can also be effectively extracted. In the Pointwise operation, the feature maps of the Depthwise operation are extracted with C out 1×1×C in filters to extract the feature information of the feature maps. The total computational amount of DwConv convolution is as follows:
[0032] FLOPs = HWC in (K 2 +C out )
[0033] The computational amount required for ordinary convolution is as follows:
[0034] FLOPs = HWCin C out K 2
[0035] The K shown in the formula refers to the size of the filter.
[0036] Preferably, the data content obtained in step S2 is input into the improved YOLOv8 model in S6, the model parameters are modified and training is started. It is characterized in that the parameter of batch size in the model is 32, the initial learning rate is 0.0001, the optimization function is selected as SGD, the number of training rounds is 300 rounds, and the image size is 640.
[0037] A lightweight method for colorectal polyp detection based on an improved YOLOv8 model according to claim 1, characterized in that in step S8, it is observed whether the various indicators in the model training process converge, and the evaluation indicators are accuracy (Precision), recall (Recall), average detection precision (mAP@0.5) under the intersection over union threshold of 0.5, average detection precision (mAP@0.5:0.95) under the intersection over union threshold from 0.5 to 0.95, floating-point operations per second in billions (GFLOPs), and number of parameters (Paramas).
[0038] The beneficial effects of the present invention are as follows: The present invention proposes a lightweight method for colorectal polyp detection based on an improved YOLOv8 model, improves the YOLOv8 model, replaces the backbone of YOLOv8 with the FastersNet network to improve the extraction efficiency of the model for partial channel information, replaces the detection head with LThead to achieve lightweight design, and replaces the convolutional module in LThead with DwConv to further improve the inference speed of the model, making real-time detection possible. Description of the Drawings
[0039] Figure 1 It is a flowchart of the method of the present invention;
[0040] Figure 2 It is the overall network architecture diagram of the present invention;
[0041] Figure 3 It is the network architecture of FasterNet in the present invention;
[0042] Figure 4 It is the structural diagram of the LThead module in the present invention;
[0043] Figure 5 It is the schematic diagram of DwConv in the present invention;
[0044] Figure 6 It is the structural diagram of the Dw-LThead module in the present invention;
[0045] Figure 7 、 Figure 8 This is the detection and classification effect diagram of colorectal polyps of the present invention; Specific implementation method
[0046] A lightweight method for colorectal polyp detection based on an improved YOLOv8 module according to the present invention is specifically implemented according to the following steps:
[0047] Step 1: Obtain a colorectal polyp dataset, including 20,000 colorectal polyp image data and their corresponding txt annotation files.
[0048] Step 2: Perform data preprocessing and augmentation operations on the colorectal polyp dataset. Create images and labels folders to store colorectal polyp images and annotation information files respectively. And the colorectal polyp images are divided into training set, validation set and test set according to the ratio of 8:1:1. Random selection, flipping and adding noise operations are performed on the colorectal polyp images for data augmentation.
[0049] Step 3: Improve the YOLOv8 model by replacing the backbone of YOLOv8 with the FastersNet architecture to improve the model's utilization efficiency of channel information.
[0050] Step 4: Replace the detection head part of the model with the newly developed LThead architecture to reduce the computational amount during model inference and improve the running efficiency of the model.
[0051] Step 5: Replace the Conv of the newly developed LThead with DwConv to reduce the computational amount and number of parameters for channel information extraction during model convolution, and achieve model lightweighting.
[0052] Step 6: Input the dataset obtained in Step 2 into the improved YOLOv8 model obtained in Step 6, adjust the parameters and perform training.
[0053] Step 7: Input the preprocessed dataset into the improved YOLOv8 model, adjust the model parameters and perform training. The model training parameters in this example are: the batch size is 32, the initial learning rate is 0.0001. The optimization function is selected as SGD, the training rounds are 300 rounds, and the image size is 640.
[0054] Step 8: Observe various indicators during the model training process to judge whether the model converges. If the training converges, save the trained colorectal polyp detection and classification model, and use the trained colorectal polyp detection and classification model to detect and output the category and location information of colorectal polyps.
[0055] To verify the effects of introducing FastersNet, LThead, and DwConv into the model, two groups of experiments were conducted in this example. One group used the YOLOv8 model, and the other group used the YOLOv8 + LThead + DwConv + FastersNet model. The training set, validation set, and test set of all models used the same set of data. The specific data is shown in Table 1 below.
[0056] Table 1
[0057] By comparison, it can be seen that compared with the YOLOv8 model, the improved YOLOv8 model of the present invention has significantly improved the number of parameters, FPS, and GFLOPs, and has a slight improvement in precision. The improved model proposed by the present invention exhibits a series of significant advantages. Through data preprocessing and enhancement operations, the robustness and generalization ability of the model are improved; the designed LThead combines the advantages of decoupled heads and coupled heads, reducing the inference calculation amount and the number of parameters required by the model. Subsequently, a more lightweight Dw-LThead was designed in combination with DwConv to further improve the detection rate of the model.
[0058] The above examples only represent the implementation manners of the present invention and should not be construed as limitations on the present invention. Within the scope of the claims, any modifications, equivalent substitutions, and improvements that conform to the core spirit and principles of the present invention should be regarded as being included within the protection scope of the present invention.
Claims
1. A lightweight method for colorectal polyp detection based on an improved YOLOv8 model, characterized in that, It includes the following steps: S1. Obtain data from the local colorectal polyp dataset; S2. Perform data preprocessing and data augmentation operations on the data in the colorectal polyp dataset; S3. Improve the YOLOv8 model by replacing the backbone of YOLOv8 with the FastersNet architecture to improve the model's utilization efficiency of channel information; S4. Replace the detection head part of the model with LThead to reduce the computational amount during model inference and improve the model's running efficiency; S5. Replace the Conv module of LThead with DwConv to reduce the computational amount of channel information extraction and the number of parameters during model convolution, and achieve model lightweighting; S6. Input the dataset data obtained in step S2 into the improved YOLOv8 model obtained in step S6, adjust the parameters and perform training; S7. Observe various indicators of the model during training to determine whether the model converges. If the model has converged, save the converged lightweight colorectal polyp detection model, use the lightweight colorectal detection model to detect and classify polyps, and calculate the inference duration, number of parameters, and computational amount of the model.
2. A lightweight method for colorectal polyp detection based on improved YOLOv8 according to claim 1, characterized in that, In step S1, the obtained colorectal polyp dataset contains 20,000 pieces of colorectal polyp data. Among them, each piece of data has a txt file with the same name as the label file for this piece of data. This label file contains the type and location information of the polyp. "0" is a hyperplastic polyp, and "1" is an adenomatous polyp. The four numbers after the category data represent the relative position of the polyp in the data.
3. A lightweight method for colorectal polyp detection based on an improved YOLOv8 model according to claim 1, characterized in that, In step S2, the data preprocessing and data augmentation operations on the data in the colorectal polyp dataset include: S21. According to the storage method of the YOLO dataset, successively create the images and labels folders to store the colorectal polyp image data and the corresponding annotation information files, and divide the colorectal polyp data into a training set, a validation set, and a test set according to the ratio of 8:1:1; S22. Perform operations such as randomly rotating, flipping, and adding various noises to the colorectal polyp image data to achieve data augmentation.
4. A lightweight method for colorectal polyp detection based on an improved YOLOv8 model according to claim 1, characterized in that In step S3, replace the backbone of YOLOv8 with FastersNet to improve the model's ability to extract channel information, so that the computational amount required for the model during inference is significantly reduced, effectively improving the model inference speed. The FastersNet includes: S31. Partial Convolution module. The partial convolution module slices the input feature map F ∈ R H×W×C into and two parts, performs normal convolution operations on one part to extract information from the channels of this part and reduce the extra operations on the redundant information of the remaining channels. In addition, the other part after slicing is retained for feature extraction in subsequent work. S32. The small block convolution module. For the remaining feature channel part after partial convolution cutting, use small block convolution to only extract the channel information in the center part of the remaining channels, obtain the information with a relatively large proportion of importance in this channel, and after getting the result, splice it with the result obtained by partial convolution to form a new feature map.
5. A lightweight method for colorectal polyp detection based on an improved YOLOv8 model according to claim 1, characterized in that, Step S4, replace the detection head with LThead. LThead combines the advantages of the decoupling head and the coupling head. One of the convolutional modules of the coupling head is used to share the tasks of localization or classification, improving the integration ability of the detection head for the existing feature information. Compared with the decoupling head architecture, LThead greatly reduces the number of model parameters. Due to the significant reduction in the required number of parameters, the computational amount required in subsequent inference tasks also decreases significantly, and the number of io operations required for memory also decreases, greatly improving the inference speed of the model.
6. A colorectal polyp detection and classification method based on an improved YOLOv8 model according to claim 1, characterized in that In step S5, replace the Conv of LThead replaced in step S4 with DWConv. DWConv simplifies the traditional convolution operation, greatly reducing the computational amount of convolution operations and also reducing the number of parameters required to save weights, achieving a lightweight effect. The DWConv includes: S51. Depthwise operation abandons the operation of using C out filters to extract the features of the feature map in ordinary convolution. Instead, it first uses one filter for each channel to extract their respective features once, obtaining the number of feature maps equal to the number of input feature channels. This can also effectively extract the feature information carried by each feature. S52. Pointwise operation, using C feature maps extracted in step S51 respectively out filters of 1×1×C in to extract feature information from the feature maps. Since the computational cost required by 1×1×C in is much smaller than that of ordinary convolution operations, and the number of parameters required is also greatly reduced. Therefore, the lightweight effect is achieved.
7. A lightweight method for colorectal polyp detection based on an improved YOLOv8 model according to claim 1, characterized in that, In step S6, input the dataset data obtained in step S2 into the improved YOLOv8 model obtained in step S6, adjust the model parameters and perform training. It is characterized in that the model parameter batch size is 32, the initial learning rate is 0.0001, the optimization function is SGD, the number of training rounds is 300 rounds, and the image size is 640.
8. A lightweight method for colorectal polyp detection based on an improved YOLOv8 model according to claim 1, characterized in that, In step S8, during the model training process, observe whether the various indicators of the model converge. The evaluation indicators include but are not limited to accuracy (Precision), recall (Recall), mean average precision at an intersection over union threshold of 0.5 (mAP@0.5), mean average precision at an intersection over union threshold from 0.5 to 0.95 (mAP@0.5:0.95), floating point operations per second in billions (GFLOPs), and number of parameters (Paramas).