Colorectoscope polyp intelligent classification and detection method based on data enhancement and improved YOLO
By combining the SCGAN model to generate highly realistic polyp-free images with the ST-YOLO model, the problems of data imbalance and insufficient feature extraction in colorectal endoscopy polyp detection are solved, achieving high-precision polyp detection and classification, and adapting to the detection needs in complex backgrounds.
Patent Information
- Application Number
- CN202511731009.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-24
- Publication Date
- 2026-03-20
- Estimated Expiration
- 2045-11-24
AI Technical Summary
Existing colorectal endoscopy methods for polyp detection suffer from high false negative and false positive rates when faced with imbalanced datasets and complex backgrounds. In particular, feature extraction capabilities are limited when there are many intestinal folds, complex backgrounds, small polyp sizes, or when polyps are obscured.
We employ a data augmentation and improved YOLO approach, generating realistic polyp-free images using the SCGAN model. By combining the Swin Transformer with the ST-YOLO model of YOLO, we achieve intelligent detection and classification of polyps. We improve the CycleGAN model using structural similarity loss to generate a balanced dataset, and embed the Swin Transformer module into the YOLOv8 model. We design a multi-task synchronization strategy for detection and classification.
It improves the detection accuracy and recall rate of small polyps, reduces the false detection rate, and enhances the accuracy and efficiency of the detection process, adapting to the detection needs of complex intestinal backgrounds.
Smart Images

Figure CN121707930A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision and medical image processing technology, and relates to a method for classifying and detecting colorectal polyps. Specifically, it relates to a method for achieving high-precision intelligent detection of colorectal polyps in imbalanced datasets by sample augmentation and improving the target detection network. Background Technology
[0002] Colorectal cancer is one of the most common malignant tumors worldwide, ranking among the top in both incidence and mortality. Clinical studies have shown that the vast majority of colorectal cancers develop from early-stage benign adenomatous polyps, and late detection is a significant contributing factor. Therefore, timely and accurate detection and removal of precancerous polyps during colonoscopy has become an effective way to prevent colorectal cancer.
[0003] Traditional polyp detection methods heavily rely on the endoscopist's personal experience and mental state, making them prone to missed diagnoses and misdiagnoses when faced with complex intestinal structures, fecal obstruction, and the varied shapes and sizes of polyps. In recent years, deep learning-based object detection methods, including R-CNN, Fast R-CNN, Faster R-CNN, YOLO, and single-shot multi-boundary detection algorithms, have offered new approaches to addressing these challenges.
[0004] However, in clinical data collection, because most patients seek medical attention when they are ill, there is a significant imbalance in medical imaging data, with far more image samples containing polyps than healthy images without polyps. This imbalance can bias model training, severely affecting its generalization ability and diagnostic accuracy. In particular, existing object detection algorithms (such as the YOLO series) have limited feature extraction capabilities when dealing with scenarios involving numerous intestinal folds, complex backgrounds, small polyp sizes, or occlusions, easily losing key information and resulting in a high false negative rate for small polyps.
[0005] Therefore, developing an intelligent method that can effectively address data imbalance and accurately detect various polyps in complex scenarios has significant clinical application value. Summary of the Invention
[0006] The purpose of this invention is to provide a method for intelligent classification and detection of polyps in colorectal endoscopy based on data augmentation and improved YOLO. This method uses the SCGAN model to perform data augmentation on colorectal endoscopy images and uses the ST-YOLO model to design a strategy for simultaneous detection and classification tasks, ultimately achieving intelligent detection and classification of polyps. This solves the problem of low polyp detection accuracy caused by data imbalance and insufficient small target detection performance in the prior art.
[0007] The objective of this invention is achieved through the following technical solution:
[0008] A method for intelligent classification and detection of colorectal polyps based on data augmentation and improved YOLO includes the following steps:
[0009] Step 1: Preprocess and enhance the diversity of colorectal endoscopy images. The specific steps are as follows:
[0010] Step 101: Perform geometric transformation on the minority samples (i.e., healthy colonoscopy images without polyps) to simulate the image changes caused by different shooting angles in actual examinations.
[0011] Step 102: Convert the RGB image to the HSV color space, adjust the saturation and brightness of the image by adjusting the S and V channels, and adjust the image contrast by linear transformation to simulate the changes in light inside the intestine, further increasing the diversity of data and improving the realism of the healthy images generated by the subsequent model.
[0012] Step 2: Improve the training loss function of the Recurrent Generative Adversarial Network (CycleGAN) by introducing Structural Similarity (SSIM) loss. and will Add to the total loss function In the process, the improved SC-GAN model was used to generate polyp-free healthy colonoscopy images. The specific steps are as follows:
[0013] Step 201: Construct an image transformation framework based on CycleGAN, and use this framework model to obtain two metrics:
[0014]
[0015] in, To counteract the loss, it is used to train the adversarial relationship between the generator and the discriminator; Cyclic consistency loss is used to ensure that the image generated by the generator is close to the original image after being generated repeatedly, thus preventing the generation of unrealistic images. For discriminator Real images of a healthy colon and rectum The expected value of the discrimination probability. A realistic image showing a colon or rectum with polyps. , Represents a generator. , This represents the discriminator, where the L1 norm is used to measure the difference between the generated image and the original image;
[0016] Step 202: Improve the loss function of CycleGAN by introducing a structural similarity (SSIM) loss term in addition to the original adversarial loss and cycle consistency loss. Quantize the brightness of two images Contrast and structure Three dimensions of similarity are used to assess the differences between images, namely:
[0017]
[0018]
[0019] in, , For image and The local mean, , Standard deviation, and For variance, For covariance, , and It is the stability constant;
[0020] Will Mapping from [1, 0] to [0, 1]:
[0021]
[0022] in, Images of meat with polyps, Images of meat without breathing surface;
[0023] Will Add it to the total loss function, that is:
[0024]
[0025] in, The weights for the cycle consistency loss, The weights for the structural consistency loss;
[0026] The local windowing mechanism of SSIM is used to capture the structural features of local regions in colorectal endoscopy images;
[0027] Step 203: Using the preprocessed minority class samples and the original majority class samples from Step 1, train the SC-GAN model improved by introducing SSMI and CycleGAN. After training, input an image with polyps and the generator will generate an image without polyps.
[0028] Step 204: Combine subjective and objective evaluation indicators to screen images that pass the quality test, obtain the effective generation rate, and use it to judge the quality of the images generated in each group of experiments. Finally, select the polyp-free images and polyp-containing images generated through the above steps to form a balanced dataset.
[0029] Step 3: Embed the Swin Transformer module before the Spatial Pyramid Pooling (SPPF) module in the backbone layer of the YOLOv8 model to obtain the ST-YOLO model, and then train and fine-tune the parameters of the constructed ST-YOLO model. The specific steps are as follows:
[0030] Step 301: Improve the backbone network structure of the YOLOv8 model by embedding a Swing Transformer module before the SPPF in its backbone layer:
[0031] ,
[0032] in, and These are the output features of the previous submodule and the current submodule, respectively. Representative level normalization, This represents a typical multi-head self-attention window. This represents multi-head self-attention in a shifted window. This represents a multilayer perceptron. Indicates residual connection;
[0033] Step 302: Using the balanced dataset generated in Step 2, train and fine-tune the constructed ST-YOLO model:
[0034] For training the classification head: The training set images are processed through several convolutional layers in the Backbone to extract image features, which are then fed into the Swin Transformer module. Multi-scale feature extraction is achieved through block merging. The intermediate feature maps extracted by convolution are averaged across all pixels in each channel using a Global Average Pooling (GAP) layer. The output of the GAP is then fed into a fully connected (FC) layer and passed through a Softmax activation function layer to generate the polyp presence prediction probability. The classification head loss is obtained according to the classification head loss formula, and then ST-YOLO updates all parameters through backpropagation;
[0035] For training the detection head: the training set images are processed through several convolutional layers in the Backbone to extract image features, which are then fed into the Swin Transformer module. Multi-scale feature extraction is achieved through block merging. The resulting global feature map is fed into SPPF for max pooling, and then into the Neck for multi-scale feature fusion. Upsampling and cross-layer connections generate output feature maps at three scales. Each scale feature map is then input into the decoupled detection head, where mini-convolutions are used to generate feature maps. The output feature map The input is fed into the decoupling head; YOLOv8 abandons the Anchor strategy and directly outputs the predicted vector. ,in, These represent the distances from the current grid center point to the left, top, right, and bottom edges of the prediction box, respectively. The probability of the target class is represented by Logits transformed by a Sigmoid algorithm. The final coordinates are then generated through decoding. , , , , , The center coordinates of the current grid point. For the grid x-coordinate, The grid's vertical coordinates, , These are the center coordinates and width and height of the final predicted bounding box, respectively.
[0036] Detection head loss function It consists of three parts:
[0037]
[0038] in: For bounding box regression loss, To detect classification loss, For confidence / targeting loss, , , yes , , The weighting coefficients; all ST-YOLO parameters can be updated through backpropagation using the head loss; finally, the total loss parameters are adjusted. Fine-tuning is used to prevent the model from overfitting during training.
[0039] Step 4: Design and implement a multi-task strategy that simultaneously performs classification and detection. Input the image to be detected into the trained ST-YOLO model to determine whether the image contains polyps or not. Then, perform differential processing to further locate the polyp target. The specific steps are as follows:
[0040] Step 401: Design a strategy that enables simultaneous detection and classification, and equip the improved ST-YOLO model with a classification head training loss. and detection head training loss ;
[0041] Step 402: Use the trained classification head to classify the input colonoscopy images into images with polyps and images without polyps;
[0042] Step 403: For the classified images with polyps, activate the model's detection head and extract the polyp region.
[0043] Compared with the prior art, the present invention has the following advantages:
[0044] 1. This invention addresses the problem of imbalanced data by generating polyp-free images with highly realistic structure and texture through a CycleGAN-based improved model. This fundamentally solves the limitation of clinical data imbalance on model performance and lays a solid data foundation for subsequent high-precision detection.
[0045] 2. The innovative ST-YOLO model of this invention, by combining the global feature capture capability of Swing Transformer with the efficient detection framework of YOLO, greatly improves the detection accuracy and recall of small, occluded polyps in complex intestinal backgrounds.
[0046] 3. This invention adopts a multi-task synchronous strategy of "classification first, detection later", which avoids unnecessary detection calculations on polyp-free images, reduces the false detection rate, makes the entire detection process more in line with clinical diagnostic logic, and improves the accuracy and efficiency of diagnosis. Attached Figure Description
[0047] Figure 1 This is a flowchart of colorectal endoscopy image data enhancement and intelligent classification and detection of colorectal endoscopy images.
[0048] Figure 2 This is a flowchart of the image preprocessing process.
[0049] Figure 3 This is a schematic diagram of the CycleGAN cyclic process.
[0050] Figure 4 This is a comparison chart of the results of the polyp sample enhancement experiment.
[0051] Figure 5 This is a schematic diagram of the ST-YOLO network structure, which is an improvement on YOLOv8.
[0052] Figure 6 This is a diagram showing the effect of simultaneous detection and classification. Detailed Implementation
[0053] The technical solution of the present invention will be further described below with reference to the accompanying drawings, but it is not limited thereto. Any modifications or equivalent substitutions to the technical solution of the present invention that do not depart from the spirit and scope of the technical solution of the present invention should be covered within the protection scope of the present invention.
[0054] This invention provides a method for intelligent classification and detection of colorectal polyps based on data augmentation and improved YOLO, such as... Figure 1 As shown, the method includes the following steps:
[0055] Step 1: Preprocess and enhance the diversity of colorectal endoscopy images. Specific steps are as follows:
[0056] Step 101: Perform geometric transformation on the minority class samples (i.e., healthy colonoscopy images without polyps), including applying a random angle rotation to the image. Horizontal flip and vertical flip , , Represents the original coordinates. , Represents the transformed coordinates. This indicates the angle of rotation, simulating the image changes caused by different shooting angles during actual inspections.
[0057] Step 102: Convert the RGB image to the HSV color space. , , By adjusting the S channel and V channel This is used to adjust the saturation and brightness of the image, and to adjust the image contrast through linear transformation to simulate changes in light inside the intestines, further increasing data diversity and improving the realism of the healthy images generated by the subsequent model. Here, RGB represents the red, green, and blue primary color model of the image. It is the red component value of the pixel. It is the green component value of the pixel. It is the blue component value of the pixel. The hue represents the phase angle of the color. For saturation, Brightness is defined as the maximum value among the three RGB components. , This is the originally calculated saturation level. For the adjusted new saturation, This is the saturation adjustment coefficient. The original calculated brightness, For the adjusted new brightness, This is the brightness adjustment factor.
[0058] Step 2: To address the issues of unrealistic image shape and edge formation in traditional algorithms, the training loss function of the Recurrent Generative Adversarial Network (CycleGAN) is improved by introducing Structural Similarity (SSIM) loss. And add it to the total loss function. To improve the realism of polyp-free images generated from samples, an improved SCGAN model was used to generate polyp-free healthy colonoscopy images. The specific steps are as follows:
[0059] Step 201: Construct an image transformation framework based on a recurrent generative adversarial network (CycleGAN), which contains two generators. , With two discriminators , Train the generator to achieve The generator can generate images of class Y from images of class X. , The generator can generate an image of class X from an image of class Y. Simultaneously train the discriminator to judge and The generated results are similar to the dataset, and two cycle consistency rules are defined to ensure that... , Images generated by X and Y respectively, using , The generated image is similar to the original image. , Using this framework model, two metrics can be obtained:
[0060]
[0061] in, To counteract the loss, it is used to train the adversarial relationship between the generator and the discriminator; Cyclic consistency loss is used to ensure that the image generated by the generator is close to the original image after being generated repeatedly, thus preventing the generation of unrealistic images. For discriminator Real images of a healthy colon and rectum The expected value of the discrimination probability. A true image representing a healthy colon and rectum. A realistic image showing a colon or rectum with polyps. , Represents a generator. , This represents the discriminator, and the L1 norm is used to measure the difference between the generated image and the original image.
[0062] Step 202: Improve the loss function of CycleGAN. Based on the original adversarial loss and cycle consistency loss, introduce a structural similarity (SSIM) loss term to quantify the brightness of the two images. Contrast and structure Three dimensions of similarity are used to assess the differences between images, namely:
[0063]
[0064] in, , For image and The local mean value reflects brightness information; , Standard deviation, and The variance reflects the contrast. Covariance represents contrast. , and To ensure stability, the denominator should not be zero.
[0065] Structural Similarity Index Measure (SSIM) is introduced into the loss function of CycleGAN to give the generated colorectal images more realistic shape, structure, and edges.
[0066]
[0067] To unify the concept that smaller loss indicates better model performance, therefore... Mapping from [1, 0] to [0, 1]:
[0068]
[0069] in, Images of meat with polyps, Image of meat without breathing. Add it to the total loss function, that is:
[0070]
[0071] in, To combat the losses, For cycle consistency loss, The weights for the cyclic similarity loss are... For structural consistency loss, The weights for the structural consistency loss.
[0072] Set the sliding window size to 11×11, and use the local window calculation mechanism in SSIM. , , This method captures the structural features of local areas in colorectal endoscopy images, avoiding the problem of global statistical bias masking local distortions in the intestine. Among other things, The number of pixels. Represents pixel set, , For the first The x and y coordinates of each pixel. , For image and The local mean value reflects brightness information; , Standard deviation, and The variance reflects the contrast. Covariance represents contrast.
[0073] Step 203: Using the preprocessed minority class samples from Step 1 and the original majority class samples, the improved model incorporating SSMI and CycleGAN, referred to as SC-GAN, is trained. After training, an image with polyps is input, and the generator produces an image without polyps.
[0074] Step 204: Based on a series of evaluation metrics, specifically including: Mean Squared Error (MSE), Peak Signal-to-Noise Ratio (PSNR), Inception Score (IS), and Structural Similarity Index (SSIM), namely:
[0075]
[0076] in, This is the total pixel value. and The first The generated image and the pixel values of the real image are compared. The smaller the value, the more accurate the prediction.
[0077]
[0078] in, The maximum pixel value of the image. A higher value indicates less image quality loss.
[0079]
[0080] in, The distribution of the Inception model's predicted categories for generated images. It is a marginal category distribution. for divergence, The higher the value, the better the quality and the greater the diversity of the generated images.
[0081] The effective generation rate is derived by combining subjective and objective evaluation indicators to select images that pass quality standards. This rate is used to judge the quality of images generated in each experimental group. The effective generation rate comprehensively considers various indicators of the generated images, as well as their visual realism, structural rationality, and polyp residue. It selects the proportion of images suitable for input into the subsequent target detection model as healthy colorectal images from all generated images, forming a balanced dataset together with the polyp samples.
[0082] Step 3: Embed the Swin Transformer module before the SPPE layer of the backbone network in the YOLOv8 model to obtain the ST-YOLO model, and then train and fine-tune the parameters of the constructed ST-YOLO model. The specific steps are as follows:
[0083] Step 301: Improve the backbone network structure of the YOLOv8 model by adding a Spatial Pyramid Pooling Module (SPPF) to its backbone layer. , Previously, a SwingTransformer module was embedded:
[0084] ,
[0085] in, It is the input before the SPPF module, consisting of multiple feature layers of different scales. It is composed of concat elements. This is the basic feature map before entering the SPPE module. , , These are pooling features at different scales. It finally passes through the convolutional layer ( The processed output, and These are the mathematical tensors, which are the output data of the previous and current submodules, respectively. Layer Normalization. This represents a typical multi-head self-attention window. This represents multi-head self-attention in a shifted window. This represents a multilayer perceptron (typically consisting of two fully connected layers and a GELU activation function), while This indicates a residual connection.
[0086] Step 302: Using the balanced dataset generated in Step 2, train and fine-tune the constructed ST-YOLO model:
[0087] For training the classification head: The training set images are processed through several convolutional layers of the Backbone to extract image features. The convolution formula for each layer is as follows:
[0088]
[0089] in, Input feature maps into the classification head. For convolution kernel, For bias, The kernel size is [size]. Input the number of channels. The output channel index is fed into the Swing Transformer module for multi-scale feature extraction through block merging, preserving more detailed image information. ,in, This indicates that the output feature map after downsampling is in coordinates. The feature vector at that location, Represents the input feature map, its subscript , , and Refers to adjacent elements in the input image. In a pixel region, Concat means concatenating the features of these four pixels along the channel dimension, and LN represents layer normalization processing. Then it refers to the matrix containing learnable weights. A linear projection layer is used to map the concatenated high-dimensional features to the target dimension, thereby achieving multi-scale feature extraction while preserving details. The intermediate feature maps extracted by convolution are then subjected to global average pooling (GAP). Average all pixels in each channel to a single value, pass through a gap, and then connect a fully connected layer (FC): Next, the Softmax function is used to convert these scores into probabilities: This yields the predicted probability for each category. Then it is fed into the classification head, according to the classification head loss formula: The losses were incurred, among which, This represents the total number of samples (or the batch size). For the index of the sample, Indicates the first The true label of each sample (with a value of 0 or 1). The representative model predicts the first The probability value of each sample belonging to the positive class is then used by ST-YOLO to update all parameters through backpropagation, thereby continuously improving the model's classification performance.
[0090] For training the detection head: After extracting image features from the training set images through several convolutional layers in the Backbone, the images are fed into the Swin Transformer module. Multi-scale feature extraction is achieved through block merging, giving the model better accuracy and robustness in polyp detection tasks. The resulting feature maps, containing richer global information, are fed into SPPF for max pooling, and then into the Neck for multi-scale feature fusion. Upsampling and cross-layer connections generate output feature maps at three scales. The feature maps at each scale are then input into the Head, where decoupled branches and small convolutions generate the output feature maps. :
[0091]
[0092] in, The feature map is input to the detection head. It is a convolution kernel. It is a bias output. Represents spatial coordinates First The predicted values (Logits) for each output channel. Output feature map. The input is fed into the decoupling head; YOLOv8 abandons the Anchor strategy and directly outputs the predicted vector. ,in, These represent the distances from the current grid center point to the left, top, right, and bottom edges of the prediction box, respectively. The probability of the target class is represented by Logits transformed by a Sigmoid algorithm. The final coordinates are then generated through decoding. , , , , , The center coordinates of the current grid point. The x-coordinate of the grid. The grid's vertical coordinates, , This determines the center coordinates and width / height of the final predicted bounding box.
[0093] The head loss function typically consists of three parts:
[0094]
[0095] in: The bounding box regression loss measures the accuracy between the predicted and ground truth bounding boxes. YOLOv8 typically uses CIoU (Complete IoU) Loss or DFL (Distribution Focal Loss). To detect the classification loss, which measures whether the object within the bounding box is correctly classified (in this invention, whether it is a "polyp"), cross-entropy loss is also typically used. The confidence / targeting loss measures the model's information about whether a target (polyp) exists at a certain location, and is typically expressed using binary cross-entropy loss. , , These are the weight coefficients of these three sub-losses. After obtaining the loss, ST-YOLO updates all parameters through backpropagation, thereby continuously improving the model's classification performance.
[0096] Finally, by analyzing the total loss parameter... Fine-tuning is performed to prevent the model from overfitting during training, ensuring robust recognition capabilities for both polyp-containing and polyp-free images.
[0097] Step 4: Design and implement a multi-task strategy that performs classification and detection simultaneously. Input the image to be detected into the trained ST-YOLO model to determine whether the image is "with polyps" or "without polyps". Then perform differential processing to further locate the polyp target.
[0098] Step 401: Design a strategy that enables simultaneous detection and classification, and equip the improved ST-YOLO model with a classification head. and detection head .
[0099] Step 402: Use the classification head trained in the above steps to classify the input colonoscopy images into images with polyps and images without polyps.
[0100] Step 403: For the "polyp-containing" images classified in the above steps, activate the detection head of the model, extract the polyp region, and finally complete the entire process.
[0101] Example
[0102] The healthy colonoscopy images used in this embodiment are all 154 images obtained from a double-blind endoscopic experiment provided by the Department of Gastroenterology, Beijing Friendship Hospital, Capital Medical University; the polyp colonoscopy images are obtained from the open-source dataset Kvasir-SEG, which contains 1000 images.
[0103] Step 1: Image preprocessing workflow as follows Figure 2As shown, the input image is subjected to random angle rotation, horizontal / vertical flipping, random pixel cropping, and color space conversion for random saturation, brightness, and contrast through geometric transformations. Specifically, after converting the RGB image to the HSV color space, saturation and brightness are adjusted by adjusting the S and V channels, and image contrast is adjusted through linear transformation.
[0104] Step 2: After preprocessing, the dataset was expanded from 154 healthy colonoscopy images to 616 images. Traditional GAN models use 1000 images with polyps and 154 images without polyps as the dataset. The improved SCGAN model of this invention uses 1000 images with polyps and 616 images without polyps as the dataset, divided according to a training set:validation set:test set ratio of 7:2:1. The improved SCGAN model was trained using the dataset, and then the divided test set was augmented (Preprocess + SSIM), and compared with the image augmentation results of the traditional GAN model (Preprocess-only). The data augmentation results based on the SCGAN model and the traditional GAN model are as follows: Figure 4 As shown in Table 1, the effective generation rate of the enhanced results is as follows.
[0105]
[0106] from Figure 4 It can be clearly seen that the polyp-free colonoscopy images generated by the SCGAN model in the polyp-free sample enhancement experiment have richer and more realistic texture details than those generated by the traditional GAN model, and conform to the structure of the human colon and rectum. As shown in Table 1, the effective generation rate of the improved SCGAN model of this invention is better than that of the traditional GAN model.
[0107] Step 3: Embed the SwinTransformer module before the SPPF layer of the Backbone layer in the backbone network structure of the YOLOv8 model, such as... Figure 5 As shown, the previously generated balanced dataset is used to train and fine-tune the constructed ST-YOLO model to prevent overfitting during training, ensuring robust recognition capabilities for both polyp-containing and polyp-free images.
[0108] Step 4: Implement a multi-task strategy that simultaneously performs classification and detection. This strategy first uses the model's classification head to classify the pre-processed, augmented colonoscopy images into "polyp-containing" and "polyp-free" images. Then, it performs differential processing based on the classification results: if an image is determined to be "polyp-free," the process ends; if it is determined to be "polyp-containing," the model's detection head is immediately activated to accurately locate the polyp target in the image and output a bounding box containing the polyp location. The effect of simultaneous detection and classification is shown in the image below. Figure 6 As shown. Precision rate is used. Recall rate The performance of the ST-YOLO model with average accuracy mAP50-95 in classification and detection tasks was compared with that of the traditional YOLOv8 in two control groups: unprocessed colorectal endoscopy images (Base-Base) and preprocessed colorectal endoscopy images (Pre-SSIM-Base). The specific results are shown in Table 2. It can be seen that the improved ST-YOLO model is the best in all quantitative indicators of classification and detection tasks.
[0109]
Claims
1. A method for intelligent classification and detection of colorectal polyps based on data augmentation and improved YOLO, characterized in that... The method includes the following steps: Step 1: Preprocess and enhance the diversity of colorectal endoscopy images; Step 2: Improve the training loss function of CycleGAN by introducing structural similarity loss. and will Add to the total loss function In the process, the improved SC-GAN model was used to generate polyp-free healthy colonoscopy images; Step 3: Embed the SwinTransformer module before the Spatial Pyramid Pooling module (SPPF) in the Backbone layer of the YOLOv8 model to obtain the ST-YOLO model, and then train and fine-tune the parameters of the constructed ST-YOLO model. Step 4: Design and implement a multi-task strategy that performs classification and detection simultaneously, and input the image to be detected into the trained ST-YOLO model to determine whether the image has polyps or not, and then perform differential processing to further locate the polyp target.
2. The intelligent classification and detection method for colorectal polyps based on data augmentation and improved YOLO according to claim 1, characterized in that... The specific steps of step 1 are as follows: Step 101: Perform geometric transformation processing on the minority sample, namely healthy colonoscopy images without polyps, to simulate the image changes caused by different shooting angles in actual examinations. Step 102: Convert the RGB image to the HSV color space. Adjust the saturation and brightness of the image by adjusting the S and V channels. Adjust the image contrast by linear transformation to simulate the changes in light inside the intestine, further increasing the diversity of data and improving the realism of the healthy images generated by the subsequent model.
3. The method for intelligent classification and detection of colorectal polyps based on data augmentation and improved YOLO according to claim 1, characterized in that... The specific steps of step 2 are as follows: Step 201: Construct an image transformation framework based on CycleGAN, and use this framework model to obtain two metrics: in, To combat the losses, For cycle consistency loss, For discriminator Real images of a healthy colon and rectum The expected value of the discrimination probability. This dataset represents a colorectal region containing polyps. , Represents a generator. , Indicates the discriminator; Step 202: Improve the loss function of CycleGAN by introducing a structural similarity loss term in addition to the original adversarial loss and cycle consistency loss. Quantize the brightness of two images Contrast and structure Three dimensions of similarity are used to assess the differences between images, namely: in, , For image and The local mean, , Standard deviation and For variance, For covariance, , and It is the stability constant; Will Mapping from [1, 0] to [0, 1]: in, Images of meat with polyps, Images of meat without breathing surface; Will Add it to the total loss function, that is: in, The weights for the cycle consistency loss, The weights for structural consistency loss; The local windowing mechanism of SSIM is used to capture the structural features of local regions in colorectal endoscopy images; Step 203: Using the preprocessed minority class samples and the original majority class samples from Step 1, train the SC-GAN model improved by introducing SSMI and CycleGAN. After training, input an image with polyps and the generator will generate an image without polyps. Step 204: Combine subjective and objective evaluation indicators to screen images that pass the quality test, obtain the effective generation rate, and use it to judge the quality of the images generated in each group of experiments. Finally, select images without polyps and images with polyps to form a balanced dataset.
4. The intelligent classification and detection method for colorectal polyps based on data augmentation and improved YOLO according to claim 1, characterized in that... The specific steps of step 3 are as follows: Step 301: Improve the backbone network structure of the YOLOv8 model by embedding a Swing Transformer module before the SPPF in its backbone layer: 、 in, and These are the output features of the previous submodule and the current submodule, respectively. Representative level normalization, This represents a typical multi-head self-attention window. This represents multi-head self-attention in a shifted window. This represents a multilayer perceptron. Indicates residual connection; Step 302: Use the balanced dataset generated in Step 2 to train and fine-tune the constructed ST-YOLO model.
5. The method for intelligent classification and detection of colorectal polyps based on data augmentation and improved YOLO according to claim 4, characterized in that... In step 302, for training the classification head: the training set images are processed through several convolutional layers of the Backbone to extract image features, which are then fed into the Swin Transformer module. Multi-scale feature extraction is achieved through block merging. The intermediate feature maps extracted by convolution are averaged into a single value using a Global Average Pooling (GAP) layer. The output of GAP is then fed into a fully connected (FC) layer and passed through a Softmax activation function layer to generate a polyp presence prediction probability. The classification head loss is obtained according to the classification head loss formula, and then ST-YOLO updates all parameters through backpropagation. For the training of the detection head: the training set images are processed through several convolutional layers of the Backbone to extract image features, which are then fed into the Swin Transformer module. Multi-scale feature extraction is achieved through block merging. The obtained global information feature map is fed into SPPF for max pooling, and then into the Neck for multi-scale feature fusion. Three-scale output feature maps are generated through upsampling and cross-layer connections. The feature map of each scale is input into the decoupled detection head, and feature maps are generated using small convolutions. The output feature map The input is fed into the decoupling head; YOLOv8 abandons the Anchor strategy and directly outputs the predicted vector. ,in, These represent the distances from the current grid center point to the left, top, right, and bottom edges of the prediction box, respectively. The probability of the target category is then used to generate the final coordinates through decoding. , , , , , The center coordinates of the current grid point. The x-coordinate of the grid. The grid's vertical coordinates, , The final predicted center coordinates and dimensions of the bounding box; Detection head loss function It consists of three parts: in: For bounding box regression loss, To detect classification loss, For confidence / targeting loss, , , yes , , The weighting coefficients; all ST-YOLO parameters can be updated through backpropagation using the head loss; finally, the total loss parameters are adjusted. Fine-tuning is used to prevent the model from overfitting during training.
6. The intelligent classification and detection method for colorectal polyps based on data augmentation and improved YOLO according to claim 1, characterized in that... The specific steps of step 4 are as follows: Step 401: Design a strategy that enables simultaneous detection and classification, and equip the improved ST-YOLO model with a classification head training loss. and training loss of the detection head ; Step 402: Use the trained classification head to classify the input colonoscopy images into images with polyps and images without polyps; Step 403: For the classified images with polyps, activate the model's detection head and extract the polyp region.
Citation Information
Patent Citations
Substation potential safety hazard detection method based on window self-attention mechanism
CN116682057A
Low-illumination small target detection method based on SCKConv multi-scale feature fusion enhancement
CN117409244A
Oral squamous cell carcinoma medical image segmentation method based on improved U-Net network
CN117437240A
Apple maturity detection method based on Swin Transform
CN117893823A
Colorectal polyp detection and classification method based on improved YOLOv7 model
CN118038161A