A method and system for detecting the size of colonic polyps under colonoscopy

By combining YOLOv5 and DepthAnything models, accurate positioning and dimensional measurement of colon polyps are achieved, which solves the problem of large detection errors in the existing technology, and improves the accuracy and automation of detection.

CN119205794BActive Publication Date: 2025-06-13ZHEJIANG RADIOLOGY INFORMATION TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411733238.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-29
Publication Date
2025-06-13
Estimated Expiration
2044-11-29

AI Technical Summary

Technical Problem

In the prior art, colon polyp size detection depends on the experience and visual evaluation of the doctor, and is subjective and has great errors, especially in the absence of suitable references, which is difficult to provide accurate measurement results.

Method used

Combining the YOLOv5 colon polyp detection model and the DepthAnything depth estimation model, the precise positioning and size measurement of polyp are achieved. By training the model to detect the polyp and estimate its depth map, determine the distance between the midgut wall and the lens of the colonoscopy, reconstruct the three-dimensional point cloud map, and calculate the actual physical size of the polyp.

Benefits of technology

Improve the accuracy of polyp size detection, reduce measurement errors due to the lack of references, realize automated detection, simplify operational processes and reduce costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119205794B_ABST
    Figure CN119205794B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for detecting the size of colon polyps under a colonoscope. The present invention combines the YOLOv5 colon polyp detection model and the DepthAnything depth estimation model, which can not only accurately detect polyps in colonoscopy images, but also estimate a precise depth map. Subsequently, the distance between the intestinal wall and the lens during colonoscopy can be determined using the polyp bounding box and the depth map, and the three-dimensional point cloud map of the colonoscopy image in the real world can be reconstructed, the size of the polyp in the point cloud can be calculated, and converted into the actual physical size of the polyp, which is displayed in real time during the operation. The present invention realizes the precise positioning and size measurement of polyps, and avoids the measurement errors caused by the lack of reference objects in traditional methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical image processing, and particularly relates to a method and system for detecting the size of colon polyps. Background Art

[0002] A colonoscope is an important tool in modern medicine for detecting and diagnosing colon diseases. By inserting a flexible tube with a camera into the rectum, doctors can observe the inside of the colon in real time and form a continuous video stream. This technology is crucial for detecting abnormal proliferative tissues (such as polyps) in the colon, as timely detection and removal of these tissues can greatly reduce the risk of colon cancer.

[0003] After detecting a colon polyp, it is classified according to its diameter as: micro polyps ≤5 mm, small polyps >5 - 10 mm, large polyps >10 - 20 mm, and giant polyps >20 mm. The treatment methods for polyps of different sizes vary. The size detection of colon polyps mainly relies on doctors' experience and visual assessment. Doctors will estimate according to the appearance of the polyp on the screen, referring to instruments of known size or the colon wall. However, this method is highly subjective and prone to errors. Especially when there is a lack of a suitable reference in the video, it is difficult to provide accurate measurement results, and incorrect size estimation of colon polyps may lead to waste of medical devices.

[0004] To solve these problems, recent research has started to explore automated methods to assist in polyp size detection. Based on the perspective phenomenon of objects being larger when closer and smaller when farther away in colonoscopy images, some early methods tried to take images at a preset distance and use a fixed zoom factor multiplied by the size of the polyp in the image at the determined distance to predict the true size of the polyp. Or, instead of taking images at a preset distance, a laser rangefinder lens was added to the colonoscope to measure the distance between the intestinal wall at the center of the image and the lens in real time, and a virtual scale was superimposed and displayed on the screen for doctors to use to measure the true size of the polyp. Additionally, some used reference objects of a determined size such as biopsy forceps or water jets, and used deep learning methods to automatically compare the polyp and the reference object to predict the true size of the polyp. However, from the perspective of practical application, some of these methods are inconvenient to operate, and some require purchasing additional lenses, and they still need to be further improved. Summary of the Invention

[0005] In view of the problems existing in the above-mentioned background technology, the present invention proposes a method and system for detecting the size of colon polyps under a colonoscope. The present invention combines the YOLOv5 colon polyp detection model and the depth estimation model based on DepthAnything, which can not only accurately detect polyps in colonoscope images, but also estimate an accurate depth map. Subsequently, the distance between the intestinal wall and the lens during colonoscopy can be determined using the polyp bounding box and the depth map, and the three-dimensional point cloud map of the colonoscope image in the real world can be reconstructed, the size of the polyp in the point cloud can be calculated, and converted into the actual physical size of the polyp, which is displayed in real time during the operation.

[0006] The method of the present invention specifically includes the following steps:

[0007] Step 1: Train the YOLOv5 colon polyp detection model

[0008] The present invention uses a dataset of colon polyps to train the YOLOv5 object detection model, which is widely used in the field of object detection for its high efficiency and accuracy.

[0009] Step 1.1: Preparation and preprocessing of the colon polyp detection dataset

[0010] The colon polyp detection dataset consists of a large number of colonoscopy video frames, which contain a large number of various types of colon polyps and have been professionally annotated to mark the presence and location of polyps. The annotation format is , where is the offset of the center of the bounding box relative to the upper left corner of the image, and are the ratio values of the width and height of the bounding box relative to the width and height of the image respectively.

[0011] Perform size normalization processing on the original video frames. First, crop and remove the black edges, then fill the square with a gray value of 128 and scale it to the size of L×L using bilinear interpolation, where L represents the side length of the input image, to meet the requirements of the YOLOv5 model. Finally, randomly divide the dataset into three parts: training set, validation set, and test set. The training set is used for model training, the validation set is used for hyperparameter tuning, and the test set is used for evaluating the model performance. At the same time, apply image enhancement techniques such as brightness adjustment, rotation, and mapping to the training set to enhance the robustness of the model.

[0012] Step 1.2: Training process of the YOLOv5 colon polyp detection model

[0013] First, load the pre-trained weights provided by YOLOv5 as the initial state of the model, and train using the training set described in Step 1.1. The model divides the feature map into blocks using a chessboard grid, and each block is associated with in the model output. Vectors, each vector representing a prediction box, where represent the offsets of the prediction box relative to the top - left corner of the block in the x and y directions respectively, represents the length and width of the prediction box, represents the probability that the prediction box contains an object, represents the one - hot encoding of the predicted class value of the prediction box, and there are a total of categories. Loss function:

[0014]

[0015] where represents the output feature map. The difference between different feature maps lies in the number of blocks being different; is used to balance the weights of each output feature map; represents the th th block of the th prediction box of the nd feature map. Whether it is a positive sample (matched with a ground truth) is 1 if it is a positive sample, otherwise 0;

[0016]

[0017] is the complete intersection - over - union loss, represent the prediction box and the ground truth box respectively; represents the distance between the centers of the prediction box and the ground truth box, is the diagonal distance of the minimum bounding rectangle of the prediction box and the ground truth box; w G , h G represent the width and height of the prediction box respectively, w P , h P represent the width and height of the ground truth box respectively.

[0018]

[0019] is the object confidence loss, where represents the binary cross - entropy loss function, is the probability predicted by the prediction box, is the value of the prediction box and the corresponding ground truth box, represents the weight of the positive sample.

[0020]

[0021] is the classification confidence loss, where is the cross - entropy loss function, is the predicted class of the bounding box, is the true class, represents the weight of the true class.

[0022] In each epoch, the model predicts the images in the training set, calculates the loss value according to the loss function, and then updates the weights through backpropagation. In addition, the performance on the validation set is used to adjust the learning rate and other hyperparameters. After training is completed, the best model weights are saved, and the model performance is evaluated on the test set, including metrics such as accuracy, recall, and mAP score.

[0023] Step 2: Train the DepthAnything depth estimation model

[0024] The present invention trains the DepthAnything depth estimation model using a colonoscopy depth estimation dataset, and performs semi - supervised training using both labeled virtual images and unlabeled real images. Through training on synthetic data with true labels, the model can learn the scale information of image depth and physical size; self - supervised learning in real clinical data can further optimize the model, improve the model's understanding ability of complex conditions in the real environment, enhance the stability of depth estimation, and make it perform better in the real clinical environment. Specifically:

[0025] Step 2.1: Preparation and pre - processing of the colonoscopy depth estimation dataset

[0026] The colonoscopy depth estimation dataset includes a labeled virtual image dataset and an unlabeled real image dataset. The virtual image dataset is a public dataset, generally including images taken of in - vitro prostheses or computer - simulated images, and the dataset provides the camera intrinsic matrix and the camera distortion coefficients . The real image dataset is a series of unlabeled colonoscopy videos obtained from hospitals.

[0027] For the real clinical dataset, the camera intrinsic matrix and the camera distortion coefficients of the endoscope need to be obtained through manual calibration. First, prepare a black - and - white checkerboard with known size and number of corner points, then place this calibration pattern at different positions and angles, use the endoscope lens to be calibrated to take multiple images containing the calibration pattern from multiple perspectives, and finally use the OpenCV library to solve for the camera intrinsic matrix and the camera distortion coefficients .

[0028] The data preprocessing part includes: if the data is a video, filter out the blurred frames of the video, retain the relatively high-quality part, then split it into pictures, and perform fisheye lens correction on the pictures using the in-lens matrix and the lens distortion coefficient. Divide the input pictures by 255 to scale them to range; the depth ground truth of the virtual pictures is distributed in millimeters, divided by to scale it to range. Finally, the depth ground truth of the input pictures and virtual pictures after removing distortion are uniformly scaled to pixels.

[0029] Step 2.2: Train the DepthAnything depth estimation model on the labeled virtual picture dataset

[0030] Train the DepthAnything depth estimation model on the labeled virtual picture dataset, and call this model the Teacher model. The input image of the model has a dimension of , and the model output has a dimension of , and the loss function:

[0031]

[0032] Step 2.3: Train the DepthAnything illumination adjustment model

[0033] The present invention uses the traditional pinhole camera model. The camera is located at the origin of the world coordinate system and looks towards the negative Z-axis, the X-axis is horizontal to the left, and the Y-axis is vertically pointing to the ground. The camera internal parameter matrix is defined as , define the point in the world coordinate system as , and through map it to the pixel coordinate .

[0034] First, use the parameters obtained by lens calibration to remove distortion from the input frames, and scale the processed images to size, then use the trained Teacher model to predict its depth map , and then use the following formula to restore the pixels representing the depth values in to the point cloud in the world coordinate system:

[0035]

[0036] Among them, for the pixel with pixel coordinates , represents its world coordinate, and represents its distance from the lens (depth value).

[0037] After obtaining the point cloud, calculate the per-pixel shading (PPS) of the image according to the following formula:

[0038]

[0039]

[0040] where is the coordinate of the co-located light of the endoscope, is the illumination direction, is the angular attenuation coefficient of the illumination. The per-pixel shading shows the brightness that the colonoscopy image should have under the illumination of only the co-located light of the lens.

[0041] Train the DepthAnything illumination adjustment model. The grayscale image of the real image and its depth map are concatenated in the channel dimension to get . Downsample by bilinear interpolation with the side length reduced by four times to become as the model input. The model predicts a correction for the grayscale image of the real image, and then upsample it by bilinear interpolation with four times to get . The loss function is:

[0042]

[0043] After training is completed, use the trained DepthAnything illumination adjustment model to predict the grayscale image of the real image, obtain and save the result for subsequent model training.

[0044] Step 2.4: Train the DepthAnything depth estimation model using both virtual image and real image datasets

[0045] Train the DepthAnything depth estimation model using both virtual image and real image datasets, and call this model the Student model. The input image has a dimension of , and the model output has a dimension of . When the input is a virtual image, the loss function is:

[0046]

[0047] When the input is a real image, the loss function is:

[0048]

[0049] wherein is input obtained by the formula.

[0050] Step 3: Use the trained YOLOv5 colon polyp detection model and the Student model to detect the polyp size

[0051] The present invention jointly uses the trained YOLOv5 colon polyp detection model in Step 1 and the trained Student model in Step 2 to detect the polyp size.

[0052] For the input image frame, first use the internal camera matrix of the endoscope obtained by the lens calibration in Step 2.1 and the lens distortion coefficients to perform lens correction, and then divide by to scale it to within the range, and finally use bilinear interpolation to scale it to pixels, denoted as .

[0053] Input into the YOLOv5 colon polyp detection model to obtain the colon polyp bounding box matrix in the picture, which contains bounding box vectors of colon polyps , and all these four values are within , wherein is the offset of the bounding box center relative to the upper left corner of the image, are respectively the proportional values of the bounding box width and height relative to the width and height of the image.

[0054] Input into the Student model to obtain the depth map of this picture. Use the formula to restore the pixels representing the depth values in to the point cloud in the world coordinate system, where the superscript represents the three coordinate axes, represents that there are a total of points.

[0055] For each bounding box vector of the colon polyps, calculate the midpoint coordinates of the four sides of the bounding box, and then take out the coordinates in . Calculate of​ , where and respectively represent in the direction, the maximum and minimum values, and respectively represent in the direction, the maximum and minimum values, and finally the size of the polyp is expressed as .

[0056] Based on the concept of the above method steps, the present invention also provides a colon polyp size detection system under a colonoscope, including:

[0057] The first training module trains the YOLOv5 object detection model using the dataset of colon polyps to obtain the YOLOv5 colon polyp detection model;

[0058] The second training module trains the DepthAnything depth estimation model using the colonoscope depth estimation dataset, and simultaneously performs semi-supervised training using labeled virtual images and unlabeled real images to obtain the Student model;

[0059] The detection module jointly uses the trained YOLOv5 colon polyp detection model and the trained Student model to detect the size of the polyp.

[0060] The beneficial effects of the present invention are as follows: The present invention combines the YOLOv5 colon polyp detection model and the DepthAnything depth estimation model to achieve accurate positioning and size measurement of polyps. This method can automatically detect polyps and estimate their actual physical sizes, improving the accuracy of image recognition and avoiding measurement errors caused by the lack of reference objects in traditional methods. In addition, this technical solution is lightweight, can be deployed immediately, does not require additional equipment support, simplifies the operation process, reduces costs, and helps to detect and process lesion tissues that may lead to cancer at an early stage. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. The drawings form a part of the specification and are used together with the embodiments of the present invention to explain the present invention, but do not constitute a limitation to the present invention. Those skilled in the art should understand that these drawings are not necessarily drawn to scale. In the drawings:

[0062] Figure 1 is the overall flowchart of the method of the present invention;

[0063] Figure 2 is the application flowchart of the method of the present invention;

[0064] Figure 3 Schematic diagram of lens calibration and distortion removal in the embodiments of the present invention;

[0065] Figure 4 Schematic diagram of clinical real image and depth map in the embodiments of the present invention. Detailed implementation manners

[0066] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application.

[0067] As Figure 1 and Figure 2 shown, an endoscopic colon polyp size detection method is provided in an embodiment of the present application. The specific steps of the method are as follows:

[0068] Step 1: Train a YOLOv5 colon polyp detection model

[0069] Step 1.1: Preparation and preprocessing of a colon polyp detection data set

[0070] The hyper-kvasir gastroscopy public data set is adopted in this embodiment, which contains 110,079 images and 374 video diagnostic processes, including the bounding boxes of 1,000 images of polyp types.

[0071] Set a pixel brightness threshold to preliminarily screen out reflective and overly dark video frames, and retain relatively clear parts. Convert the bounding box annotation format to .

[0072] Finally, fill the pictures with a gray value of 128 to form squares, and then use the bilinear interpolation method to scale them to a size of 640×640.

[0073] Randomly divide the data set into three parts: a training set, a validation set, and a test set at a ratio of . The training set is used for model training, the validation set is used for hyperparameter tuning, and the test set is used to evaluate the model performance. At the same time, image enhancement techniques such as brightness adjustment, rotation, and mapping are applied to the training set to enhance the robustness of the model.

[0074] Step 1.2: Training process of the YOLOv5 colon polyp detection model

[0075] First, load the pre-trained weights provided by YOLOv5 as the initial state of the model, and train using the training set described in Step 1.1. The model divides the feature map into types of block sizes with a checkerboard grid into blocks, each block is associated with the vector in the model output. Each vector represents a prediction box, where respectively represent the offsets of the prediction box in the x and y directions relative to the upper left corner of the block, represent the length and width of the prediction box, represents the probability that the prediction box contains a target, represents the one-hot encoding of the class prediction value of the prediction box. There are a total of classes, namely polyps and background. Loss function:

[0076]

[0077] where represents the output feature maps, that is, the feature maps divided by chunk sizes, represents the number of chunks of the th feature map; is used to balance the weights of each output feature map. For the feature map of size, are respectively represents whether the th prediction box of the th block of the th feature map is a positive sample (matched with a ground truth). If it is a positive sample, it is 1, otherwise it is 0; is the weight of the corresponding item's loss function, which is here.

[0078] In one embodiment, is expressed as follows:

[0079]

[0080] where respectively represent the prediction box and the ground truth box; represents the distance between the centers of the prediction box and the ground truth box, is the diagonal distance of the minimum bounding rectangle of the prediction box and the ground truth box; w G , h G respectively represent the width and height of the prediction box, w P , h P respectively represent the width and height of the ground truth box.

[0081] In one embodiment, is expressed as follows:

[0082]

[0083] where represents the binary cross-entropy loss function, is the probability predicted by the prediction box, is the value between the prediction box and the corresponding ground truth box, represents the weight of the positive samples.

[0084] In one embodiment, it is expressed as follows:

[0085]

[0086] where is the cross-entropy loss function, is the predicted class of the prediction box, is the true class, represents the weight of the true class, which is here.

[0087] In each epoch, the model makes predictions on the images in the training set, calculates the loss value according to the loss function, and then updates the weights through backpropagation. After the training is completed, the best model weights are saved, and the model performance is evaluated on the test set, including metrics such as accuracy, recall, and mAP score. The final results are shown in Table 1.

[0088] Table 1 Test Results of YOLOv5 Colorectal Polyp Detection Model

[0089]

[0090] Step 2: Train the DepthAnything depth estimation model

[0091] Step 2.1: Preparation and preprocessing of the colonoscopy depth estimation dataset

[0092] The colonoscopy depth estimation dataset includes an annotated virtual image dataset and an unannotated real image dataset.

[0093] The virtual image dataset used is C3VD (Colonoscopy 3D Video Dataset). The C3VD dataset is a 3D video dataset specifically designed for colonoscopy, aiming to provide high-precision registration data for the application of computer vision technology in the medical field. This dataset integrates 22 colonoscopy videos processed by multi-modal 2D-3D registration technology. The videos were taken in vitro models and include 2,363 sigmoid colon image frames, 2,736 transverse colon image frames, 148 descending colon image frames, and 4,768 ascending colon image frames, totaling 10,015 frames of images. Each frame is accompanied by paired true depth, surface normal, optical flow, occlusion, six degrees of freedom (HDoF) camera pose information, and a coverage map. The dataset also provides a 3D model file of the colon and open-source code for registration and rendering. In addition, the dataset includes 4 colonoscopy videos and corresponding camera pose annotations. The C3VD dataset also provides the camera intrinsic matrix and the lens distortion coefficients .

[0094] The real image dataset is a series of 30 unannotated colonoscopy videos obtained from a hospital. For the real clinical dataset, the camera intrinsic matrix and the lens distortion coefficients of the endoscope need to be obtained through manual calibration. The calibration process is as follows: First, prepare a black and white checkerboard with known dimensions and the number of corner points. Then, place this calibration pattern at different positions and angles, and use the endoscope lens to be calibrated to take images containing the calibration pattern from multiple perspectives. Finally, use the OpenCV library to solve for the camera intrinsic matrix of the endoscope and the lens distortion coefficients . The situations of the checkerboard and the lens before and after distortion removal are shown in Figure 3 .

[0095] The data preprocessing part includes: screening out blurred video frames, retaining the relatively high-quality part, then splitting them into pictures, and using the camera intrinsic matrix and lens distortion coefficients to correct the fisheye lens of the pictures. The input pictures are divided by to be scaled to the range of ; the true depth values of the virtual pictures are distributed in millimeters and are divided by to be scaled to the range of . Finally, the undistorted input pictures and the true depth values of the virtual pictures are bilinearly interpolated and uniformly scaled to pixels.

[0096] Step 2.2: Train the DepthAnything depth estimation model on the annotated virtual image dataset

[0097] Fifteen videos in the C3VD dataset are designated as the training set, and seven videos are designated as the validation set. Train the DepthAnything depth estimation model on the training set, and call this model the Teacher model. The input image of the model has dimensions of , and the model output has dimensions of . The loss function:

[0098]

[0099] Finally, the average loss of the Teacher model on the C3VD validation set is millimeters.

[0100] Step 2.3: Train the DepthAnything illumination adjustment model

[0101] In this embodiment, a traditional pinhole camera model is used. The camera is located at the origin of the world coordinate system and looks towards the negative Z-axis, with the X-axis horizontal to the left and the Y-axis vertically pointing to the ground. The camera internal parameter matrix is defined as , and the point in the world coordinate system is defined as , and is mapped to the pixel coordinate through .

[0102] First, use the parameters obtained from lens calibration to remove distortion from the input frame, and scale the processed image to size. Then, use the trained Teacher model to predict its depth map , and then use the following formula to restore the pixels representing depth values in to a point cloud in the world coordinate system:

[0103]

[0104] where, for a pixel coordinate of pixels, represents its world coordinate, and represents its distance from the lens (depth value).

[0105] After obtaining the point cloud, calculate the per-pixel shading (PPS) of the image according to the following formula:

[0106]

[0107] where is the coordinate of the co-located light of the endoscope, which is , and is the illumination direction, which is the negative Z-axis direction. The angular attenuation coefficient for light is .

[0108] Use all real images to train the DepthAnything light adjustment model. The grayscale image of the real image and its depth map The concatenation result in the channel dimension is . Take The side length of is bilinearly interpolated and downsampled by a factor of four to become as the model input. The model predicts a correction to the grayscale image of the real image , and then bilinearly interpolates and upsamples by a factor of four to obtain . The loss function:

[0109]

[0110] After training, use the trained DepthAnything light adjustment model to predict the grayscale image of the real image , obtain and save the results for use in subsequent model training.

[0111] Finally, the average brightness loss on the training set is .

[0112] Step 2.4: Train the DepthAnything depth estimation model using both virtual and real image datasets

[0113] Train the DepthAnything depth estimation model using both the training set of the C3VD dataset and the real image dataset. Call this model the Student model. The model input image has a dimension of , and the model output has a dimension of . When the input is a virtual image, the loss function is:

[0114]

[0115] When the input is a real image, the loss function is:

[0116]

[0117] where is the input obtained from the formula.

[0118] Finally, the average loss of the Student model on the C3VD validation set is Millimeters.

[0119] Step 3: Use the trained YOLOv5 colon polyp detection model and Student model to detect the size of polyps, see Figure 4 .

[0120] The present invention jointly uses the YOLOv5 colon polyp detection model trained in step 1 and the Student model trained in step 2 to detect the size of polyps.

[0121] For the input image frame, first use the endoscope lens intrinsic parameter matrix obtained by step 2.1 lens calibration And the lens distortion coefficient Perform lens correction and divide by To zoom to In the range, bilinear interpolation is used to scale to Pixel, denoted as .

[0122] Will Input the YOLOv5 colon polyp detection model to get the colon polyp bounding box matrix in the image , which includes Bounding box vector of a colon polyp , these four values ​​are in Among them is the offset of the center of the bounding box relative to the upper left corner of the image, are the ratios of the bounding box width and height relative to the image width and height, respectively.

[0123] Will Input the Student model to get the depth map of the image , using the formula Will The pixels representing the depth value are restored to the point cloud in the world coordinate system ,in express Three coordinate axes, Indicates a total of points.

[0124] for For each colon polyp in the bounding box vector, calculate the midpoint coordinates of the four sides of the bounding box , then remove exist Coordinates in .calculate of ,in and Respectively exist The maximum and minimum values in the and respectively represent the maximum and minimum values in the direction. Finally, the size of the polyp is expressed as .

[0125] An embodiment of the present application further provides a system for detecting the size of a colon polyp under a colonoscope, including:

[0126] A first training module that trains a YOLOv5 object detection model using a dataset of colon polyps to obtain a YOLOv5 colon polyp detection model;

[0127] A second training module that trains a DepthAnything depth estimation model using a colonoscope depth estimation dataset and simultaneously performs semi-supervised training using labeled virtual images and unlabeled real images to obtain a Student model;

[0128] A detection module that jointly uses the trained YOLOv5 colon polyp detection model and the trained Student model to detect the size of the polyp.

[0129] Finally, it should be noted that the above are only the preferred embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be covered by the protection scope of the present invention.

Claims

1. A method for detecting the size of colon polyps under colonoscopy, characterized in that The method comprises the following steps: Step 1: Use the colon polyp dataset to train the YOLOv5 target detection model to obtain the YOLOv5 colon polyp detection model; Step 2: Use the colonoscopy depth estimation dataset to train the DepthAnything depth estimation model, and use both labeled virtual images and unlabeled real images for semi-supervised training to obtain the Student model; Step 3: Combine the YOLOv5 colon polyp detection model trained in step 1 and the Student model trained in step 2 to detect the size of polyps; The step 2 specifically includes: Step 2.1: Preparation and preprocessing of colonoscopy depth estimation dataset The colonoscopy depth estimation dataset includes a labeled virtual image dataset and an unlabeled real image dataset; Preprocessing includes: if the data is a video, filter out the blurred frames of the video, split it into pictures, and use the lens intrinsic parameter matrix and lens distortion coefficient to correct the pictures with fisheye lens; Step 2.2: Train the first DepthAnything depth estimation model on the labeled virtual image dataset to obtain the Teacher model; Step 2.3: Train the DepthAnything illumination adjustment model to predict the grayscale image of the real image; Use the trained Teacher model to predict the depth map D of the image teacher , and the depth map D teacher The pixels representing the depth values ​​are restored to the point cloud in the world coordinate system: X(u,v)=D(u,v)K -1 (u,v,1) T (17) Among them, for a pixel with pixel coordinates (u, v), X(u, v) represents its world coordinates, D(u, v) represents its distance from the lens, that is, the depth value; K is the definition of the camera intrinsic parameter matrix; After obtaining the point cloud, calculate the pixel-by-pixel light and shadow PPS(X) of the image: PPS(X)=Att(X)×[LD(X)·Nor(X)] (18) in The coordinates of the co-location light of the endoscope are (0,0), The lighting direction is the negative direction of the Z axis. The angular attenuation coefficient of the illumination is 1; Use all real images to train the DepthAnything illumination adjustment model, grayscale image of the real image With its depth map D teacher The splicing result in the channel dimension is Ref in The side length of bilinear interpolation is downsampled by four times to As the model input, the model predicts a correction to the grayscale image of the real image Then upsample four times by bilinear interpolation to get The loss function is: After training, the trained DepthAnything illumination adjustment model is used to predict the grayscale image of the real image. gray , get and save the result PPS ref =Ref out I gray For subsequent model training; Step 2.4: Use the virtual image and real image datasets to train the second DepthAnything depth estimation model to obtain the Student model; When the input is a virtual image, the loss function is: When the input is a real picture, the loss function is: Among them, PPS stu Yes D pred_stu Input equations (18) and (19); The step 3 is specifically: The preprocessed image is input into the YOLOv5 colon polyp detection model to obtain the colon polyp bounding box matrix Polyps in the image bbox ; Input the preprocessed image into the Student model to obtain the depth map D of the image student ; The depth map D student The pixels representing the depth values ​​are restored to the point cloud in the world coordinate system; For the colon polyp bounding box matrix Polyps bbox The bounding box vector of each colon polyp in the point cloud is calculated, and the midpoint coordinates of the four sides of the bounding box are taken as the coordinates PC of the midpoint coordinates in the point cloud. center ; Calculate the coordinates PC center (x min ,x max ,y min ,y max ), where x max With x min Represents the coordinates PC center The maximum and minimum values ​​in the x direction, y max With y min Respectively represent PC center Maximum and minimum values ​​in the y direction; Denote the size of the polyp as (x max -x min ,y max -y min ).

2. The method for detecting the size of colon polyps under colonoscopy according to claim 1, characterized in that: The step 1 specifically includes: Step 1.1: Prepare and preprocess the colon polyp detection dataset; The colon polyp detection dataset consists of a large number of colonoscopy video frames, which contain a large number of various types of colon polyps and are professionally annotated to mark the presence and location of polyps; Normalize the size of the original video frames to meet the requirements of the YOLOv5 target detection model; The data set is randomly divided into three parts: training set, validation set and test set; Step 1.2: YOLOv5 colon polyp detection model training; Load the pre-trained weights provided by YOLOv5 as the initial state of the model, and train using the training set described in step 1.

1. In each epoch, the model predicts the images in the training set, calculates the loss value based on the loss function, and then back-propagates to update the weights. In addition, the performance on the validation set is used to adjust the learning rate and other hyperparameters. After training is complete, the best model weights are saved and the model performance, including precision, recall, and F1 score, is evaluated on the test set.

3. A method for detecting the size of colon polyps under colonoscopy according to claim 2, characterized in that: The annotation format is (x center ,y center ,w,h), where (x center ,y center ) is the offset of the center of the bounding box relative to the upper left corner of the image, and (w,h) are the ratios of the bounding box width and height to the image width and height, respectively.

4. A method for detecting the size of colon polyps under colonoscopy according to claim 2 or 3, characterized in that: It also includes image augmentation by applying brightness adjustment, rotation, and mapping to the training set to enhance the robustness of the model.

5. The method for detecting the size of colon polyps under colonoscopy according to claim 2, characterized in that: During the training process of the YOLOv5 colon polyp detection model, the model divides the feature map into three blocks using a chessboard grid. Each block is associated with three vectors in the model output, and each vector represents a prediction box.

6. The method for detecting the size of colon polyps under colonoscopy according to claim 1, characterized in that: For the real image dataset, the endoscope lens intrinsic parameter matrix and lens distortion coefficient are obtained through calibration.

7. A colon polyp size detection system under colonoscopy, characterized in that: include: In the first training module, a YOLOv5 target detection model is trained using a colon polyp dataset to obtain a YOLOv5 colon polyp detection model. In the second training module, the DepthAnything depth estimation model was trained using the colonoscopy depth estimation dataset, and semi-supervised training was performed using labeled virtual images and unlabeled real images to obtain the Student model. The detection module uses the trained YOLOv5 colon polyp detection model and the trained Student model to detect the size of polyps. The second training module performs the following operations: Step 2.1: Preparation and preprocessing of colonoscopy depth estimation dataset The colonoscopy depth estimation dataset includes a labeled virtual image dataset and an unlabeled real image dataset; Preprocessing includes: if the data is a video, filter out the blurred frames of the video, split it into pictures, and use the lens intrinsic parameter matrix and lens distortion coefficient to correct the pictures with fisheye lens; Step 2.2: Train the first DepthAnything depth estimation model on the labeled virtual image dataset to obtain the Teacher model; Step 2.3: Train the DepthAnything illumination adjustment model to predict the grayscale image of the real image; Use the trained Teacher model to predict the depth map D of the image teacher , and the depth map D teacher The pixels representing the depth values ​​are restored to the point cloud in the world coordinate system: X(u,v)=D(u,v)K -1 (u,v,1) T (17) Among them, for a pixel with pixel coordinates (u, v), X(u, B) represents its world coordinates, D(u, v) represents its distance from the lens, that is, the depth value; K is the definition of the camera intrinsic parameter matrix; After obtaining the point cloud, calculate the pixel-by-pixel light and shadow PPS(X) of the image: PPS(X)=Att(X)×[LD(X)·Nor(X)] (18) in The coordinates of the co-location light of the endoscope are (0,0), The lighting direction is the negative direction of the Z axis. The angular attenuation coefficient of the illumination is 1; Use all real images to train the DepthAnything illumination adjustment model, grayscale image of the real image With its depth map D teacher The splicing result in the channel dimension is Ref in The side length of bilinear interpolation is downsampled by four times to As the model input, the model predicts a correction to the grayscale image of the real image Then upsample four times by bilinear interpolation to get The loss function is: After training, the trained DepthAnything illumination adjustment model is used to predict the grayscale image of the real image. gray , get and save the result PPS ref =Ref out I gray For subsequent model training; Step 2.4: Use the virtual image and real image datasets to train the second DepthAnything depth estimation model to obtain the Student model; When the input is a virtual image, the loss function is: When the input is a real picture, the loss function is: Among them, PPS stu Yes D pred_stu Input equations (18) and (19); The detection module performs the following operations: The preprocessed image is input into the YOLOv5 colon polyp detection model to obtain the colon polyp bounding box matrix Polyps in the image bbox ; Input the preprocessed image into the Student model to obtain the depth map D of the image student ; The depth map D student The pixels representing the depth values ​​are restored to the point cloud in the world coordinate system; For the colon polyp bounding box matrix Polyps bbox The bounding box vector of each colon polyp in the point cloud is calculated, and the midpoint coordinates of the four sides of the bounding box are taken as the coordinates PC of the midpoint coordinates in the point cloud. center ; Calculate the coordinates PC center (x min ,x max ,y min ,y max ), where x max With x min Represents the coordinates PC center The maximum and minimum values ​​in the x direction, y max With y min Respectively represent PC center Maximum and minimum values ​​in the y direction; Denote the size of the polyp as (x max -x min ,y max -y min ).

Citation Information

Patent Citations

  • Marker size measuring method, device and equipment based on electronic endoscope and medium

    CN115294191A

  • Fruit space coordinate positioning method based on monocular vision

    CN116071424A

  • Hysteroscope image target detection method based on depth information and knowledge distillation

    CN118644460A