A method and system for evaluating the size of colon polyps based on multi-frame matching

Through the colonoscopic polyp size evaluation method based on multi-frame matching, the intestinal polyp detection model and depth estimation model are used to calculate the size of polyp in the colonoscopic video stream, solving the problem of large-subjective influence and high cost of auxiliary tools in the prior art, and achieving more accurate polyp size measurement.

CN119722775BActive Publication Date: 2025-06-13SUZHOU LINGYING YUNNUO MEDICAL TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510225467.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-13
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

The prior art is subjectively affected by human subjectively in the evaluation of intestinal polyps size, and the auxiliary tools are costly and inconvenient to operate. The accuracy of the method of estimating the depth of intestinal lesions depends on the simulation of the intestinal body model, making it difficult to accurately estimate in real colonoscopy.

Method used

The colonoscopic polyp size evaluation method based on multi-frame matching is adopted. By obtaining the colonoscopic video stream, the intestinal polyp detection model is used to determine whether there are polyps in the video frame, the image similarity of subsequent video frames is calculated, the depth map is generated, and the polyp size is calculated based on the polyp position and depth information, and the output standard polyp size is weighted.

Benefits of technology

It improves the accuracy of prediction of the depth information of real colonoscopy images, eliminates the impact of polyp blurring, smooths out the polyp size output, provides more accurate polyp size measurements, and helps endoscopy doctors make better decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119722775B_ABST
    Figure CN119722775B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of image processing technology, and particularly to a method and system for evaluating the size of colon polyps based on multi-frame matching. A video stream to be evaluated of a colonoscope is obtained, and a colon polyp detection model is used to determine whether there are polyps in the video frames to be evaluated therein. When there are polyps, the position information of the polyps is obtained and the image similarity between two adjacent frames of N subsequent video frames of the video frame to be evaluated is calculated. When the image similarity meets a preset condition, a depth estimation model is used to generate a corresponding depth map. According to the position information of the polyps and the polyp depth information in the depth map, the size of the polyps in the subsequent video frames is calculated. The sizes of the polyps in each subsequent video frame are weighted and calculated, and the standard polyp size of the video stream to be evaluated of the colonoscope is output. The present invention can better predict the depth information of images in a real colonoscope, eliminate the influence of polyp blurring, smooth the display of polyp size output, measure the size of polyps more accurately, and help endoscopic doctors make better decisions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly relates to a method and system for evaluating the size of colon polyps based on multi-frame matching. Background Art

[0002] As one of the common malignant tumors globally, the early diagnosis of colorectal cancer is of crucial importance. Intestinal polyps are recognized as an important precancerous lesion. The size of intestinal polyps is not only related to the choice of treatment methods, but more importantly, has a profound impact on the formulation of subsequent intervention measures and disease management strategies.

[0003] Currently, most doctors estimate the size of polyps based on their own experience. In the patent No. CN213588351U, a biopsy forceps with a spring ruler is used to measure the size of polyps. In the patent No. CN114680809A, a laser emission lamp is integrated in the colonoscope, and the size of the polyp is measured with the laser spot as a reference. The patent No. CN111091562B calculates the size of the lesion with the forward water column or the biopsy forceps as a reference. The patent No. CN115578437B provides a method for obtaining the depth data of intestinal lesions. By simulating and modeling the intestinal model, the depth map of the real colonoscope image is estimated, so as to facilitate the subsequent estimation of the lesion size.

[0004] Currently, the method of doctors visually estimating the size of intestinal polyps depends on the experience judgment of endoscopists, which is greatly affected by human subjectivity and is prone to deviation. The existing methods measure the lesion by introducing auxiliary tools or references such as biopsy forceps and laser probes as references, but the implementation cost is high and the operation is inconvenient. The correctness of the method for estimating the depth of intestinal lesions depends on the simulation of the intestinal model for the real colonoscope scene. Since there are significant differences between the intestinal model and the real scene environment, and no real colonoscope data is added for calibration, it is difficult to accurately estimate the size of intestinal polyps in real colonoscopy examinations. Summary of the Invention

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] According to the first aspect of the present invention, the present invention claims protection for a method for evaluating the size of colon polyps based on multi-frame matching, characterized by comprising:

[0007] Obtain the video stream to be evaluated of the colonoscope, and determine whether there is a polyp in the video frame to be evaluated in the video stream to be evaluated of the colonoscope through an intestinal polyp detection model;

[0008] When there is a polyp in the video frame to be evaluated, obtain the position information of the polyp and calculate the image similarity between two adjacent frames of N subsequent video frames of the video frame to be evaluated by using the perceptual hashing algorithm;

[0009] When the image similarity between two adjacent frames of N subsequent video frames meets a preset condition, a depth estimation model is used to generate a corresponding depth map;

[0010] According to the position information of the polyp and the polyp depth information in the depth map, calculate the size of the polyp in the subsequent video frames;

[0011] Perform weighted calculation on the polyp sizes of each subsequent video frame, and output the standard polyp size of the colonoscopy video stream to be evaluated.

[0012] Further, the obtaining of the colonoscopy video stream to be evaluated, and determining whether there is a polyp in the video frame to be evaluated in the colonoscopy video stream to be evaluated through an intestinal polyp detection model, further includes:

[0013] The intestinal polyp detection model uses lower digestive tract polyp data as the training set;

[0014] For the input colonoscopy video stream to be evaluated, the intestinal polyp detection model outputs the upper left, lower right coordinates and confidence level of the position of the polyp in the video frame to be evaluated in the colonoscopy video stream to be evaluated;

[0015] When the confidence level is greater than a preset threshold, it is considered that there is a polyp in the video frame to be evaluated.

[0016] Further, when there is a polyp in the video frame to be evaluated, obtaining the position information of the polyp and calculating the image similarity between two adjacent frames of N subsequent video frames of the video frame to be evaluated by using the perceptual hash algorithm, further includes:

[0017] Perform image preprocessing on N subsequent video frames of the video frame to be evaluated, and calculate the DCT coefficient matrix of each subsequent video frame;

[0018] Extract the sub - matrix of each DCT coefficient matrix as the image low - frequency information of each subsequent video frame, and calculate the mean value of each sub - matrix;

[0019] Based on the comparison result between the image low - frequency information of each subsequent video frame and the mean value of each sub - matrix, calculate the matching hash value of each subsequent video frame;

[0020] Calculate the Hamming distance between two adjacent frames of each subsequent video frame according to the matching hash value, and determine the image similarity between two adjacent frames of each subsequent video frame based on the Hamming distance.

[0021] Further, when the image similarity between two adjacent frames of N subsequent video frames meets a preset condition, using a depth estimation model to generate a corresponding depth map, further includes:

[0022] When the Hamming distance between two adjacent frames of the N subsequent video frames is less than a preset threshold, it is determined that the image similarity between two adjacent frames of the N subsequent video frames meets the preset condition;

[0023] Train a student depth prediction model using simulated rendered images and real colonoscopy images, and generate a corresponding depth map based on the student depth prediction model.

[0024] Further, the training of the student depth prediction model using simulated rendered images and real colonoscopy images further includes:

[0025] Input the simulated rendered images into the encoder 1 and the decoder in sequence, and calculate the loss in the training of the rendered intestinal images according to the real depth labels;

[0026] Perform data augmentation on the real colonoscopy images and then input them into the encoder 1, and calculate the loss in the training of the real intestinal images according to the pseudo-depth labels;

[0027] Input the real colonoscopy images into the encoder 2 to introduce semantic segmentation features as auxiliary supervision information, and use the output features of the encoder 2 as auxiliary semantic supervision;

[0028] The pseudo-depth labels are predicted by the teacher model for the real colonoscopy images. Among them, the teacher model is only trained using rendered images with real depth labels, and the model structure of the teacher model is the same as that of the student model.

[0029] Further, according to the position information of the polyp and the polyp depth information in the depth map, calculating the polyp size of the subsequent video frames further includes:

[0030] Obtain the central coordinate depth information of the polyp according to the student depth prediction model;

[0031] Calculate the initial polyp size of the subsequent video frame according to the central coordinate depth information and the position information;

[0032] Correct the initial polyp size to obtain the final polyp size of the subsequent video frame.

[0033] Further, the weighted calculation of the polyp sizes of each subsequent video frame and the output of the standard polyp size of the colonoscopy video stream to be evaluated further includes:

[0034] Construct a circular queue based on the central depth of the polyp in each subsequent video frame within the depth of field range of the shot;

[0035] Based on the IoU value of the polyp frames between adjacent frames of each subsequent video frame and the state of the depth at the polyp center within the depth of field range of the shot, determine whether the front and rear frames are the same polyp and perform a direct enqueue or empty enqueue operation on the circular queue;

[0036] When the circular queue is full, a weighted moving average is used to output the standard polyp size of the colonoscopy video stream to be evaluated.

[0037] Further, when determining whether the front and rear frames are the same polyp based on the IoU value of the adjacent frame polyp boxes of each subsequent video frame and the depth at the polyp center within the depth of field range of the current shot, and performing a direct enqueue or cleared enqueue process on the circular queue, it further includes:

[0038] If the IoU value of the adjacent frame polyp boxes of each subsequent video frame is greater than or equal to the first threshold and the depth at the polyp center is within the depth of field range of the current shot, it is determined that the front and rear frames are the same polyp, and the current polyp size is added to the circular queue;

[0039] If the IoU value of the adjacent frame polyp boxes of each subsequent video frame is less than the first threshold and the depth at the polyp center is within the depth of field range of the current shot, it is determined that the front and rear frames are not the same polyp, then the circular queue is cleared, and the polyp size of the video frame to be evaluated is re-added to the cleared circular queue.

[0040] According to a second aspect of the present invention, the present invention claims protection for a colonoscopy polyp size evaluation system based on multi-frame matching, including:

[0041] One or more processors;

[0042] A memory storing one or more programs thereon, when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the described method for evaluating the size of colonoscopy polyps based on multi-frame matching.

[0043] The present invention relates to the technical field of image processing, and particularly relates to a method and system for evaluating the size of colonoscopy polyps based on multi-frame matching. The method includes obtaining a colonoscopy video stream to be evaluated, determining whether there is a polyp in the video frame to be evaluated through a colon polyp detection model; when there is a polyp, obtaining the position information of the polyp and calculating the image similarity between two adjacent frames of N subsequent video frames of the video frame to be evaluated; when the image similarity meets a preset condition, using a depth estimation model to generate a corresponding depth map; calculating the polyp size of the subsequent video frames according to the position information of the polyp and the polyp depth information in the depth map; performing weighted calculation on the polyp sizes of each subsequent video frame, and outputting the standard polyp size of the colonoscopy video stream to be evaluated. The present invention can better predict the depth information of images in a real colonoscopy, eliminate the influence of polyp blurring, smooth the display of the output of the polyp size, measure the size of the polyp more accurately, and help endoscopic doctors make better decisions. Description of the Drawings

[0044] Figure 1The flowchart of a method for evaluating the size of colon polyps based on multi-frame matching, which is claimed in the embodiments of the present invention;

[0045] Figure 2 The Unity simulation rendering diagram of a method for evaluating the size of colon polyps based on multi-frame matching, which is claimed in the embodiments of the present invention;

[0046] Figure 3 The schematic diagram of the model training process of a method for evaluating the size of colon polyps based on multi-frame matching, which is claimed in the embodiments of the present invention;

[0047] Figure 4 The specific structural schematic diagram of encoder 1 and decoder of a method for evaluating the size of colon polyps based on multi-frame matching, which is claimed in the embodiments of the present invention;

[0048] Figure 5 The schematic diagram of the fusion block structure and the depth estimation prediction head structure of a method for evaluating the size of colon polyps based on multi-frame matching, which is claimed in the embodiments of the present invention. Detailed implementation manners

[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0050] The terms "first", "second", and "third" in the present invention are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", and "third" may explicitly or implicitly include at least one of such features. In the description of the present invention, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined. All directional indications (such as up, down, left, right, front, back...) in the embodiments of the present invention are only used to explain the relative position relationship and movement conditions between components in a specific posture (as shown in the drawings). If the specific posture changes, the directional indications will also change accordingly. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes unlisted steps or units, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.

[0051] As used herein, the mention of "embodiment" means that the specific features, structures, or characteristics described in connection with the embodiment may be included in at least one embodiment of the present invention. The phrase appears at various positions in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.

[0052] According to a first aspect of the present invention, the present invention claims a method for evaluating the size of intestinal polyps based on multi-frame matching, with reference to Figure 1 , including:

[0053] Obtain the video stream to be evaluated of the colonoscope, and determine whether there are polyps in the video frame to be evaluated in the video stream to be evaluated of the colonoscope through an intestinal polyp detection model;

[0054] When there are polyps in the video frame to be evaluated, obtain the position information of the polyps and calculate the image similarity between adjacent two frames of N subsequent video frames of the video frame to be evaluated by using the perceptual hashing algorithm;

[0055] When the image similarity between adjacent two frames of N subsequent video frames meets a preset condition, use a depth estimation model to generate a corresponding depth map;

[0056] According to the position information of the polyps and the polyp depth information in the depth map, calculate the size of the polyps in the subsequent video frames;

[0057] Perform weighted calculation on the polyp sizes of each subsequent video frame, and output the standard polyp size of the video stream to be evaluated of the colonoscope.

[0058] Furthermore, the step of obtaining the video stream to be evaluated of the colonoscope and determining whether there are polyps in the video frame to be evaluated in the video stream to be evaluated of the colonoscope through an intestinal polyp detection model further includes:

[0059] The intestinal polyp detection model uses lower digestive tract polyp data as the training set;

[0060] For the input video stream to be evaluated of the colonoscope, the intestinal polyp detection model outputs the upper left, lower right coordinates and confidence of the polyp position in the video frame to be evaluated in the video stream to be evaluated of the colonoscope;

[0061] When the confidence is greater than a preset threshold, it is considered that there are polyps in the video frame to be evaluated.

[0062] Wherein, in this embodiment, the intestinal polyp detection model can use a single-stage detection network based on CNN such as YOLOv8, or a single-stage detection network based on transformer such as DETR.

[0063] The detection model uses lower digestive tract polyp data as the training set. For the input video stream images, the detection model outputs the upper left, lower right coordinates of the polyp positions in the corresponding images and the confidence level conf. When the confidence level is greater than the set threshold, it is considered that there is a polyp in this frame.

[0064] In a specific actual detection example, the detection model outputs the upper left, lower right coordinates as [269, 215], [509, 416] respectively, and the confidence level is 0.88. The currently set confidence threshold is 0.3. Therefore, it is considered that there is a polyp in the current image frame.

[0065] Further, when there is a polyp in the video frame to be evaluated, obtaining the position information of the polyp and calculating the image similarity between two adjacent frames of the N subsequent video frames of the video frame to be evaluated by using the perceptual hashing algorithm further includes:

[0066] Performing image preprocessing on the N subsequent video frames of the video frame to be evaluated, and calculating the DCT coefficient matrix of each subsequent video frame;

[0067] Extracting the sub-matrix of each DCT coefficient matrix as the low-frequency image information of each subsequent video frame, and calculating the mean value of each sub-matrix;

[0068] Based on the comparison results of the low-frequency image information of each subsequent video frame and the mean value of each sub-matrix, calculating the matching hash value of each subsequent video frame;

[0069] Calculating the Hamming distance between two adjacent frames of each subsequent video frame according to the matching hash value, and determining the image similarity between two adjacent frames of each subsequent video frame based on the Hamming distance.

[0070] Among them, in this embodiment, performing image preprocessing on the N subsequent video frames of the video frame to be evaluated and calculating the DCT coefficient matrix of each subsequent video frame further includes:

[0071] Image shrinking: Shrinking the image to a fixed size (such as 32x32 pixels) to remove detailed information.

[0072] Grayscale processing: Converting the shrunk image into a grayscale image;

[0073] Calculating the discrete cosine transform matrix (DCT): Using the discrete cosine transform to convert the image from the spatial domain to the frequency domain to obtain the DCT coefficient matrix.

[0074] Specifically, extracting the sub-matrix of each DCT coefficient matrix as the low-frequency image information of each subsequent video frame and calculating the mean value of each sub-matrix further includes:

[0075] Retain low-frequency coefficients: Retain the upper-left 8x8 submatrix of the DCT coefficient matrix, and these coefficients represent the low-frequency information of the image;

[0076] Calculate the mean value: Calculate the mean value of the 8x8 submatrix.

[0077] Specifically, based on the comparison results between the low-frequency image information of each subsequent video frame and the mean value of each submatrix, calculate the matching hash value of each subsequent video frame, and it further includes:

[0078] Compare the low-frequency image information of each subsequent video frame with the mean value of each submatrix. Those greater than the mean value are recorded as 1, and those less than or equal to the mean value are recorded as 0 to generate a 64-bit binary hash value.

[0079] Specifically, calculate the Hamming distance between two adjacent frames of each subsequent video frame according to the matching hash value, and determine the image similarity between two adjacent frames of each subsequent video frame based on the Hamming distance, and it further includes:

[0080] Calculate the Hamming distance between two images: The number of different bits in the 64-bit hash values of two images is the Hamming distance. The smaller the Hamming distance, the higher the similarity between the two images.

[0081] In a specific embodiment, specifically, N is taken as 5, and the hash values of the subsequent 5 frames of images are converted to hexadecimal and respectively represented as 0x9c62512251686b6d, 0x9c625122502a6a6d, 0x9c625122512a6a6d, 0x9c625122510a626c, 0x9c6251225108626c. Then the Hamming distances between adjacent frames are 4, 1, 3, 1 respectively.

[0082] Furthermore, when the image similarity between two adjacent frames of the N subsequent video frames meets the preset conditions, use a depth estimation model to generate a corresponding depth map, and it further includes:

[0083] When the Hamming distance between two adjacent frames of the N subsequent video frames is less than the preset threshold, it is determined that the image similarity between two adjacent frames of the N subsequent video frames meets the preset conditions;

[0084] Train a student depth prediction model with simulated rendered images and real colonoscopy images, and generate a corresponding depth map based on the student depth prediction model.

[0085] Among them, in this embodiment, taking the hash values of 5 images as an example, the set threshold is 5. Then the Hamming distances of 5 consecutive frames of images (4, 1, 3, 1) are all less than 5, so it is considered that the current 5 frames of images have high similarity, indicating that the current picture is in a relatively static or slightly shaking state. Therefore, the picture is relatively clear, reducing the interference of motion blur on the depth estimation model.

[0086] The depth estimation model does not only use the depth map obtained from the simulated intestinal model. In this solution, the scenario of a real colonoscope is added as an auxiliary to improve the accuracy of depth estimation.

[0087] The dataset used in this solution generates a three-dimensional model using CT colon imaging and renders images using the Unity simulation environment to obtain the effect closest to real colonoscopy examination. As Figure 2 shown are the Unity simulation rendered images and the true depth of the corresponding rendered images.

[0088] Furthermore, training the student depth prediction model using the simulated rendered images and real colonoscope images further includes:

[0089] Sequentially inputting the simulated rendered images into the encoder 1 and the decoder, and calculating the loss in the training of the rendered intestinal images based on the true depth labels;

[0090] Performing data augmentation on the real colonoscope images and then inputting them into the encoder 1, and calculating the loss in the training of the real intestinal images based on the pseudo-depth labels;

[0091] Inputting the real colonoscope images into the encoder 2 to introduce semantic segmentation features as auxiliary supervision information, and using the output features of the encoder 2 as auxiliary semantic supervision;

[0092] The pseudo-depth labels are obtained by the teacher model predicting the real colonoscope images. Among them, the teacher model is only trained using the rendered images with true depth labels, and the model structure of the teacher model is the same as that of the student model.

[0093] Among them, in this embodiment, referring to Figure 3 , the solid line part is the process of calculating the loss in the training of the rendered intestinal images with true depth labels, and the dashed line part is the process of calculating the loss in the training of the real intestinal images without true depth labels. The teacher model is trained only using the rendered intestinal images with true depth labels and predicts the real intestinal images to generate pseudo-depth labels.

[0094] Performing data augmentation operations on the real intestinal images, including CutMix, color perturbation, Gaussian blur, exposure enhancement, etc. Through the above operations, the true label loss and the pseudo-label loss can be obtained.

[0095] In order to enable the student model to capture more semantic information, semantic segmentation features are introduced as auxiliary supervision information. The output features of the encoder 2 are used as auxiliary semantic supervision. Here, the encoder 2 uses the pre-trained and frozen DINOv2 encoder.

[0096] Among them, DINOv2 learns general semantic information from a large amount of unlabeled images through self-supervised learning technology, so it is applicable to different tasks. Then its semantic feature alignment loss is defined as follows:

[0097]

[0098] where MLP is a trainable multi-layer perceptron operation, and cos is a cosine similarity operation. are the features output by encoder 1 and encoder 2 respectively, α is a threshold constant, and H and W are the height and width of

[0099] Therefore, the total loss of the student depth prediction model is , where , , correspond to the hyperparameters of the three losses respectively. Encoder 1 of the student depth prediction model also adopts the ViT-B structure, and the decoder adopts the DPT decoder structure. Finally, the student depth prediction model is used to predict the depth information in real colonoscopy images.

[0100] Refer to Figure 4 , which shows the structures of encoder 1 and the decoder. In encoder 1, non-overlapping patch embedding is adopted, and it is obtained by linearly projecting the flattened representation of non-overlapping patches.

[0101] Each Transformer block in encoder 1 consists of N repeated Transformer layers and outputs tokens. p represents the size of the original patch, and the relationship between and the length and width H and W of the original image is . In this solution, N in the encoder is 3.

[0102] In the decoding stage, it mainly consists of a recombination block and a fusion block. The recombination operation is expressed as follows:

[0103]

[0104] where represents the output feature dimension, s is the ratio of the slice dimension output by the Transformer block to the size of the input image, and t represents tokens, and its dimension is . The Read operation maps tokens to tokens, that is, . The Concatenate operation converts the dimension from to , thus converting it into an image representation similar to the original image. Finally, through operation, the dimension is changed from to . The specific operation is to change the channels through 1×1 convolution. When s≥p, 3×3 convolution is used for downsampling. When s is less than p, 3×3 transposed convolution is used for upsampling. means performing Read, Concatenate, and Resample operations on t in sequence.

[0105] The fusion block structure and the depth estimation prediction head structure are as shown in Figure 5 . In this solution, p = 16. The resolution of the output of the final fusion block 1 is half of the image, and the output of the depth estimation prediction head is the same size as the original image.

[0106] Furthermore, according to the position information of the polyp and the polyp depth information in the depth map, calculating the size of the polyp in subsequent video frames further includes:

[0107] Obtaining the depth information of the center coordinates of the polyp according to the student depth prediction model;

[0108] Calculating the initial size of the polyp in the subsequent video frames based on the center coordinate depth information and the position information;

[0109] Correcting the initial size of the polyp to obtain the final size of the polyp in the subsequent video frames.

[0110] Among them, in this embodiment, according to the upper left, lower right coordinates of the polyp rectangle frame output by the intestinal polyp detection model ;

[0111] Calculating the center image coordinates of the polyp according to the A and B coordinates , and calculating the pixel size of the polyp as , where dx and dy respectively represent the actual size of each pixel in the x and y directions (unit: mm / pixel);

[0112] Obtaining the depth at the center coordinate C of the current image polyp according to the student depth prediction model ;

[0113] Calculating the initial size of the polyp , where the current focal length of the endoscope lens;

[0114] Since is the distance from the surface of the polyp center to the endoscope lens, it is necessary to correct the initial size of the polyp to obtain the final size of the polyp in the current image , where λ is a correction constant, affected by the polyp morphology, and the value range is between 2 and 100. In this solution, it is tentatively set to 3.

[0115] Taking the polyp coordinates in step 1 as an example, the upper-left coordinate A is [269, 215], the lower-right coordinate B is [509, 416], then C is [389, 315]. Given that both dx and dy are 0.005 mm / pixel, the pixel length of the polyp is calculated to be 1.57 mm, and the depth at the coordinate C [389, 315] is obtained from the depth estimation model. is 21.6 mm, and the endoscope focal length f is 5 mm. It can be calculated that = 6.78 mm. After correction = 7.49 mm.

[0116] Furthermore, the weighted calculation of the polyp sizes of each subsequent video frame to output the standard polyp size of the video stream to be evaluated by the colonoscope further includes:

[0117] Constructing a circular queue based on the central depth of the polyp in each subsequent video frame within the depth of field range of the current lens;

[0118] When the IoU value of the adjacent frame polyp boxes of each subsequent video frame and the depth at the polyp center are within the depth of field range of the current lens, determining whether the front and rear frames are the same polyp and performing a direct enqueue or empty enqueue operation on the circular queue;

[0119] When the circular queue is full, using weighted moving average to output the standard polyp size of the video stream to be evaluated by the colonoscope.

[0120] Furthermore, when the IoU value of the adjacent frame polyp boxes of each subsequent video frame and the depth at the polyp center are within the depth of field range of the current lens, determining whether the front and rear frames are the same polyp and performing a direct enqueue or empty enqueue operation on the circular queue further includes:

[0121] If the IoU value of the adjacent frame polyp boxes of each subsequent video frame is greater than or equal to the first threshold and the depth at the polyp center is within the depth of field range of the current lens, it is determined that the front and rear frames are the same polyp, and the current polyp size is added to the circular queue;

[0122] If the IoU value of the adjacent frame polyp boxes of each subsequent video frame is less than the first threshold and the depth at the polyp center is within the depth of field range of the current lens, it is determined that the front and rear frames are not the same polyp, then the circular queue is cleared, and the polyp size of the video frame to be evaluated is re-added to the cleared circular queue.

[0123] Among them, in this embodiment, in order to eliminate the influence of large errors in depth estimation caused by polyp blurring, when the depth at the center C of the polyp in the i-th frame is within the depth of field range of the current lens inside, l low and lhigh are the minimum and maximum values of the depth of field of the lens, respectively. Add the current polyp size to a circular queue of length N ;

[0124] If the IoU of the polyp bounding boxes between adjacent frames is ≥ 0.3 and the depth at the center of the polyp is within the depth of field range of the current lens it is determined that the front and back frames are the same polyp, and add the current polyp size to a queue of length N ;

[0125] If the IoU of the polyp bounding boxes between adjacent frames is < 0.3 and the depth at the center of the polyp is within the depth of field range of the current lens it is determined that the front and back frames are not the same polyp, then clear the queue and add the polyp size of the current frame to a queue of length N ;

[0126] When the queue is full, use weighted moving average to output and display the polyp size:

[0127] .

[0128] Taking N = 5 as an example, the depth of field range of the endoscope lens is [9mm, 100mm], and the depths of the polyps in 5 consecutive frames are [16.6mm, 16.8mm, 17.3mm, 17.8mm, 18.1mm], all within the depth of field of the endoscope, and the IoU is greater than 0.3. The polyp sizes calculated through step 4 are Q = [6.71mm, 7.12mm, 7.23mm, 6.78mm, 6.88mm], then through 6.94mm.

[0129] According to the second embodiment of the present invention, the present invention claims protection for a colon polyp size evaluation system based on multi-frame matching, including:

[0130] One or more processors;

[0131] A memory storing one or more programs, which when executed by the one or more processors cause the one or more processors to implement the method for evaluating the size of colon polyps based on multi-frame matching as described above.

[0132] In several embodiments provided by the present invention, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces, indirect couplings or communication connections of devices or units, and can be in electrical, mechanical, or other forms.

[0133] In addition, each functional unit in various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units. The above is only the implementation manner of the present invention and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

[0134] The specific implementation manners of the invention have been described in detail above, but they are only examples. The present invention is not limited to the specific implementation manners described above. For those skilled in the art, any equivalent modification or substitution to the invention is also within the scope of the present invention. Therefore, all equal transformations, modifications, improvements, etc. made without departing from the spirit and principles of the present invention should be covered by the scope of the present invention.

Claims

1. A colonoscopic polyp size assessment method based on multi-frame matching, characterized in that: include: Obtain a colonoscopy video stream to be evaluated, and determine whether there is a polyp in the video frame to be evaluated in the colonoscopy video stream to be evaluated by using a colon polyp detection model; When a polyp exists in the video frame to be evaluated, obtaining location information of the polyp and using a perceptual hash algorithm to calculate image similarity between two adjacent frames of N subsequent video frames of the video frame to be evaluated; When the image similarity of two adjacent frames of N subsequent video frames meets a preset condition, a corresponding depth map is generated using a depth estimation model; Calculating the size of the polyp in subsequent video frames according to the position information of the polyp and the depth information of the polyp in the depth map; A circular queue is constructed in the depth of field of the lens according to the center depth of the polyp in each subsequent video frame; If the IoU values ​​of the polyp frames of adjacent frames of each subsequent video frame are greater than or equal to the first threshold and the depth of the center of the polyp is within the depth of field of the lens, the previous and next frames are considered to be the same polyp, and the current polyp size is added to the circular queue; If the IoU value of the polyp frame of the adjacent frames of each subsequent video frame is less than the first threshold and the depth of the polyp center is within the depth of field of the lens, it is determined that the previous and next frames are not the same polyp, then the circular queue is cleared, and the polyp size of the video frame to be evaluated is added back to the cleared circular queue; When the circular queue is full, a weighted moving average is used to output the standard polyp size of the colonoscopy video stream to be evaluated.

2. The method for evaluating the size of colonoscopic polyps based on multi-frame matching according to claim 1, characterized in that: The method of obtaining the colonoscopy video stream to be evaluated and determining whether there is a polyp in the video frame to be evaluated in the colonoscopy video stream to be evaluated by using the intestinal polyp detection model also includes: The intestinal polyp detection model uses lower digestive tract polyp data as a training set; For the input colonoscopy video stream to be evaluated, the intestinal polyp detection model outputs the upper left and lower right coordinates and confidence of the polyp position in the video frame to be evaluated in the colonoscopy video stream to be evaluated; When the confidence level is greater than a preset threshold, it is considered that polyps exist in the video frame to be evaluated.

3. The method for evaluating the size of colonoscopic polyps based on multi-frame matching according to claim 1, characterized in that: When a polyp exists in the video frame to be evaluated, obtaining the location information of the polyp and using a perceptual hash algorithm to calculate the image similarity of two adjacent frames of N subsequent video frames of the video frame to be evaluated also includes: Performing image preprocessing on N subsequent video frames of the video frame to be evaluated, and calculating and obtaining a DCT coefficient matrix of each subsequent video frame; Extracting submatrices of each DCT coefficient matrix as image low-frequency information of each subsequent video frame, and calculating the mean of each submatrix; Based on the comparison result of the image low-frequency information of each subsequent video frame and the mean value of each sub-matrix, the matching hash value of each subsequent video frame is calculated; The Hamming distance between two adjacent frames of each subsequent video frame is calculated according to the matching hash value, and the image similarity between two adjacent frames of each subsequent video frame is determined based on the Hamming distance.

4. The method for evaluating the size of colonoscopic polyps based on multi-frame matching according to claim 3, characterized in that: When the image similarity of two adjacent frames of the N subsequent video frames meets a preset condition, using the depth estimation model to generate a corresponding depth map also includes: When the Hamming distances between two adjacent frames of the N subsequent video frames are all less than a preset threshold, determining that the image similarity between two adjacent frames of the N subsequent video frames meets a preset condition; The simulated rendered images and real colonoscopy images are used to train a student depth prediction model, and a corresponding depth map is generated based on the student depth prediction model.

5. The method for evaluating the size of colonoscopic polyps based on multi-frame matching according to claim 4, characterized in that: The method of using the simulated rendering image and the real colonoscopy image to train the student depth prediction model also includes: The simulated rendered image is sequentially input into the encoder 1 and the decoder, and the loss is calculated in the rendered intestinal image training according to the real depth label; The real colonoscopy image is subjected to data enhancement processing and then input into encoder 1, and the loss is calculated in the real intestinal image training according to the pseudo depth label; The real colonoscopy image is input into encoder 2, and semantic segmentation features are introduced as auxiliary supervision information, and the output features of encoder 2 are used as auxiliary semantic supervision; The pseudo depth label is predicted by a teacher model on a real colonoscopy image, wherein the teacher model is trained only with rendered images having real depth labels, and the model structure of the teacher model is the same as that of the student model.

6. The method for evaluating the size of colonoscopic polyps based on multi-frame matching according to claim 5, characterized in that: Calculating the size of the polyp in the subsequent video frame according to the position information of the polyp and the depth information of the polyp in the depth map, further comprising: Obtaining the center coordinate depth information of the polyp according to the student depth prediction model; Calculating the initial polyp size of the subsequent video frame according to the center coordinate depth information and the position information; The initial polyp size is corrected to obtain a final polyp size of the subsequent video frame.

7. A colonoscopic polyp size assessment system based on multi-frame matching, characterized in that: include: one or more processors; A memory having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement a colonoscopic polyp size assessment method based on multi-frame matching according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • A method and system for measuring the size of digestive tract lesions

    CN111091562B

  • Measuring system for measuring polyp size under enteroscope by using light spot

    CN114680809A

  • Methods, devices, electronic equipment, and storage media for acquiring depth data of intestinal lesions

    CN115578437B

  • Biopsy forceps for measuring polyp size

    CN213588351U

  • Real-time endoscope enteroscope polyp detection system

    CN111383214A