Endoscope video detection method and device, readable medium and electronic equipment
By extracting keyframes from endoscopic videos for polyp detection and predicting detection boxes for non-keyframes, the problem of high computational cost and low accuracy in existing technologies is solved, achieving efficient and real-time polyp detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2023-09-21
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies require a large amount of computation for model detection of each frame of images during colonoscopy, resulting in wasted computing resources and compromised accuracy, and failing to effectively utilize the coherent information between video frames.
Polyp detection is performed by extracting keyframes from endoscopic videos. Keyframes are used to predict polyp detection boxes for non-keyframes. Detection boxes for non-keyframes are only confirmed when the prediction score meets a threshold, reducing the number of detections performed by the model.
It reduces computational load, improves detection efficiency and accuracy, enables real-time polyp detection, and avoids model performance loss.
Smart Images

Figure CN117115139B_ABST
Abstract
Description
Endoscopic video inspection methods, devices, readable media and electronic equipment Technical Field
[0001] This disclosure relates to the field of video inspection technology, and more specifically, to an endoscopic video inspection method, apparatus, readable medium, and electronic device. Background Technology
[0002] During a colonoscopy, doctors will remove any polyps that are detected. Therefore, if a polyp is detected, it will remain within the field of view for a considerable period of time and will not move to a distant location.
[0003] In related technologies, image-based model detection techniques are commonly used. This involves performing individual model detection on each frame of a colonoscopy video and outlining detected polyps to reduce the workload for doctors. However, this approach requires model detection for every single frame, resulting in high computational costs and many unnecessary calculations. Furthermore, to avoid further increasing the computational load, small models are typically used, but this leads to a significant loss of accuracy and introduces biases. Summary of the Invention
[0004] This summary section is provided to briefly introduce the concepts, which will be described in detail in the detailed description section below. This summary section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0005] In a first aspect, this disclosure provides an endoscopic video detection method, the method comprising:
[0006] Acquire the video of the endoscope to be tested;
[0007] Polyp detection is performed on the endoscopic video using a polyp detection model to obtain polyp detection boxes corresponding to each frame of the endoscopic video. The polyp detection model is used to detect polyps in the endoscopic video in the following manner:
[0008] Keyframes are extracted from the endoscopic video, and polyp detection is performed on the keyframes to obtain polyp detection boxes in the keyframes;
[0009] Based at least on the polyp detection boxes in the keyframes, predictive polyp detection boxes in non-keyframes of the endoscopic video (excluding the keyframes) are determined, wherein each predictive polyp detection box corresponds to a prediction score, and the prediction score is proportional to the prediction accuracy of the corresponding predictive polyp detection box.
[0010] When the prediction score corresponding to the predicted polyp detection box is greater than a preset threshold, the predicted polyp detection box is used as the target polyp detection box corresponding to the non-key frame.
[0011] Secondly, this disclosure provides an endoscopic video detection device, the device comprising:
[0012] The acquisition module is used to acquire the video of the endoscope to be inspected;
[0013] The first detection module is used to detect polyps in the endoscopic video using a polyp detection model, obtaining polyp detection boxes corresponding to each frame of the endoscopic video. The polyp detection model is used to detect polyps in the endoscopic video in the following manner:
[0014] Keyframes are extracted from the endoscopic video, and polyp detection is performed on the keyframes to obtain polyp detection boxes in the keyframes;
[0015] Based at least on the polyp detection boxes in the keyframes, predictive polyp detection boxes in non-keyframes of the endoscopic video (excluding the keyframes) are determined, wherein each predictive polyp detection box corresponds to a prediction score, and the prediction score is proportional to the prediction accuracy of the corresponding predictive polyp detection box.
[0016] When the prediction score corresponding to the predicted polyp detection box is greater than a preset threshold, the predicted polyp detection box is used as the target polyp detection box corresponding to the non-key frame.
[0017] Thirdly, this disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in any one of the first aspects.
[0018] Fourthly, this disclosure provides an electronic device, comprising:
[0019] A storage device on which computer programs are stored;
[0020] A processing device for executing the computer program in the storage device to implement the steps of the method according to any one of the first aspects.
[0021] The above technical solution first acquires the endoscopic video to be detected, and then uses a polyp detection model to detect polyps in the endoscopic video, obtaining polyp detection boxes for each frame of the endoscopic video. Specifically, the polyp detection model extracts keyframes from the endoscopic video and performs polyp detection on these keyframes, obtaining polyp detection boxes within them. Then, based at least on the polyp detection boxes in the keyframes, it determines predicted polyp detection boxes in non-keyframes of the endoscopic video. When the prediction score of a predicted polyp detection box is greater than a preset threshold, the predicted polyp box is used as the target polyp detection box for the non-keyframe, thus obtaining the polyp detection boxes for each frame of the endoscopic video. This solution eliminates the need for model detection on every single frame of the endoscopic video. It predicts polyp detection boxes for non-keyframes based on polyp detection boxes in keyframes, thus obtaining the polyp detection boxes for each frame. It is computationally efficient and has low computational cost, ensuring both model performance and accuracy while avoiding bias, enabling real-time polyp detection in polyp inspection scenarios.
[0022] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description
[0023] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings:
[0024] Figure 1 is a flowchart of an endoscopic video detection method according to an exemplary embodiment;
[0025] Figure 2 is a schematic diagram of an endoscopic video detection process according to an exemplary embodiment;
[0026] Figure 3 is a schematic diagram of a process for polyp detection in keyframes according to an exemplary embodiment;
[0027] Figure 4 is a block diagram of an endoscope video detection device according to an exemplary embodiment;
[0028] Figure 5 is a block diagram of an electronic device according to an exemplary embodiment. Detailed Implementation
[0029] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0030] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0031] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0032] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0033] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0034] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0035] In polyp detection scenarios, common approaches include image-based and video-based model detection techniques. Image-based model detection performs individual model detection on each frame of a colonoscopy video, identifying and bounding down the polyps to reduce the workload for doctors. However, this approach requires model detection for every single frame, resulting in high computational cost and unnecessary calculations. Furthermore, to avoid further increasing computational load, small models are often used, but this severely compromises accuracy and introduces bias. It also fails to utilize the continuity information between video frames, leading to inaccurate detection. Video-based model detection, on the other hand, inputs multiple frames for model detection, bounding down the detected polyps and plotting their trajectories on the video. However, this approach also requires model detection for each single frame, resulting in high computational cost and low efficiency.
[0036] In view of this, the present disclosure provides an endoscopic video detection method, apparatus, readable medium, and electronic device to solve the above-mentioned technical problems.
[0037] The embodiments of this disclosure will be further explained and described below with reference to the accompanying drawings.
[0038] Figure 1 is a flowchart illustrating an endoscopic video detection method according to an exemplary embodiment of the present disclosure. Referring to Figure 1, the method may include the following steps:
[0039] S101. Obtain the endoscope video to be tested.
[0040] S102. Perform polyp detection on the endoscopic video using a polyp detection model to obtain the polyp detection box corresponding to each frame of the endoscopic video.
[0041] The polyp detection model is used to detect polyps in endoscopic videos as follows: keyframes are extracted from the endoscopic video, and polyp detection is performed on the keyframes to obtain polyp detection boxes in the keyframes; based at least on the polyp detection boxes in the keyframes, predicted polyp detection boxes are determined in non-keyframes of the endoscopic video, excluding the keyframes, wherein each predicted polyp detection box corresponds to a prediction score, and the prediction score is proportional to the prediction accuracy of the corresponding predicted polyp detection box; when the prediction score corresponding to the predicted polyp detection box is greater than a preset threshold, the predicted polyp box is used as the target polyp detection box corresponding to the non-keyframe.
[0042] The above technical solution eliminates the need for model detection on every frame of the endoscopic video. It can predict the polyp detection boxes in non-key frames based on the polyp detection boxes in key frames, thereby obtaining the polyp detection boxes corresponding to each frame. This approach is computationally efficient and does not require limiting the model size. It can ensure model performance and accuracy while avoiding deviations, thus enabling real-time polyp detection in polyp inspection scenarios.
[0043] In some possible approaches, the method may further include: when the prediction score corresponding to the predicted polyp detection box is less than or equal to a preset threshold, performing polyp detection on non-key frames to obtain target polyp detection boxes in the non-key frames.
[0044] For example, referring to Figure 2, polyp detection can be performed on keyframes in the endoscopic video via branch 1. Keyframes are those frames in the video where significant image changes occur due to large camera movement. First, keyframes are extracted from the endoscopic video. Then, polyp detection is performed on these keyframes to obtain polyp detection boxes. This allows for the determination of predicted polyp detection boxes in non-keyframes of the endoscopic video, i.e., trajectory prediction. This yields the predicted trajectory of the polyp detection boxes in the endoscopic video, representing the positional change of the polyp detection boxes in each frame over time. When the prediction score of the predicted polyp detection box is greater than a preset threshold (thr), the preset output condition is met, indicating that the prediction accuracy of the predicted polyp detection box meets the requirements. The predicted polyp detection box is then output and used as the target polyp detection box for the non-keyframe. The polyp detection boxes corresponding to keyframes and the target polyp detection boxes corresponding to non-keyframes constitute the polyp detection box for each frame of the endoscopic video.
[0045] It should be understood that in the polyp detection scenario, the location of the polyp does not change over long distances in the endoscopic video, and the frames in the video are relatively close together, resulting in extremely high image similarity between frames. Therefore, a lightweight polyp detection model can be used to predict the location of the polyp detection box in each frame of the entire video using keyframes. The number of keyframes is very small; for example, a 100-frame video may only have 3 keyframes, which can greatly reduce the amount of computation.
[0046] However, if the predicted score corresponding to the predicted polyp detection box is less than or equal to a preset threshold, i.e., the preset output condition is not met, it indicates that the difference between the two frames is large, and the polyp position may have changed over a long distance. In this case, using keyframes for trajectory prediction will fail. Therefore, branch 2 can be used to perform polyp detection on non-keyframes to obtain the target polyp detection box in the non-keyframes. However, this situation rarely occurs in polyp detection scenarios. Therefore, in most cases, the target polyp detection box corresponding to the non-keyframe can be predicted using keyframes, resulting in low overall computational load, short processing time, and high efficiency.
[0047] In one possible approach, polyp detection on keyframes to obtain polyp detection boxes in the keyframes can include: extracting features from the keyframes using a feature extraction module in the polyp detection model to obtain keyframe features, and then detecting polyps on the keyframe features using a detection module in the polyp detection model to obtain polyp detection boxes in the keyframes. Similarly, polyp detection on non-keyframes to obtain target polyp detection boxes in the non-keyframes can include: extracting features from the non-keyframes using a feature extraction module to obtain non-keyframe features, and then detecting polyps on the non-keyframe features using a detection module to obtain target polyp detection boxes in the non-keyframes.
[0048] For example, continuing with Figure 2, polyp detection is performed on non-keyframes via branch 2. First, the feature extractor module extracts features from the non-keyframes to obtain non-keyframe features. Then, the detector module performs polyp detection on these non-keyframe features to obtain target polyp detection boxes in the non-keyframes. The detector module typically has two branches: classification and regression. Regression is used to improve the accuracy of the output polyp detection boxes, while classification is used to output the detected object category.
[0049] For example, Figure 3 is a schematic flowchart of polyp detection on keyframes provided in an embodiment of this disclosure, which also includes a feature extraction module and a detection module. That is, the process of polyp detection on keyframes in branch 1 is the same as the process of polyp detection on non-keyframes in branch 2. Therefore, branches 1 and 2 can reuse the same feature extraction module and detection module to improve the reusability of modules in the polyp detection model and reduce the memory resources occupied by the polyp detection model.
[0050] Of course, a separate feature extraction module and detection module can be designed for polyp detection in keyframes, and another feature extraction module and detection module can be designed for polyp detection in non-keyframes. This disclosure does not impose any restrictions on this.
[0051] In one possible approach, determining the predicted polyp detection boxes in non-key frames (excluding key frames) of the endoscopic video based at least on the polyp detection boxes in key frames may include: encoding the coordinates of the polyp detection boxes in the key frames according to the temporal information carried by the key frames to obtain the coordinate features corresponding to the polyp detection boxes; extracting the original image features corresponding to the polyp detection boxes in the key frames; performing feature fusion based at least on the original image features and coordinate features to obtain the target fused features; and inputting the target fused features into the multilayer perceptron module in the polyp detection model to obtain the predicted polyp detection boxes in non-key frames (excluding key frames) of the endoscopic video.
[0052] For example, continuing to refer to Figure 3, keyframes [N,H,W] represent N keyframes in the endoscopic video, H represents the image height, and W represents the image width.
[0053] For example, continuing with Figure 3, the feature extraction module extracts features from the input keyframe, and then the detection module detects these features, outputting the coordinates of the polyp detection box. The box position encoding module then encodes the coordinates of the polyp detection box, thus encoding the temporal location information carried by the keyframe into the polyp detection box's coordinates, rather than simply the spatial location information. By introducing temporal information, subsequent trajectory prediction becomes more accurate.
[0054] Then, the original image features (box features) corresponding to the polyp detection boxes in the keyframes are extracted. The original image features and coordinate features are then fused (concatenated) to obtain the target fused features. The target fused features are then input into the multilayer perceptron module (two MLP modules) in the polyp detection model to obtain the predicted polyp detection boxes in the non-keyframes of the endoscopic video, excluding the keyframes.
[0055] It should be understood that in polyp detection scenarios, endoscopic videos are acquired in real time, meaning the number of image frames increases in real time, and correspondingly, the number of keyframes also increases in real time.
[0056] Therefore, in a possible manner, the method may further include: acquiring historical polyp detection boxes in historical keyframes during historical detection. Determining predicted polyp detection boxes in non-keyframes of the endoscopic video, excluding the keyframes, based at least on the polyp detection boxes in the keyframes, may include: determining predicted polyp detection boxes in non-keyframes of the endoscopic video, excluding the keyframes, based on real-time polyp detection boxes and historical polyp detection boxes.
[0057] For example, taking the acquisition of a new keyframe (the current keyframe) as an example, the feature extraction module extracts features from the input current keyframe, and then the detection module detects the features of the current keyframe, outputting the coordinates of the polyp detection box in the current keyframe, plus the coordinates of the polyp detection boxes in previously detected keyframes. This can be used... This represents the coordinate position of the polyp detection box in the t-th keyframe of the endoscopic video, where 4 indicates that the coordinate position of the polyp detection box is represented by the coordinates of two vertices, such as the coordinates of the upper left and lower right vertices of the polyp detection box.
[0058] For example, the coordinates of the polyp detection box can be encoded using the following encoding method:
[0059]
[0060] Among them, PE(B t This indicates that the keyframe t is positionally encoded, and the polyp detection box after adding time information can be used... It means that d m This represents the dimensional parameters, specifically the five dimensions of time and coordinate location. This combined spatial and temporal feature is then fed into a Multi-Layer Perceptron (MLP) for further feature transformation to obtain coordinate features, which can be represented using C. b This means that in practical applications, polyp detection can be performed based on real-time endoscopic video, offering high real-time performance.
[0061] In possible ways, feature fusion is performed based on at least the original image features and coordinate features to obtain the target fused features, including: enlarging the polyp detection box in the keyframe by a preset factor to obtain an enlarged detection box, and extracting the enlarged image features corresponding to the enlarged detection box in the keyframe; fusing the original image features and the enlarged image features to obtain intermediate fused features; and fusing the intermediate fused features and coordinate features to obtain the target fused features.
[0062] For example, continuing with parameter Figure 3, the polyp detection box in the current keyframe can be enlarged by a preset factor to obtain an enlarged detection box, and the enlarged image features corresponding to the enlarged detection box in the current keyframe can be extracted. The preset factor can be determined according to requirements, and this disclosure does not impose any restrictions on it. Then, the original image features and the enlarged image features are fused and fed into a multilayer perceptron for further feature transformation to obtain intermediate fused features, which can be used with I... b This means that not only are the image features of the polyp detection box present, but also the image features of its surrounding area, making subsequent trajectory prediction more accurate.
[0063] Furthermore, C b and I b Feature fusion is performed, and after passing through two MLP modules, predicted polyp detection boxes are obtained in non-key frames of the endoscopic video, excluding key frames. Finally, the corresponding detection box trajectory (Boxtrajectory) of the endoscopic video is obtained, which can be used as a B... t ,t∈[1,2,…,T] represents the polyp detection box of the t-th frame image, and T represents the total number of frames in the endoscopic video.
[0064] In one possible approach, the training process of the polyp detection model may include: acquiring sample endoscopic videos, each frame of which is labeled with a sample polyp detection box; inputting the sample endoscopic videos into the polyp detection model to extract sample keyframes from the sample endoscopic videos and perform polyp detection on the sample keyframes to obtain a first polyp detection box corresponding to the sample keyframe; determining a second polyp detection box in the sample non-keyframes of the sample endoscopic videos, excluding the sample keyframes, based on the first polyp detection box; calculating a loss function value based on the first polyp detection box and the sample polyp detection box corresponding to the sample keyframe, as well as the second polyp detection box and the sample polyp detection box corresponding to the sample non-keyframe; and adjusting the parameters of the polyp detection model based on the loss function value.
[0065] For example, the sample endoscopic videos can be obtained from open-source data, as long as the videos are continuous; this disclosure does not impose any restrictions on this. Furthermore, the sample endoscopic videos can be divided into three categories: a training set, a validation set, and a test set, for use in model training, validation, and testing, respectively.
[0066] For example, the model training environment can be configured, such as selecting a suitable operating system, CPU (Central Processing Unit), memory size, CPU (Graphics Processing Unit), and graphics card computing platform, etc. Specific configurations can be made according to requirements, and this disclosure does not impose any limitations on this. Model training parameters can also be configured, for example, the following model training parameters can be configured:
[0067] Epochs (maximum number of training iterations): 30;
[0068] Batch size (number of samples in a single training session): 256;
[0069] Initial learning rate: 5e-4;
[0070] Learning rate decay: Decrease every 10 epochs with a decay factor of 0.1;
[0071] Weight decay: 1e-5;
[0072] Gradient clipping: 1.0;
[0073] Input video resolution: 224×224;
[0074] Optimization algorithm: Adam.
[0075] It should be understood that the training parameters of the above model can be adjusted according to needs, and this disclosure does not impose any restrictions on this.
[0076] By adopting the above technical solution, the trajectory of the polyp detection box can be predicted based on key frames in the endoscopic video through the polyp detection model, which greatly reduces the amount of model calculation, runs faster, facilitates real-time operation of the solution, and helps doctors determine the location of polyps in actual polyp examination scenarios.
[0077] Based on the same inventive concept, this disclosure provides an endoscopic video detection device. Referring to FIG4, the device 400 includes:
[0078] The acquisition module 401 is used to acquire the endoscope video to be tested;
[0079] The first detection module 402 is used to detect polyps in the endoscopic video using a polyp detection model, and to obtain a polyp detection box corresponding to each frame of the endoscopic video. The polyp detection model is used to detect polyps in the endoscopic video in the following manner:
[0080] Keyframes are extracted from the endoscopic video, and polyp detection is performed on the keyframes to obtain polyp detection boxes in the keyframes;
[0081] Based at least on the polyp detection boxes in the keyframes, predictive polyp detection boxes in non-keyframes of the endoscopic video (excluding the keyframes) are determined, wherein each predictive polyp detection box corresponds to a prediction score, and the prediction score is proportional to the prediction accuracy of the corresponding predictive polyp detection box.
[0082] When the prediction score corresponding to the predicted polyp detection box is greater than a preset threshold, the predicted polyp detection box is used as the target polyp detection box corresponding to the non-key frame.
[0083] With the above-mentioned device, it is not necessary to perform model detection on every frame of the endoscopic video. The polyp detection box in non-key frames can be predicted based on the polyp detection box in the key frame, so that the polyp detection box corresponding to each frame can be obtained. The computation is small and the efficiency is high. Furthermore, there is no need to limit the size of the model. The model performance can be guaranteed while ensuring the model accuracy and avoiding deviation. Thus, real-time polyp detection can be performed in polyp inspection scenarios.
[0084] Optionally, the device 400 further includes:
[0085] The second detection module is used to perform polyp detection on the non-key frame when the prediction score corresponding to the predicted polyp detection box is less than or equal to the preset threshold, so as to obtain the target polyp detection box in the non-key frame.
[0086] Optionally, the first detection module 402 is used for:
[0087] The keyframe is feature extracted by the feature extraction module in the polyp detection model to obtain keyframe features, and polyps are detected by the detection module in the polyp detection model to obtain polyp detection boxes in the keyframe.
[0088] The second detection module is used for:
[0089] The feature extraction module extracts features from the non-key frames to obtain non-key frame features, and the detection module performs polyp detection on the non-key frame features to obtain target polyp detection boxes in the non-key frames.
[0090] Optionally, the first detection module 402 is used for:
[0091] Based on the time information carried by the keyframe, the coordinates of the polyp detection box in the keyframe are positionally encoded to obtain the coordinate features corresponding to the polyp detection box.
[0092] Extract the original image features corresponding to the polyp detection box in the keyframe;
[0093] At least based on the original image features and the coordinate features, feature fusion is performed to obtain the target fused features;
[0094] The target fusion features are input into the multilayer perceptron module of the polyp detection model to obtain the predicted polyp detection boxes in the non-key frames of the endoscopic video, excluding the key frames.
[0095] Optionally, the first detection module 402 is used for:
[0096] The polyp detection box in the key frame is enlarged by a preset factor to obtain an enlarged detection box, and the enlarged image features corresponding to the enlarged detection box in the key frame are extracted.
[0097] The original image features and the enlarged image features are fused to obtain intermediate fused features;
[0098] The intermediate fusion feature and the coordinate feature are fused together to obtain the target fusion feature.
[0099] Optionally, the device 400 further includes an acquisition submodule, the acquisition submodule being used for:
[0100] Obtain historical polyp detection bounding boxes from historical keyframes during the historical detection process;
[0101] Optionally, the first detection module 402 is used for:
[0102] Based on the real-time polyp detection bounding box and the historical polyp detection bounding box, the predicted polyp detection bounding box in the non-key frames of the endoscopic video, excluding the key frames, is determined.
[0103] Optionally, the training process of the polyp detection model includes:
[0104] Acquire endoscopic video of the sample, wherein each frame of the endoscopic video of the sample is marked with a sample polyp detection box;
[0105] The sample endoscopy video is input into the polyp detection model to extract sample keyframes from the sample endoscopy video through the polyp detection model, and polyp detection is performed on the sample keyframes to obtain the first polyp detection box corresponding to the sample keyframe. Based on the first polyp detection box, the second polyp detection box is determined in the sample non-keyframes in the sample endoscopy video other than the sample keyframe.
[0106] The loss function value is calculated based on the first polyp detection box and sample polyp detection box corresponding to the sample keyframe, and the second polyp detection box and sample polyp detection box corresponding to the sample non-keyframe;
[0107] The parameters of the polyp detection model are adjusted based on the loss function value.
[0108] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0109] Based on the same concept, embodiments of this disclosure also provide a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the above-described endoscopic video detection method.
[0110] Based on the same concept, embodiments of this disclosure also provide an electronic device, including:
[0111] A storage device on which computer programs are stored;
[0112] A processing device is used to execute the computer program in the storage device to implement the steps of the above-described endoscopic video detection method.
[0113] Referring now to FIG5, a schematic diagram of the structure of an electronic device 500 suitable for implementing embodiments of the present disclosure is shown. The electronic device in embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in FIG5 is merely an example and should not impose any limitation on the functionality and scope of use of embodiments of the present disclosure.
[0114] As shown in Figure 5, the electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0115] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 shows an electronic device 500 with various devices, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0116] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.
[0117] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0118] In some implementations, communication can be conducted using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol), and can be interconnected with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0119] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0120] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire an endoscopic video to be detected; perform polyp detection on the endoscopic video using a polyp detection model to obtain polyp detection boxes corresponding to each frame of the endoscopic video, wherein the polyp detection model is used to perform polyp detection on the endoscopic video in the following manner: extract keyframes from the endoscopic video and perform polyp detection on the keyframes to obtain polyp detection boxes in the keyframes; determine predicted polyp detection boxes in non-keyframes of the endoscopic video other than the keyframes, at least based on the polyp detection boxes in the keyframes, wherein the predicted polyp detection boxes correspond to a prediction score, and the prediction score is proportional to the prediction accuracy of the corresponding predicted polyp detection box; when the prediction score corresponding to the predicted polyp detection box is greater than a preset threshold, the predicted polyp box is used as the target polyp detection box corresponding to the non-keyframe.
[0121] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0122] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0123] The modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules are not, in some cases, intended to limit the functionality of the module itself.
[0124] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0125] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0126] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0127] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0128] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.
Claims
1. An endoscopic video detection method, characterized in that, The method includes: acquiring an endoscopic video to be detected; performing polyp detection on the endoscopic video using a polyp detection model to obtain polyp detection boxes corresponding to each frame of the endoscopic video, wherein the polyp detection model is used to perform polyp detection on the endoscopic video in the following manner: extracting keyframes from the endoscopic video and performing polyp detection on the keyframes to obtain polyp detection boxes in the keyframes; encoding the coordinates of the polyp detection boxes in the keyframes according to the time information carried by the keyframes to obtain the coordinate features corresponding to the polyp detection boxes; and extracting the polyp detection data from the keyframes. The original image features corresponding to the bounding box are used; feature fusion is performed based on at least the original image features and the coordinate features to obtain target fused features; the target fused features are input into the multilayer perceptron module in the polyp detection model to obtain predicted polyp detection boxes in non-key frames of the endoscopic video, excluding the key frames, wherein the predicted polyp detection box corresponds to a prediction score, and the prediction score is proportional to the prediction accuracy of the corresponding predicted polyp detection box; when the prediction score corresponding to the predicted polyp detection box is greater than a preset threshold, the predicted polyp detection box is used as the target polyp detection box corresponding to the non-key frame.
2. The method according to claim 1, characterized in that, The method further includes: when the prediction score corresponding to the predicted polyp detection box is less than or equal to the preset threshold, performing polyp detection on the non-key frame to obtain the target polyp detection box in the non-key frame.
3. The method according to claim 2, characterized in that, The step of detecting polyps in the keyframes to obtain polyp detection boxes in the keyframes includes: extracting features from the keyframes using the feature extraction module in the polyp detection model to obtain keyframe features, and then detecting polyps in the keyframe features using the detection module in the polyp detection model to obtain polyp detection boxes in the keyframes; the step of detecting polyps in the non-keyframes to obtain target polyp detection boxes in the non-keyframes includes: extracting features from the non-keyframes using the feature extraction module to obtain non-keyframe features, and then detecting polyps in the non-keyframe features using the detection module to obtain target polyp detection boxes in the non-keyframes.
4. The method according to any one of claims 1-3, characterized in that, The step of fusing features based at least on the original image features and the coordinate features to obtain target fused features includes: enlarging the polyp detection box in the keyframe by a preset factor to obtain an enlarged detection box, and extracting the enlarged image features corresponding to the enlarged detection box in the keyframe; fusing the original image features and the enlarged image features to obtain intermediate fused features; and fusing the intermediate fused features and the coordinate features to obtain target fused features.
5. The method according to any one of claims 1-3, characterized in that, The method further includes: acquiring historical polyp detection boxes in historical keyframes during historical detection; determining predicted polyp detection boxes in non-keyframes other than the keyframes in the endoscopic video based at least on the polyp detection boxes in the keyframes includes: determining predicted polyp detection boxes in non-keyframes other than the keyframes in the endoscopic video based on real-time polyp detection boxes and the historical polyp detection boxes.
6. The method according to any one of claims 1-3, characterized in that, The training process of the polyp detection model includes: acquiring sample endoscopic videos, each frame of the sample endoscopic video being labeled with a sample polyp detection box; inputting the sample endoscopic videos into the polyp detection model to extract sample keyframes from the sample endoscopic videos using the polyp detection model, and performing polyp detection on the sample keyframes to obtain a first polyp detection box corresponding to the sample keyframe; determining a second polyp detection box in the sample non-keyframes of the sample endoscopic videos, excluding the sample keyframes, based on the first polyp detection box and the sample polyp detection box corresponding to the sample keyframes, and the second polyp detection box and the sample polyp detection box corresponding to the sample non-keyframes; and adjusting the parameters of the polyp detection model based on the loss function value.
7. An endoscopic video detection device, characterized in that, The device includes: an acquisition module for acquiring an endoscopic video to be detected; and a first detection module for performing polyp detection on the endoscopic video using a polyp detection model to obtain polyp detection boxes corresponding to each frame of the endoscopic video, wherein the polyp detection model performs polyp detection on the endoscopic video in the following manner: extracting keyframes from the endoscopic video and performing polyp detection on the keyframes to obtain polyp detection boxes in the keyframes; encoding the coordinates of the polyp detection boxes in the keyframes according to the time information carried by the keyframes to obtain the coordinate features corresponding to the polyp detection boxes; and extracting the keyframes... The original image features corresponding to the polyp detection box in the frame; feature fusion is performed based on at least the original image features and the coordinate features to obtain target fused features; the target fused features are input into the multilayer perceptron module in the polyp detection model to obtain predicted polyp detection boxes in non-key frames of the endoscopic video, excluding the key frames, wherein each predicted polyp detection box corresponds to a prediction score, and the prediction score is proportional to the prediction accuracy of the corresponding predicted polyp detection box; when the prediction score corresponding to the predicted polyp detection box is greater than a preset threshold, the predicted polyp detection box is used as the target polyp detection box corresponding to the non-key frame.
8. A computer-readable medium having a computer program stored thereon, characterized in that, When executed by the processing device, the program implements the steps of the method according to any one of claims 1-6.
9. An electronic device, characterized in that, include: A storage device having a computer program stored thereon; a processing device for executing the computer program in the storage device to implement the steps of the method according to any one of claims 1-6.
Citation Information
Patent Citations
Video object acceleration detection method and device, server and storage medium
CN110427800A
Training method of polyp detection model, polyp detection method and related device
CN116228715A