A Method, Device, Equipment and Storage Medium for Extracting Key Frames of Colonoscopy Videos

By adopting a multi-module collaborative method in colonoscopic video, including artifact detection, ResNet50 model, similarity analysis network and iAFF-HSFYOLO network, the automation and accuracy of colonoscopic video keyframe extraction in the prior art is solved, and efficient and accurate polyp detection and keyframe extraction are achieved.

CN119399666BActive Publication Date: 2025-06-10SOUTHWEST JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411293724.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-14
Publication Date
2025-06-10
Estimated Expiration
2044-09-14

AI Technical Summary

Technical Problem

The prior art lacks automation and accuracy in colonoscopic video keyframe extraction, is susceptible to artifacts, and relies on manual settings, making it difficult to effectively identify polyps and other pathological features.

Method used

Multi-module collaborative working methods are adopted, including artifact area detection, ResNet50 model for polyp detection, similarity analysis network and iAFF-HSFYOLO network for polyp localization, and keyframes are extracted from it in combination with a comprehensive evaluation mechanism.

Benefits of technology

It significantly improves the accuracy and efficiency of polyp detection in colonoscopy videos, reduces the examination burden of doctors, can more accurately identify pathological characteristics, and has a wide range of clinical application prospects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119399666B_ABST
    Figure CN119399666B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of video key frame extraction, and specifically discloses a method, device, equipment and storage medium for extracting key frames of a colonoscopy video, including: detecting the artifact area of the colonoscopy video frame sequence; using a ResNet50 model to detect polyp frames, and sending the colonoscopy video frames containing polyps into a similarity analysis network composed of two VGG16 sub-networks with shared weights to calculate the similarity scores between video frames, and grouping according to the similarity scores; inputting the colonoscopy video frame sequence into the iAFF-HSFYOLO network for polyp localization; introducing a comprehensive evaluation mechanism based on the artifact area ratio, polyp area ratio, distance of the polyp from the center and confidence score to select key frames from each group of similar video frames, and completing the extraction of key frames. The present invention combines a polyp detection model, a ResNet architecture, a similarity analysis network and an attention mechanism, which can significantly reduce the examination burden of doctors and improve the accuracy and efficiency of detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video key frame extraction, and particularly relates to a method, device, equipment and storage medium for extracting key frames of a colonoscopy video. Background Art

[0002] The acquisition of colonoscopy videos is of great significance in the diagnosis, retrospective analysis and comprehensive examination of colorectal cancer. Endoscopists need to find pathological features such as lesions, polyps, erosions, ulcers, etc. from the complete colonoscopy videos, and professional and cautious judgments by experienced endoscopists are required to identify the location, nature and morphology of the lesions. However, traditional methods rely on the professional experience of endoscopists and are easily affected by factors such as bubbles, feces, motion blur, etc. Endoscopists may also miss detections due to interference such as fatigue and stress during long-term work. Therefore, finding a key frame extraction method to extract key information in colonoscopy examinations is very important for improving the diagnostic efficiency of doctors and reducing the diagnostic burden on doctors.

[0003] Existing methods for extracting key frames of colonoscopy videos usually only rely on manually set features, such as threshold setting, entropy-based frame quality assessment, and clustering-based key frame selection. These methods do not fully consider the basic attributes of colonoscopy videos themselves, nor do they consider the influence of artifacts (bubbles, feces, reflections, motion blur, etc.) in colonoscopy examinations on colonoscopy videos. Summary of the Invention

[0004] To solve the problems existing in the prior art, the present invention provides a method, device, equipment and storage medium for extracting key frames of a colonoscopy video, which is used for automatically extracting key frames of a colonoscopy video, including multiple modules, and each module works in cooperation to achieve efficient and accurate polyp detection and key frame extraction, and solves the problems mentioned in the above background art.

[0005] To achieve the above object, the present invention provides the following technical solution: A method for extracting key frames of a colonoscopy video, including the following steps:

[0006] S1. Detection of the artifact area in the colonoscopy video frame sequence:

[0007] S2. Detect polyp frames using the ResNet50 model, train the model using the cross-entropy loss function and the Adam optimizer, and input the colonoscopy video frame sequence retained after deletion in step S1 into the trained ResNet50 model to identify video frames containing polyps and discard frames with no information;

[0008] S3. Send the retained colonoscopy video frames containing polyps into a similarity analysis network composed of two VGG16 sub-networks with shared weights, calculate the similarity scores between video frames, and group them according to the similarity scores;

[0009] S4. Input the colonoscopy video frame sequence of each group into the iAFF-HSFYOLO network for polyp localization;

[0010] S5. Based on the polyp localization results of the colonoscopy video frames, introduce a comprehensive evaluation mechanism according to the area ratio of the artifact region, the area ratio of the polyp region, the distance of the polyp from the center, and the confidence score, and select key frames from the similar video frames of each group in the grouping results to complete key frame extraction.

[0011] Preferably, in step S1, it specifically includes:

[0012] S11. Use the EndoCV2020 dataset to construct a video frame artifact detection dataset and train a video frame artifact detection model based on the EndoNet model;

[0013] S12. Input the colonoscopy video frames to be processed into the artifact detection model EndoNet, calculate the area of the artifact detection box and its proportion in the image through the model, and delete the video frames with the area proportion of the artifact region higher than the threshold.

[0014] Preferably, in step S11, the video frame artifact detection model based on the EndoNet model uses ResNet as the feature extraction backbone, extracts multi-scale features through residual connections, and uses the Feature Pyramid Network (FPN) to fuse features at different levels; the model predicts the bounding box and class label of the artifact through the regression head and classification head respectively. During the detection process, multiple candidate boxes are generated on each feature layer, and the artifact detection boxes are selected after classification and regression, and non-maximum suppression (NMS) is used for deduplication.

[0015] Preferably, in step S3, the similarity analysis network consists of two VGG16 sub-networks with shared weights. Each sub-network is based on the VGG16 architecture, removes the average pooling layer and the classifier, and only retains the convolutional layers for feature extraction;

[0016] Extract features from two adjacent frames of the colonoscopy video containing polyps that are retained through the VGG16 sub-networks respectively, then calculate their L1 distance, and compare the feature differences through the L1 distance between the two feature vectors;

[0017] L1 = ∑|x1 i - x2 i |

[0018] where x1 i and x2 i respectively represent the elements in the feature vectors of two adjacent frames, and the obtained feature differences are processed through two fully connected layers. The first fully connected layer has 512 neurons, and the second fully connected layer outputs the similarity score;

[0019] The similarity score is converted into a probability value between 0 and 1 through the Sigmoid function. A higher probability value indicates high similarity, and the calculation formula is as follows:

[0020]

[0021] Among them, s is the original similarity score generated by the similarity analysis network for the input adjacent frames, and the result of grouping the frame sequence is obtained based on the similarity score.

[0022] Preferably, in step S4, the iAFF-HSFYOLO network is specifically: the HSFPN module is used to replace the SPPF layer in the YOLOv8 backbone network. At the same time, the iterative attention feature fusion module iAFF is introduced in the feature fusion stage;

[0023] In the HSFPN module, the input feature maps are respectively subjected to adaptive average pooling and adaptive max pooling operations to generate the global average feature and the maximum feature; these feature maps are then dimension-reduced through the 1x1 convolutional layer conv1, and after passing through the ReLU activation function ReLU, the number of channels is restored through another 1x1 convolutional layer conv2; the processed average feature and maximum feature are added and fused to generate a comprehensive feature map; the comprehensive feature map generates an attention weight map through the Sigmoid activation function; finally, the weight map is multiplied element-wise with the original input feature map to generate the final feature representation;

[0024] In the iterative attention feature fusion module iAFF, the local and global context information is combined to generate a weighted feature map; by dynamically adjusting the weights of the input features, the model can better capture and locate polyps at different scales; finally, the enhanced feature map is processed through the detection head of the model to generate a high-precision polyp localization result.

[0025] Preferably, in step S5, the calculation formula of the comprehensive evaluation mechanism is as follows:

[0026] Score = ω 1 ·(1 - A artifact ) + ω 2 ·A polyp + ω 3 ·(1 - D center ) + ω 4 ·C confidence

[0027] Among them, ω 1 、ω 2 、ω 3 and ω 4 are weight coefficients, A artifactDenote the area ratio of the artifact region, A polyp Denote the area ratio of the polyp region, D center Denote the distance of the polyp from the center, and C confidence Denote the confidence score.

[0028] On the other hand, to achieve the above object, the present invention also provides the following technical solution: A colonoscopy video key frame extraction device, including the following modules:

[0029] Artifact region detection module, detect the artifact region of the colonoscopy video frame sequence;

[0030] Polyp detection module, use the ResNet50 model to detect polyp frames, use the cross-entropy loss function and the Adam optimizer to train the model, and input the remaining colonoscopy video frame sequence after deletion in step S1 into the trained ResNet50 model to identify video frames containing polyps and discard frames with no information;

[0031] Similarity analysis module, send the remaining colonoscopy video frames containing polyps into a similarity analysis network composed of two VGG16 sub-networks with shared weights, calculate the similarity scores between video frames, and group them according to the similarity scores;

[0032] Polyp localization module, input the colonoscopy video frame sequence of each group in the grouping into the iAFF-HSFYOLO network for polyp localization;

[0033] Key frame extraction module, based on the polyp localization result of the colonoscopy video frame, introduce a comprehensive evaluation mechanism according to the area ratio of the artifact region, the area ratio of the polyp region, the distance of the polyp from the center, and the confidence score, and select key frames from the similar video frames of each group in the grouping result to complete key frame extraction.

[0034] On the other hand, to achieve the above object, the present invention also provides the following technical solution: An electronic device, the electronic device includes: a processor; and a memory for storing one or more programs;

[0035] When the one or more programs are executed by the processor, the processor is caused to execute the colonoscopy video key frame extraction method.

[0036] On the other hand, to achieve the above object, the present invention also provides the following technical solution: A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the colonoscopy video key frame extraction method is implemented.

[0037] The beneficial effects of the present invention are as follows: The method for extracting key frames of colonoscopy videos of the present invention is used for extracting key frames in automated colonoscopy videos. By combining a polyp detection model, a ResNet architecture, a similarity analysis network, and an attention mechanism, this method can significantly reduce the examination burden on doctors, improve the accuracy and efficiency of detection, and has broad clinical application prospects. Description of the Drawings

[0038] Figure 1 It is a schematic flowchart of the steps of the method for extracting key frames of colonoscopy videos in an embodiment of the present invention;

[0039] Figure 2 It is a schematic diagram of the module of the device for extracting key frames of colonoscopy videos in an embodiment of the present invention;

[0040] Figure 3 It is a schematic diagram of the structure of an electronic device in an embodiment of the present invention;

[0041] In the figure, 110 - artifact region detection module; 120 - polyp detection module; 130 - similarity analysis module; 140 - polyp localization module; 150 - key frame extraction module; 210 - processor; 220 - storage. Detailed Embodiments

[0042] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0043] Please refer to Figures 1 - 3 , the present invention provides a technical solution: A method for extracting key frames of colonoscopy videos, as Figure 1 shown, includes the following steps:

[0044] S1. Detection of the artifact region in the colonoscopy video frame sequence

[0045] Build a video frame artifact detection dataset using the EndoCV2020 dataset and train a video frame artifact detection model based on the EndoNet model. This model uses ResNet as the feature extraction backbone, extracts multi-scale features through residual connections, and uses a Feature Pyramid Network (FPN) to fuse features at different levels to enhance the accuracy of artifact detection. The model predicts the bounding boxes and class labels of artifacts through a regression head and a classification head respectively. During the detection process, multiple candidate boxes are generated on each feature layer, and after classification and regression, the artifact detection boxes are selected and duplicate removal is performed through Non-Maximum Suppression (NMS). Finally, the model calculates the area of the artifact detection box and its proportion in the image, and deletes the video frames where the proportion of the artifact area is higher than the threshold. That is, the frames with poor quality are filtered out, and the frames with high quality and clinical diagnostic value are retained.

[0046] S2. Use the ResNet50 model to detect polyp frames, and use the cross-entropy loss function and the Adam optimizer to train the model. Input the sequence of colonoscopy video frames retained after deletion in step S1 into the trained ResNet50 model to identify the video frames containing polyps and discard the frames with no information.

[0047] S3. Send the retained colonoscopy video frames containing polyps into a similarity analysis network composed of two VGG16 sub-networks with shared weights, calculate the similarity scores between the video frames, and group them according to the similarity scores.

[0048] The similarity analysis network is composed of two VGG16 sub-networks with shared weights. Each sub-network is based on the VGG16 architecture, removes the average pooling layer and the classifier, and only retains the convolutional layers for feature extraction;

[0049] Extract the features of two adjacent frames of the retained colonoscopy video frames containing polyps through the VGG16 sub-network respectively, then calculate their L1 distance, and compare the feature differences through the L1 distance between the two feature vectors;

[0050] L1 = ∑|x1 i -x2 i |

[0051] where x1 i and x2 i represent the elements in the feature vectors of two adjacent frames respectively. The obtained feature differences are processed through two fully connected layers. The first fully connected layer has 512 neurons, and the second fully connected layer outputs the similarity score;

[0052] Convert the similarity score into a probability value between 0 and 1 through the Sigmoid function. A higher probability value indicates high similarity. The calculation formula is as follows:

[0053]

[0054] Among them, s is the original similarity score generated by the similarity analysis network for the input adjacent frames, and the result of grouping the frame sequence is obtained based on the similarity score.

[0055] S4. Input the colonoscopy video frame sequence of each group in the grouping into the iAFF-HSFYOLO network for polyp localization.

[0056] In step S4, through steps S1, S2, and S3, high-quality colonoscopy video frames are obtained, and the iAFF-HSFYOLO model is used to localize polyps in the colonoscopy video. SPPF extracts spatial pyramid features through max pooling operations, but this method may lose some fine feature information. To reduce redundant features, combine high-level semantic information, and better capture the details of polyps of different sizes, the iAFF-HSFYOLO network specifically is: The HSFPN module is used to replace the SPPF layer in the YOLOv8 backbone network to improve the accuracy and sensitivity of polyp localization. At the same time, an iterative attention feature fusion module (iAFF) is introduced in the feature fusion stage. By combining local and global context information, a weighted feature map is generated to further enhance the fusion effect of multi-scale features. The HSFPN (Hierarchical Screening-Feature Pyramid Network) module optimizes the model's detection ability for polyps of different scales through channel attention mechanisms and feature fusion.

[0057] In the HSFPN module, the input feature maps are respectively subjected to Adaptive Average Pooling and Adaptive Max Pooling operations to generate global average features and maximum features; these feature maps are then dimensionally reduced through a 1x1 convolutional layer conv1, and after passing through the ReLU activation function ReLU, the number of channels is restored through another 1x1 convolutional layer conv2; the processed average feature and maximum feature are added and fused to generate a comprehensive feature map; the comprehensive feature map generates an attention weight map through the Sigmoid activation function; finally, the weight map is multiplied element-wise with the original input feature map to generate the final feature representation. This adaptive feature selection and fusion strategy enables the HSFPN module to effectively improve the performance of the model in the polyp detection task, especially when dealing with polyps of different sizes and shapes, providing more refined and accurate feature expressions, and improving the overall detection accuracy and robustness.

[0058] In the iterative attention feature fusion module iAFF, local and global context information is combined to generate weighted feature maps. By dynamically adjusting the weights of the input features, the model can better capture and locate polyps at different scales. Finally, the enhanced feature maps are processed by the detection head of the model to generate high-precision polyp localization results, significantly improving the detection performance of the model in complex scenarios.

[0059] In the specific implementation, for the input feature maps x and y, the fused feature map x is first generated by adding the principal elements. a Subsequently, the preliminary fused feature map x a is processed by the local attention module (LA) and the global attention module (GA) respectively to generate the local feature map x l and the global feature map x g . The first-round attention weight map m is generated through the following calculation 1 .

[0060] m 1 = σ(LA(x a ) + GA(x a ))

[0061] where the calculation formulas of LA and GA are:

[0062] LA(x a ) = BN(Conv 2 (SiLU(BN(Conv 1 (x a ))))

[0063] GA(x a ) = BN(Conv 2 (SiLU(BN(Conv 1 (GAP(x a ))))))

[0064] where GAP is global average pooling. The initial feature maps x and y are weighted and combined through the attention weight map m 1 to generate the intermediate fused feature map x u :

[0065] x u = x · m 1 + y · (1 - m 1 )

[0066] In the second-round fusion, the intermediate fused feature map x u passes through the local and global attention modules again to generate the second-round attention weight map m 2 . These enhanced feature maps are processed by the detection head of the model to generate high-precision polyp localization results.

[0067] S5. Based on the polyp localization results in the colonoscopy video frames, introduce a comprehensive evaluation mechanism according to the area ratio of the artifact region, the area ratio of the polyp region, the distance of the polyp from the center, and the confidence score, and select key frames from each group of similar video frames in the grouping results to complete the extraction of key frames.

[0068] The comprehensive evaluation mechanism calculates the final score of each cluster based on the following four parameters: the area ratio of the artifact region (A artifact ), the area ratio of the polyp region (A polyp ), the distance from the center (D center ), and the confidence score (C confidence ). The calculation formula is as follows:

[0069] Score = ω 1 ·(1 - A artifact ) + ω 2 ·A polyp + ω 3 ·(1 - D center ) + ω 4 ·C confidence

[0070] Among them, ω 1 , ω 2 , ω 3 , and ω 4 are weight coefficients, which can be set as ω 1 = 0.25, ω 2 = 0.35, ω 3 = 0.15, and ω 4 = 0.25, and they can be adjusted according to the needs of actual applications to optimize the selection effect of key frames; A artifact represents the area ratio of the artifact region, A polyp represents the area ratio of the polyp region, D center represents the distance of the polyp from the center, and C confidence represents the confidence score.

[0071] Based on the same inventive concept as the above method embodiment, the embodiment of the present application also provides a colonoscopy video key frame extraction device, which can implement the functions provided by the above method embodiment. As Figure 2 shown, the device includes the following modules:

[0072] Artifact region detection module 110, detecting the artifact region of the colonoscopy video frame sequence;

[0073] The polyp detection module 120 uses the ResNet50 model to detect polyp frames, trains the model using the cross-entropy loss function and the Adam optimizer, inputs the sequence of colonoscopy video frames retained after deletion in step S1 into the trained ResNet50 model, identifies the video frames containing polyps, and discards the frames with no information;

[0074] The similarity analysis module 130 sends the retained colonoscopy video frames containing polyps into a similarity analysis network composed of two VGG16 sub-networks with shared weights, calculates the similarity scores between the video frames, and groups them according to the similarity scores;

[0075] The polyp localization module 140 inputs the sequence of colonoscopy video frames in each group of the grouping into the iAFF-HSFYOLO network for polyp localization;

[0076] The key frame extraction module 150 introduces a comprehensive evaluation mechanism based on the polyp localization results of the colonoscopy video frames, the ratio of the area of the artifact region, the ratio of the area of the polyp region, the distance of the polyp from the center, and the confidence score, selects key frames from the similar video frames in each group of the grouping results, and completes the key frame extraction.

[0077] Based on the same inventive concept as the above method embodiment, the embodiment of the present application also provides an electronic device, as Figure 3 shown, the device includes: a processor 210; and a memory 220 for storing one or more programs;

[0078] When the one or more programs are executed by the processor 210, the processor is caused to execute the method for extracting key frames of colonoscopy videos.

[0079] The method for extracting key frames of colonoscopy videos specifically includes the following:

[0080] Detection of the artifact region of the colonoscopy video frame sequence:

[0081] Use the ResNet50 model to detect polyp frames, train the model using the cross-entropy loss function and the Adam optimizer, input the retained sequence of colonoscopy video frames into the trained ResNet50 model, identify the video frames containing polyps, and discard the frames with no information;

[0082] Send the retained colonoscopy video frames containing polyps into a similarity analysis network composed of two VGG16 sub-networks with shared weights, calculate the similarity scores between the video frames, and group them according to the similarity scores;

[0083] Input the sequence of colonoscopy video frames in each group of the grouping into the iAFF-HSFYOLO network for polyp localization;

[0084] Based on the polyp localization results of colonoscopy video frames, a comprehensive evaluation mechanism is introduced according to the area ratio of the artifact region, the area ratio of the polyp region, the distance of the polyp from the center, and the confidence score. Key frames are selected from each group of similar video frames in the grouping results to complete key frame extraction.

[0085] Based on the same inventive concept as the above method embodiment, an embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor 210, the colonoscopy video key frame extraction method described above is implemented.

[0086] The colonoscopy video key frame extraction method specifically includes the following:

[0087] Detection of the artifact region in the colonoscopy video frame sequence:

[0088] Use the ResNet50 model to detect polyp frames, use the cross-entropy loss function and the Adam optimizer to train the model, input the remaining colonoscopy video frame sequence into the trained ResNet50 model to identify video frames containing polyps, and discard frames with no information;

[0089] Send the remaining colonoscopy video frames containing polyps into a similarity analysis network composed of two VGG16 sub-networks with shared weights, calculate the similarity scores between video frames, and group them according to the similarity scores;

[0090] Input the colonoscopy video frame sequence of each group in the grouping into the iAFF-HSFYOLO network for polyp localization;

[0091] Based on the polyp localization results of colonoscopy video frames, a comprehensive evaluation mechanism is introduced according to the area ratio of the artifact region, the area ratio of the polyp region, the distance of the polyp from the center, and the confidence score. Key frames are selected from each group of similar video frames in the grouping results to complete key frame extraction.

[0092] In several embodiments provided by the embodiments of the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device and method embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0093] In addition, in each embodiment of the present invention, the various functional modules can be integrated together to form an independent part, or each module can exist separately, or two or more modules can be integrated to form an independent part.

[0094] If the described functions are implemented in the form of software function modules and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, an electronic device, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes. It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements that are not explicitly listed, or also includes elements inherent to this process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article, or device including the said element.

[0095] The terms used in the embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a", "said" and "the" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise.

[0096] It should be understood that the term "and / or" used herein is merely a description of an association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent: the sole existence of A, the simultaneous existence of A and B, and the sole existence of B. Additionally, the character " / " herein generally represents an "or" relationship between the associated objects before and after.

[0097] Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detected (stated condition or event)" or "in response to detecting (stated condition or event)".

[0098] The "first / second" mentioned in the embodiments is only to distinguish similar objects and does not represent a specific order for the objects. It can be understood that the "first / second" can be interchanged in a specific order or sequence when permitted. It should be understood that the objects distinguished by the "first / second" can be interchanged appropriately so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.

[0099] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A colonoscopy video key frame extraction method, characterized in that: The steps include: S1. Detection of artifact areas in colonoscopy video frame sequences, including: S11. Use the EndoCV2020 dataset to build a video frame artifact detection dataset and train a video frame artifact detection model based on the EndoNet model; S12, inputting the colonoscopy video frame to be processed into the artifact detection model EndoNet, calculating the area of ​​the artifact detection box and its proportion in the image through the model, and deleting the video frames with the proportion of the artifact area exceeding the threshold; S2, using the ResNet50 model to detect polyp frames, using the cross entropy loss function and the Adam optimizer to train the model, inputting the colonoscopy video frame sequence retained after being deleted in step S1 into the trained ResNet50 model, identifying the video frames containing polyps, and discarding frames without information; S3, sending the retained colonoscopy video frames containing polyps to a similarity analysis network composed of two VGG16 sub-networks with shared weights, calculating the similarity scores between the video frames, and grouping them according to the similarity scores; S4, inputting the colonoscopy video frame sequence of each group in the group into the iAFF-HSFYOLO network for polyp localization; the iAFF-HSFYOLO network specifically adopts the HSFPN module to replace the SPPF layer in the YOLOv8 backbone network, and at the same time, introduces the iterative attention feature fusion module iAFF in the feature fusion stage; In the HSFPN module, the input feature maps are subjected to adaptive average pooling and adaptive maximum pooling operations to generate global average features and maximum features; these feature maps are then reduced in dimension by a 1x1 convolutional layer conv1, and after passing through the ReLU activation function ReLU, they are then restored to the number of channels by another 1x1 convolutional layer conv2; the processed average features and maximum features are added and fused to generate a comprehensive feature map; the comprehensive feature map is used to generate an attention weight map through the Sigmoid activation function; finally, the weight map is dot-multiplied with the original input feature map to generate the final feature representation; In the iterative attention feature fusion module iAFF, local and global context information are combined to generate weighted feature maps; the weights of input features are dynamically adjusted to locate polyps of different scales; the enhanced feature maps are processed by the model’s detection head to generate polyp localization results; S5. Based on the polyp localization results of colonoscopy video frames, a comprehensive evaluation mechanism is introduced according to the area ratio of the artifact area, the area ratio of the polyp area, the distance of the polyp from the center and the confidence score. Key frames are selected from each group of similar video frames in the grouping results to complete the key frame extraction.

2. The colonoscopy video key frame extraction method according to claim 1, characterized in that: In step S11, the video frame artifact detection model based on the EndoNet model uses ResNet as the feature extraction skeleton, extracts multi-scale features through residual connections, and uses the feature pyramid network FPN to fuse features at different levels; the model predicts the bounding box and category label of the artifact respectively through the regression head and the classification head. The detection process generates multiple candidate boxes on each feature layer. After classification and regression, the artifact detection box is screened out and deduplication is performed through non-maximum suppression NMS.

3. The colonoscopy video key frame extraction method according to claim 1, characterized in that: In step S3, the similarity analysis network consists of two VGG16 sub-networks with shared weights. Each sub-network is based on the VGG16 architecture, with the average pooling layer and classifier removed, leaving only the convolutional layer for feature extraction. The two adjacent frames of the retained colonoscopy video containing polyps are respectively extracted with the VGG16 sub-network, and then their L1 distance is calculated. The feature difference is compared by the L1 distance between the two feature vectors. L1=∑|x1 i -x2 i | Among them, x1 i and x2 i They represent the elements in the feature vectors of two adjacent frames respectively. The obtained feature differences are processed by two levels of fully connected layers. The first fully connected layer has 512 neurons, and the second fully connected layer outputs the similarity score; The similarity score is converted into a probability value between 0 and 1 through the Sigmoid function. A higher probability value indicates a high similarity. The calculation formula is as follows: Among them, s is the original similarity score generated by the similarity analysis network for the input adjacent frames, and the result of grouping the frame sequence is obtained based on the similarity score.

4. The colonoscopy video key frame extraction method according to claim 1, characterized in that: In step S5, the calculation formula of the comprehensive evaluation mechanism is as follows: Score=ω1·(1-A artifact )+ω2·A polyp +ω3·(1-D center )+ω4·C confidence Among them, ω1, ω2, ω3 and ω4 are weight coefficients, A artifact Represents the area ratio of the artifact area, A polyp represents the polyp area ratio, D center represents the distance of the polyp from the center, and C confidence Represents the confidence score.

5. A key frame extraction device according to the colonoscopy video key frame extraction method according to any one of claims 1 to 4, characterized in that: Includes the following modules: Artifact region detection module (110), colonoscopy video frame sequence artifact region detection; A polyp detection module (120) uses a ResNet50 model to detect polyp frames, uses a cross entropy loss function and an Adam optimizer to train the model, inputs the colonoscopy video frame sequence retained after being deleted in step S1 into the trained ResNet50 model, identifies video frames containing polyps, and discards frames without information; A similarity analysis module (130) sends the retained colonoscopy video frames containing polyps to a similarity analysis network composed of two VGG16 sub-networks with shared weights, calculates similarity scores between the video frames, and groups them according to the similarity scores; A polyp localization module (140) inputs the colonoscopy video frame sequence of each group in the grouping into the iAFF-HSFYOLO network for polyp localization; The key frame extraction module (150) introduces a comprehensive evaluation mechanism based on the polyp positioning results of the colonoscopy video frame according to the area ratio of the artifact area, the area ratio of the polyp area, the distance of the polyp from the center and the confidence score, selects key frames from each group of similar video frames in the grouping results, and completes the key frame extraction.

6. An electronic device, characterized in that: The electronic device comprises: a processor (210); and a memory (220) for storing one or more programs; When the one or more programs are executed by the processor (210), the processor is caused to perform the colonoscopy video key frame extraction method according to any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by the processor (210), the method for extracting key frames from a colonoscopy video as described in any one of claims 1 to 4 is implemented.