An artificial intelligence-based colonoscopy observation integrity evaluation method and system
By employing an AI-based method for assessing the integrity of colonoscopy observations, and utilizing multi-scale feature extraction and spatiotemporal attention mechanisms, this method addresses the problem of accurately quantifying the integrity of colonoscopy observations in existing technologies. It enables objective and quantifiable quality assessment of colonoscopy withdrawal videos, thereby improving assessment accuracy and clinical relevance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-10
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies cannot objectively, in real time, and accurately quantify the effective mucosal observation coverage during colonoscopy withdrawal, lack direct assessment of the integrity of colonoscopy observation, and the time required for manual statistics and pathology feedback is long, making it impossible to provide an immediate, objective, and quantifiable quality score for a single examination.
An AI-based method for assessing the integrity of colonoscopy observations was adopted. By acquiring video frame sequences of the colonoscope withdrawal process, multi-scale feature extraction was performed using a shared feature extraction backbone network trained by hierarchical vision. Combined with a spatiotemporal attention mechanism, high-resolution enhanced features were generated. A multi-task head was designed to perform clear mucosal region segmentation, fold region assessment, and intestinal segment location identification, and finally, quantitative scoring was achieved.
This method enables comprehensive quality assessment of colonoscopy withdrawal videos, improves the robustness and consistency of feature expression, significantly enhances assessment accuracy and clinical relevance, and avoids missed adenomas due to insufficient observation in traditional methods.
Smart Images

Figure CN121686167B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to an artificial intelligence-based method and system for assessing the integrity of colonoscopy observation. Background Technology
[0002] In existing technologies, quality control of colonoscopy mainly relies on manual post-procedure quality control, semi-automatic endoscope withdrawal time monitoring, and artificial intelligence-assisted real-time quality control.
[0003] Post-operative quality control is achieved through manual statistical methods using the endoscopy reporting system and pathology system. After the examination, endoscopists manually record information such as cecal arrival, bowel preparation score, and withdrawal time; mandatory uploading of cecal / ileal terminal photographs as evidence of insertion; and statistical calculation of indicators such as adenoma detection rate (ADR) and polyp detection rate (PDR) for each doctor based on pathology reports. The quality control department manually compiles and provides feedback quarterly or semi-annually.
[0004] Semi-automatic withdrawal time monitoring automatically records the time from the "cecal arrival mark" to the "withdrawal from the anus" through the endoscope host. It can only record the total time and cannot determine whether the doctor is really "observing slowly". Even with quick in and out, it can still take up to 6 minutes.
[0005] AI-assisted real-time quality control mainly uses automated methods such as recognizing the cecum to calculate the withdrawal time, monitoring the withdrawal speed in real time, scoring the intestinal cleanliness in a single frame, and providing real-time detection prompts for polyps to ensure quality control.
[0006] Current technologies cannot accurately and objectively quantify the quality of endoscopy withdrawal observation, lack a direct assessment of "effective mucosal coverage," and existing AI-based endoscopy withdrawal speed monitoring cannot avoid the problem of quickly traversing most of the mucosa. Manual statistics and pathology feedback take a long time, making it impossible to provide an immediate, objective, and quantifiable quality score for a single examination. Summary of the Invention
[0007] In view of the above problems, the present invention is proposed to provide an artificial intelligence-based method and system for assessing the integrity of colonoscopy observation to overcome the above problems.
[0008] This invention provides an artificial intelligence-based method for assessing the integrity of colonoscopy observations, the method comprising:
[0009] Acquire a video frame sequence of the colonoscopy withdrawal process, and divide the image frames in the video frame sequence into pixel blocks;
[0010] A shared feature extraction backbone network based on hierarchical visual training is used to extract features at multiple scales for each pixel block of the image frame, resulting in local texture semantic feature maps, global structure semantic feature maps, and global context semantic feature maps with the resolution decreasing sequentially and the number of channels increasing sequentially.
[0011] Semantic feature maps of the same scale from a video frame sequence are stacked temporally. Spatial and temporal features are extracted from the stacked features of semantic feature maps at each scale. The spatial and temporal features of the semantic feature maps at each scale are multiplied pixel-by-pixel with the corresponding semantic feature maps to obtain enhanced feature maps containing spatiotemporal information, including local texture semantic feature maps, global structure semantic feature maps, and global context semantic feature maps. The enhanced feature map of the global context semantic feature map is upsampled and added pixel-by-pixel to the enhanced feature map of the global structure semantic feature map to obtain a preliminary fusion feature map. The preliminary fusion feature map is then upsampled and added pixel-by-pixel to the enhanced feature map of the local texture semantic feature map to obtain the final fusion feature map.
[0012] The fused feature map is input into the pre-trained multi-task head, which outputs the clear mucosal region mask, the wrinkled region mask, the clarity classification result, and the intestinal segment location recognition result for each image frame.
[0013] The integrity of colonoscopy observation is quantitatively scored based on the clear mucosal area mask, folded area mask, clarity classification results, and intestinal segment location identification results for each image frame.
[0014] Another aspect of the present invention provides an artificial intelligence-based system for assessing the integrity of colonoscopy observation, the system comprising:
[0015] The data preprocessing module is used to acquire the video frame sequence of the colonoscopy withdrawal process and divide the image frames in the video frame sequence into pixel blocks;
[0016] The semantic feature extraction module is used to perform multi-scale feature extraction on each pixel block of the image frame using a shared feature extraction backbone network based on hierarchical vision training, to obtain local texture semantic feature maps, global structure semantic feature maps and global context semantic feature maps with the resolution of the image frame decreasing sequentially and the number of channels increasing sequentially.
[0017] The spatiotemporal attention module is used to stack semantic feature maps of the same scale in the temporal dimension of a video frame sequence. Spatial and temporal features are extracted from the stacked features of semantic feature maps at each scale. The spatial and temporal features of the semantic feature maps at each scale are multiplied pixel by pixel with the corresponding semantic feature maps to obtain enhanced feature maps containing spatiotemporal information, including local texture semantic feature maps, global structure semantic feature maps, and global context semantic feature maps. The enhanced feature map of the global context semantic feature map is upsampled and added pixel by pixel with the enhanced feature map of the global structure semantic feature map to obtain a preliminary fusion feature map. The preliminary fusion feature map is then upsampled and added pixel by pixel with the enhanced feature map of the local texture semantic feature map to obtain a final fusion feature map.
[0018] The feature recognition module is used to input the fused feature map into the pre-trained multi-task head and output the clear mucosal region mask, the wrinkled region mask, the clarity classification result, and the intestinal segment location recognition result for each image frame.
[0019] The scoring module is used to quantify the integrity of colonoscopy observation based on the clear mucosal area mask, wrinkled area mask, clarity classification results, and intestinal segment location identification results of each image frame.
[0020] The present invention provides an AI-based method and system for assessing the integrity of colonoscopy observation. By combining the temporal stacking of multi-scale semantic features and the spatiotemporal attention mechanism, it achieves effective modeling of dynamic interference in videos (such as camera shake and rapid retraction) and the unfolding process of folds. High-resolution enhanced features are generated through multi-scale fusion and used as shared inputs for multiple task heads, improving the robustness and consistency of feature representation. A dedicated task head is designed to achieve joint recognition of multiple tasks, such as clear frame classification, clear mucosal segmentation, assessment of fold blind areas, and intestinal segment location identification. Through feature sharing and loss complementarity, the overall assessment accuracy and clinical relevance are significantly improved, enabling comprehensive quality assessment of colonoscopy retraction videos.
[0021] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0022] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. In the drawings:
[0023] Figure 1 This is a flowchart of an artificial intelligence-based method for assessing the integrity of colonoscopy observation according to an embodiment of the present invention;
[0024] Figure 2 A schematic diagram of the overall network structure for implementing the AI-based colonoscopy observation integrity assessment method of this invention.
[0025] Figure 3 This is a schematic diagram of the network structure of the spatiotemporal attention submodule in an embodiment of the present invention;
[0026] Figure 4 This is a structural block diagram of an artificial intelligence-based colonoscopy observation integrity assessment system according to an embodiment of the present invention. Detailed Implementation
[0027] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0028] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art and should not be interpreted in an idealized or overly formal sense unless specifically defined.
[0029] To address the limitations of existing technologies in objectively, in real-time, and accurately quantifying the "effective mucosal observation coverage" during colonoscopy withdrawal, and the lack of direct assessment of the integrity of colonoscopy observation, this invention provides an artificial intelligence-based method for assessing the integrity of colonoscopy observation. This method can objectively and quantify the observation quality score of a single colonoscopy withdrawal process, avoiding missed adenomas due to insufficient observation during withdrawal (such as blurred vision, omission of folded blind areas, and uneven distribution of observation time) in traditional colonoscopy. Figure 1 As shown, the artificial intelligence-based colonoscopy observation integrity assessment method proposed in this invention includes the following steps:
[0030] S11. Obtain the video frame sequence of the colonoscopy withdrawal process, and divide the image frames in the video frame sequence into pixel blocks;
[0031] S12. A shared feature extraction backbone network based on hierarchical visual training is used to extract features at multiple scales for each pixel block of the image frame, resulting in local texture semantic feature maps, global structure semantic feature maps, and global context semantic feature maps with the resolution of the image frame decreasing sequentially and the number of channels increasing sequentially.
[0032] S13. Stack semantic feature maps of the same scale in the video frame sequence in the temporal dimension, and extract spatial and temporal features from the stacked features of semantic feature maps at each scale respectively; multiply the spatial and temporal features of the semantic feature maps at each scale pixel by pixel to obtain enhanced feature maps containing spatiotemporal information of local texture semantic feature maps, global structure semantic feature maps, and global context semantic feature maps; upsample the enhanced feature map of the global context semantic feature map and add it pixel by pixel to the enhanced feature map of the global structure semantic feature map to obtain a preliminary fusion feature map; upsample the preliminary fusion feature map and add it pixel by pixel to the enhanced feature map of the local texture semantic feature map to obtain a fusion feature map.
[0033] S14. Input the fused feature map into the pre-trained multi-task head, and output the clear mucosal region mask, wrinkled region mask, clarity classification result and intestinal segment location recognition result for each image frame respectively.
[0034] S15. The integrity of colonoscopy observation is quantitatively scored based on the clear mucosal area mask, wrinkled area mask, clarity classification results, and intestinal segment location identification results of each image frame.
[0035] The present invention provides an AI-based method for assessing the integrity of colonoscopy observations. By combining the temporal stacking of multi-scale semantic features and a spatiotemporal attention mechanism, it achieves effective modeling of dynamic interferences in videos (such as camera shake and rapid retraction) and the unfolding process of folds. High-resolution enhanced features are generated through multi-scale fusion and used as shared inputs for multiple task heads, improving the robustness and consistency of feature representation. A dedicated task head is designed to achieve joint recognition of multiple tasks, such as clear frame classification, clear mucosal segmentation, assessment of fold blind areas, and intestinal segment location identification. Through feature sharing and loss complementarity, the overall assessment accuracy and clinical relevance are significantly improved, enabling comprehensive quality assessment of colonoscopy retraction videos.
[0036] In this embodiment of the invention, step S11, which involves dividing the image frames in the video frame sequence into pixel blocks, specifically includes: normalizing the RGB pixel values of each image frame in the video frame sequence according to preset normalization parameters; dividing the normalized image frames into pixel blocks, configuring relative position encoding for each pixel block, and mapping the pixel blocks to a high-dimensional embedding space through a fully connected layer.
[0037] The overall network of this invention has an end-to-end structure, such as... Figure 2 As shown, the input is a continuous video frame sequence (T consecutive frames, e.g., T=8) of the colonoscopy withdrawal process, and the output is multiple task predictions.
[0038] Specifically, input:
[0039] * Video Sequence: A sequence of video frames with a resolution of 512×512. The frame bitmap data (RGB) is normalized, with mean=(0.485, 0.456, 0.406) and std=(0.229, 0.224, 0.225).
[0040] *Patch Embedding: The patch embedding network divides the image into 4*4 pixel blocks, 512 / 4=128, 128*128=16348 blocks. Each block is a small local image (size 4*4*3=48 values), which is mapped to a high-dimensional embedding space (96 dimensions, C=96) through a fully connected layer.
[0041] Position Embedding: A position encoding network that adds relative position bias (RPB) to 16348 blocks (tokens).
[0042] In this embodiment of the invention, the shared feature extraction backbone network includes a first feature extraction layer, a second feature extraction layer, a third feature extraction layer, and a fourth feature extraction layer arranged in a cascaded manner, as well as a block merging layer distributed between any two feature extraction layers. The first feature extraction layer, the second feature extraction layer, the third feature extraction layer, and the fourth feature extraction layer all use a cascaded Swing Block network to perform feature extraction. The block merging layer reduces the resolution of the feature map output by the previous feature extraction layer by half and doubles the number of channels, and then inputs it into the next feature extraction layer.
[0043] For details, see Figure 2 The Shared Backbone is a feature extraction backbone network derived using SwinTransformer Tiny (pre-trained on ImageNet-1K). Implemented through four feature extraction stages and three block merging layers, SwinTransformer Tiny outputs multi-scale features (C2, C3, C4, C5), with progressively decreasing resolution (×2) and progressively increasing channels (×2). These multi-scale features provide comprehensive feature support from details (clear mucosal boundaries) to the global picture (intestinal segment location, blind zone coverage).
[0044] Stage 1, the first feature extraction layer, consists of two [Swin Blocks] leading to output C2 (H / 4 × W / 4, C=96). This stage initially extracts low-level visual features (such as edges, texture, and color) and downsamples the input image to a medium resolution. Two consecutive Swing Transformer Blocks capture feature dependencies within local windows through Window Self-Attention (Window MSA) and Shifted Window MSA, establishing a preliminary global context. Output C2 represents the highest resolution deep features, preserving a significant amount of detail.
[0045] Stage 2, the second feature extraction layer, [Swin Block] ×2 → Output C3 (H / 8 × W / 8, C=192), further extracts intermediate semantic features (such as local structure and mucosal texture patterns) from the medium-resolution feature map. Window self-attention calculation continues on the downsampled features. Since the window size is fixed (default 7×7), the perceived size after downsampling is relatively larger, capturing a wider range of context. Output C3 balances resolution and semantic intensity.
[0046] The third feature extraction layer, Stage 3, [Swin Block] ×6 → Output C4 (H / 16 × W / 16, C=384), extracts high-level semantic features (such as overall intestinal structure, fold distribution, and abnormal patterns). Due to the lower resolution, window self-attention can efficiently capture long-range dependencies. Output C4 has the richest semantic information and is often used for global understanding tasks (such as classification and blind spot evaluation).
[0047] Stage 4, the fourth feature extraction layer, [Swin Block] ×2 → Output C5 (H / 32 × W / 32, C=768), extracts the highest-level global semantic features for the most abstract tasks (such as image-level classification or global context modeling). It has the lowest resolution and the largest receptive field, almost covering the entire image. Output C5 has the most channels and the highest information density, suitable for deep supervision or global pooling classification.
[0048] Patch Merge: After merging blocks and downsampling (DS), the resolution is reduced by half, and the number of channels (CH) is increased by half.
[0049] See Figure 2 The Temporal-Spatial Attention Module (TSA) is inserted after the backbone. After the shared feature extraction backbone network extracts local texture semantic feature maps, global structural semantic feature maps, and global context semantic feature maps, the TSA stacks features of the same scale along the temporal dimension. These features are then input into an independent TSA submodule, which sequentially multiplies the spatial and temporal features with the feature maps pixel-by-pixel to obtain an enhanced feature map containing spatiotemporal information. The three enhanced features are then upsampled and summed pixel-by-pixel to obtain a fused feature map, which serves as the module's output. Wherein:
[0050] Temporal Stacking module: Temporal stacking, stacks multiple feature maps (T=8) in the time dimension to form a 3D tensor, which is then input into the spatiotemporal attention submodule.
[0051] The Temporal-Spatial Attention submodule independently applies spatial attention (focusing on clear mucosal regions) and temporal attention (capturing inter-frame dynamics, such as wrinkle unfolding or camera movement) at each scale to obtain spatial features (M_spatial) and temporal features (M_temporal).
[0052] Decoder: Uses transposed convolution upsampling (US) to increase the resolution by 2 times, and uses 1×1 convolution to reduce the number of channels (CH) by 2 times.
[0053] Feature fusion module: The feature map is multiplied by the attention weights to output the enhanced feature F_fused.
[0054] In this embodiment of the invention, spatial feature extraction is performed on the stacked features of semantic feature maps at various scales, including: spatial feature extraction is performed on the stacked features of semantic feature maps at each scale using a preset spatiotemporal attention submodule; the spatiotemporal attention submodule includes a spatial attention branch, which is used to obtain the maximum value in the stacked features in the temporal dimension to obtain a representative feature map of the video frame sequence, obtain the channel-dimensional average feature and the channel-dimensional maximum feature of the representative feature map, stack the channel-dimensional average feature and the channel-dimensional maximum feature in the channel dimension to obtain channel fusion features, perform convolution transformation on the channel fusion features to reduce the channel dimension to 1, and perform batch regularization and activation processing on the convolutional features to obtain spatial features.
[0055] In this embodiment of the invention, temporal feature extraction is performed on the stacked features of semantic feature maps at various scales, including: extracting temporal features from the stacked features of semantic feature maps at each scale using a preset spatiotemporal attention submodule; the spatiotemporal attention submodule includes a temporal attention branch, which is used to perform zero-parameter temporal adjustment on the stacked features to obtain forward-shifted features, backward-shifted features, and unshifted features; stacking the forward-shifted features, backward-shifted features, and original features in the channel dimension to obtain temporally adjusted features; performing global average pooling on the H×W dimension of the temporally adjusted features to obtain basic temporal features; performing two channel reduction transformations on the basic temporal features using preset different weight matrices to construct the query and key of attention; calculating the similarity matrix between time steps based on the query and key of attention; weighting the basic temporal features using the similarity matrix to obtain weighted features that integrate global temporal information; restoring the number of channels by performing 1×1 convolution on the weighted features and performing residual addition with the basic temporal features to obtain the final temporal features; and performing Sigmoid activation on the final temporal features to obtain the temporal weight matrix. Among them, the forward shift feature is the feature that shifts 1 / 4 of channel C forward by 1 frame in the time dimension, the backward shift feature is the feature that shifts the other 1 / 4 of channel C backward by 1 frame in the time dimension, and the no-offset feature is the feature that keeps the remaining 1 / 2 of channel C unchanged.
[0056] Figure 3 A schematic diagram of the network structure of the spatiotemporal attention submodule in an embodiment of the present invention is shown. See also Figure 3 The spatiotemporal attention submodule includes a spatial attention branch and a temporal attention branch. Specifically:
[0057] (1) The network structure of the Spatial Attention Branch is as follows:
[0058] Feature map selection network Max Pooling on T: Take the maximum value in the temporal dimension to obtain the representative feature map F_spatial (H × W × C) of frame T.
[0059] Average Pooling on C: Channel-dimensional average, yielding F_avg (H × W × 1).
[0060] Max Pooling on C: Maximize the channel dimension to obtain F_max (H × W × 1).
[0061] Stacking module: Stack F_avg and F_max in the channel dimension to obtain F_concat(H × W × 2).
[0062] The convolutional network Conv 7x7, the batch regularization network BatchNorm, and the activation module Sigmoid are used to perform a convolutional transformation on F_concat, changing the channel dimension from 2 to 1. Then, batch regularization and activation are performed to obtain M_spatial (H × W × 1).
[0063] (2) The network structure of the Temporal Attention Branch is as follows:
[0064] Time Shifting Network: Zero-parameter time shifting. One-quarter of channel C is shifted forward by one frame (Forward) to obtain F_future (T × H × W × C / 4). One-quarter of channel C is shifted backward by one frame (Backward) to obtain F_past (T × H × W × C / 4). The remaining half of the channel remains unchanged (Stationary) to obtain F_original (T × H × W × C / 2). For boundary frames (t=0 and t=T-1), they are padded with 0s or adjacent frames are copied.
[0065] Stacking module: Stack F_future, F_past and F_original in the channel dimension to obtain F_shifted (T × H × W × C).
[0066] Global Average Pooling: Perform global average pooling on the H × W dimension to obtain F_gap (T × C).
[0067] Query Embedding: Embed the query, θ(F_gap) = W_θ(F_gap), to get Query (T× C / r), where W_θ is the preset first weight matrix.
[0068] Key Embedding: Embed the key, φ(F_gap) = W_φ(F_gap), to obtain Value (T ×C / r), where W_φ is the preset second weight matrix.
[0069] To calculate the correlation between different time steps, F_gap is first subjected to two channel reduction transformations (using two weight matrices, W_θ and W_φ, respectively). One transformation yields the "Query" and the other yields the "Key". Both transformations reduce the number of channels from C to C / r, where r is a preset reduction ratio, with the aim of reducing the amount of computation.
[0070] Similarity matrix calculation: Attention_map = softmax(Query × Key / √(C / r) ), which performs matrix multiplication on the embedded query matrix and key matrix to obtain F_attention (T × T).
[0071] Value Embedding: Embed the value, g(F_gap) = W_g(F_gap), to obtain Value (T × C). W_g is the preset third weight matrix. Then, perform another transformation on the original F_gap to obtain the "Value", which is the core feature of each time step.
[0072] Weighted aggregation: F_nonlocal = F_attention × Value. Perform matrix multiplication on the similarity matrix (F_attention) and the value matrix to obtain F_nonlocal (T × C). Multiplying the previously calculated similarity matrix and the "value" is equivalent to "taking the feature at each time step and summing it according to its correlation with all other time steps," resulting in a feature (F_nonlocal) that incorporates global temporal information.
[0073] Residual fusion is performed and channels are restored via 1×1 convolution: F_temporal = F_gap + W_out(F_nonlocal), resulting in F_temporal (T × C). W_out is the preset fourth weight matrix.
[0074] Generating time series weights using Sigmoid: Perform Sigmoid on F_temporal to obtain time series weights M_temporal (T x 1).
[0075] In this embodiment, spatial information is first removed, focusing only on temporal features. Then, the correlation between time steps is calculated through "query-key". The correlation is used to weight the features, and finally, the importance weight of each time step is generated.
[0076] In embodiments of the present invention, such as Figure 2 As shown, the multi-task head includes a clear mucosal region segmentation head, a fold blind zone coverage evaluation head, a clear frame classification head, and an intestinal segment location identification head. Among them:
[0077] Clear Mask Head: A clear mucosal region segmentation head implemented based on the UNet++ architecture. It can use F_fused as encoder input to output a pixel-level mask of the clear mucosal region (clear mucosa vs. blurry / dirty mucosa) based on the fused feature map.
[0078] Fold Mask Head: A wrinkle blind area coverage evaluation head that uses lightweight ASPP (Atrous Spatial Pyramid Pooling) to capture multi-scale wrinkles and output a pixel-level mask (wrinkle region) of the wrinkle area based on the fused feature map.
[0079] Clear Classification Head: The clear frame classification head includes a first global average pooling layer and a first fully connected layer. It uses global average pooling + fully connected layer to identify whether the image frame is clear or not based on the fused feature map in a binary classification (clear / not clear).
[0080] Seg Classification Head: The intestinal segment location identification head includes a second global average pooling layer and a second fully connected layer. It uses global average pooling + fully connected layer to identify the six categories of intestinal segment location (ileocecal / ascending colon / transverse colon / descending colon / sigmoid colon / rectum) of image frame pairs based on the fused feature map.
[0081] In one specific embodiment, the clear mucosal region segmentation head includes a first encoder, a second encoder, a third encoder, a second decoder, a first decoder, and an output layer. The first encoder, second encoder, and third encoder are used to extract shallow, mid-level, and deep features from the fused feature map, respectively. The second decoder is used to upsample the deep features to concatenate the sampled features with the mid-level features, and then perform two consecutive convolutions on the concatenated feature map to extract the first deep fused feature. The first decoder is used to upsample the first deep fused feature to concatenate the sampled features with the shallow features, and then perform two consecutive convolutions on the concatenated feature map to extract the second deep fused feature. The output layer is used to output a pixel-level mask of the clear mucosal region based on the second deep fused feature. The decoder can be implemented using a MaxPool + DoubleConv structure, or it can be implemented using an UpConv + Cat + DoubleConv structure.
[0082] In one specific embodiment, the wrinkle blind zone coverage evaluation head includes a global feature extraction network consisting of parallel 1x1 convolutions, a mid-scale wrinkle feature extraction network with 3x3 dilated convolutions and a dilation rate of 3, a large-scale wrinkle feature extraction network with 3x3 dilated convolutions and a dilation rate of 5, a feature fusion layer, a lightweight decoder, and an output layer. The feature fusion layer is used to concatenate the global, mid-scale, and large-scale wrinkle features extracted from the fused feature map by the global, mid-scale, and large-scale wrinkle feature extraction networks, and then convolve the concatenated feature map to extract fused multi-scale features. The lightweight decoder is used to extract local detail features from the fused multi-scale features, and the output layer is used to output a pixel-level mask of the wrinkle region based on the local detail features. The lightweight decoder can be implemented using a 3x3Conv + 1x1Conv structure.
[0083] The joint training loss of the overall network of the artificial intelligence-based colonoscopy observation integrity assessment method in this embodiment of the invention is L_total:
[0084] L_total = λ1 L class + λ2 L seg + λ3 L blind + λ4 L position + λ attn L consistency ;
[0085] L class : Remove the classification loss from the Clear Classification Head.
[0086] L seg : Segmentation loss of Clear Mask Head.
[0087] L blind : Segmentation loss of Fold Mask Head.
[0088] L position : Classification loss of Seg Classification Head;
[0089] L consistency Cross-task consistency loss, such as consistency between segmentation mask and classification.
[0090] λ1, λ2, λ3, λ4 and λ attn The preset weights for each loss are initially set to 1 and are dynamically adjusted based on uncertainty weighting.
[0091] In this embodiment of the invention, the integrity of colonoscopy observation is quantitatively scored based on the clear mucosal region mask, the wrinkled region mask, the clarity classification result, and the intestinal segment location identification result of each image frame, including:
[0092] A pre-defined weighted scoring model was used to quantify the integrity of colonoscopy observations. The weighted scoring model is shown below:
[0093] Score = σ( w1 S clear + w2 S time + w3 S blind + w4 S dynamic + w5 S segment ) ×100;
[0094] Score is a quantitative rating;
[0095] w1, w2, w3, w4, and w5 are the corresponding sub-items S clear S time S blind S dynamic and S segment The weights are static values of 0-1. On multiple colonoscopy videos, the correlation between expert scores and model scores is calculated to try to find the optimal weight allocation.
[0096] S clear Sclear is the average percentage of clear mucous membrane pixels across all frames in the video frame sequence, calculated as: Sclear = (1 / N)∑ (number of clear pixels / total number of pixels). i The clarity of colonoscopy images reflects the overall clarity of the images and is a core traditional quality indicator. A higher percentage of clarity indicates less interference from blurriness, contamination, and air bubbles.
[0097] S time S is the normalized score for the uniformity and concentration of sharp image frames in a video sequence, classified according to their sharpness. time = 0.5 × (1 - normalized_skewness) + 0.5 × normalized_kurtosis. Ideally, normalized_skewness should be close to 0, and normalized_kurtosis should ideally be close to a Gaussian distribution of 3. Normalization to 0-1 is necessary to assess the uniformity of the distribution of clear observation time during the lens withdrawal process. Excessive skewness indicates that observations are concentrated in the later stages of lens withdrawal (missed early stages), while abnormal kurtosis indicates that observations are either too fragmented or overly concentrated.
[0098] S blind S is the average percentage of pixels in the wrinkled region across all frames in the video frame sequence. blind= Average blind zone exposure ratio, quantifying the degree of mucosal exposure behind intestinal folds. Intestinal folds are a high-risk area for missed adenomas. This indicator directly addresses the most important clinical pain point. A low score indicates a high risk of missed detection, while >80% is considered excellent.
[0099] S dynamic S represents the percentage of effective image frames in a video frame sequence determined based on temporal features. dynamic = (∑ M_temporal i × Clear Frame Marker i ) / T, M_temporal i The temporal attention weight for the i-th frame excludes invalid frames (flag = 0) caused by dynamic interference such as rapid camera exits and lens shake, reflecting the proportion of truly effective observation time and compensating for the deficiency of static proportion ignoring dynamic quality.
[0100] S segment To normalize the entropy of the clear observation time of each intestinal segment based on the results of intestinal segment location identification, S segment =Entropy(p) / log(K), where p is the percentage distribution of clear observation time for K intestinal segments, and K is the number of categories of the intestinal segment location identification head, K=6; the higher the entropy, the more balanced it is. It assesses whether the main intestinal segments such as the ileocecal junction, ascending colon, transverse colon, descending colon, sigmoid colon, and rectum are observed evenly, avoiding over-observation of some intestinal segments while other segments are missed.
[0101] σ is the Sigmoid function, which is non-linearly mapped to 0-1;
[0102] × 100: Linear mapping to 0-100.
[0103] This invention provides an AI-based method for assessing the integrity of colonoscopy observation. It utilizes a multi-task deep learning network and a spatiotemporal attention mechanism to score the integrity of colonoscopy observation. Through an end-to-end MucosaComplete-Net network, it achieves comprehensive quality assessment of colonoscopy withdrawal videos, overcoming the shortcomings of existing technologies that rely solely on the proportion of clear regions and simple time statistics, which cannot accurately reflect clinical pain points such as missed fold blind spots, uneven dynamic observation, and imbalanced intestinal segment coverage. Specifically, this invention uses a Swing Transformer as a shared backbone, combining temporal stacking of multi-scale features (C3, C4, C5) and a spatiotemporal attention mechanism (spatial attention branch + temporal attention branch) to effectively model dynamic interference in the video (such as camera shake and rapid withdrawal) and the fold unfolding process. A high-resolution enhanced feature F_fused is generated through a top-down multi-scale fusion method, serving as a shared input for multiple task heads, improving the robustness and consistency of feature representation. This invention designs dedicated task heads, particularly a fold blind spot coverage assessment head, introducing multi-scale context capture for intestinal folds to achieve joint output of blind spot masks and coverage scores. The multi-task joint training framework proposed in this invention includes a clear mucosal region segmentation head, a fold blind zone coverage assessment head, a clear frame classification head, and an intestinal segment location identification head. It also includes a multi-dimensional scoring index system based on the output of these tasks (especially the newly added fold blind zone coverage rate, dynamic observation integrity index, and intestinal segment coverage balance). The multi-item weighted fusion makes the final score closer to the clinical effect. Through feature sharing and loss complementarity, it significantly improves the overall assessment accuracy and clinical relevance.
[0104] For the sake of simplicity, the method embodiments are described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.
[0105] This invention provides an artificial intelligence-based colonoscopy observation integrity assessment system, the system including functional modules for implementing the artificial intelligence-based colonoscopy observation integrity assessment method as described in any of the preceding claims. Figure 4 The schematic diagram illustrates the structure of an artificial intelligence-based colonoscopy observation integrity assessment system provided by an embodiment of the present invention. (Refer to...) Figure 4 The system described in this embodiment of the invention includes:
[0106] The data preprocessing module 401 is used to acquire the video frame sequence of the colonoscopy withdrawal process and divide the image frames in the video frame sequence into pixel blocks;
[0107] The semantic feature extraction module 402 is used to perform multi-scale feature extraction on each pixel block of the image frame using a shared feature extraction backbone network based on hierarchical visual training, to obtain local texture semantic feature maps, global structure semantic feature maps and global context semantic feature maps with the resolution of the image frame decreasing sequentially and the number of channels increasing sequentially.
[0108] The spatiotemporal attention module 403 is used to stack semantic feature maps of the same scale in the temporal dimension of the video frame sequence, and to extract spatial features and temporal features from the stacked features of semantic feature maps at each scale respectively; the spatial features and temporal features of the semantic feature maps at each scale are multiplied pixel by pixel with the corresponding semantic feature maps to obtain enhanced feature maps containing spatiotemporal information, namely, local texture semantic feature maps, global structure semantic feature maps, and global context semantic feature maps; the enhanced feature map of the global context semantic feature map is upsampled and added pixel by pixel with the enhanced feature map of the global structure semantic feature map to obtain a preliminary fusion feature map; the preliminary fusion feature map is upsampled and added pixel by pixel with the enhanced feature map of the local texture semantic feature map to obtain a fusion feature map.
[0109] The feature recognition module 404 is used to input the fused feature map into the pre-trained multi-task head and output the clear mucosal region mask, the wrinkled region mask, the clarity classification result and the intestinal segment location recognition result for each image frame.
[0110] The scoring module 405 is used to quantify the integrity of colonoscopy observation based on the clear mucosal area mask, folded area mask, clarity classification results, and intestinal segment location identification results for each image frame.
[0111] As the system implementation is basically similar to the method implementation, the description is relatively simple, and relevant parts can be found in the description of the method implementation.
[0112] Furthermore, another embodiment of the present invention provides a computer program product storing a computer program, which, when executed by a processor, implements the steps described in the above embodiment of the artificial intelligence-based colonoscopy observation integrity assessment method, for example... Figure 1 Steps S11-S15 are shown.
[0113] Furthermore, another embodiment of the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When executed by the processor, the computer program implements the steps described in the above embodiment of the artificial intelligence-based colonoscopy observation integrity assessment method, for example... Figure 1 Steps S11-S15 are shown.
[0114] Furthermore, those skilled in the art will understand that although some embodiments herein include certain features included in other embodiments but not others, combinations of features from different embodiments are intended to be within the scope of the invention and form different embodiments. For example, any of the claimed embodiments can be used in any combination.
[0115] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An artificial intelligence-based colonoscopy observation completeness evaluation method, characterized by, The method includes: Acquire a video frame sequence of the colonoscopy withdrawal process, and divide the image frames in the video frame sequence into pixel blocks; A shared feature extraction backbone network based on hierarchical visual training is used to extract features at multiple scales for each pixel block of the image frame, resulting in local texture semantic feature maps, global structure semantic feature maps, and global context semantic feature maps with the resolution decreasing sequentially and the number of channels increasing sequentially. Semantic feature maps of the same scale from a video frame sequence are stacked temporally. Spatial and temporal features are extracted from the stacked features of semantic feature maps at each scale. The spatial and temporal features of the semantic feature maps at each scale are multiplied pixel-by-pixel with the corresponding semantic feature maps to obtain enhanced feature maps containing spatiotemporal information, including local texture semantic feature maps, global structure semantic feature maps, and global context semantic feature maps. The enhanced feature map of the global context semantic feature map is upsampled and added pixel-by-pixel to the enhanced feature map of the global structure semantic feature map to obtain a preliminary fusion feature map. The preliminary fusion feature map is then upsampled and added pixel-by-pixel to the enhanced feature map of the local texture semantic feature map to obtain the final fusion feature map. The fused feature map is input into the pre-trained multi-task head, which outputs the clear mucosal region mask, the wrinkled region mask, the clarity classification result, and the intestinal segment location recognition result for each image frame. The integrity of colonoscopy observation is quantitatively scored based on the clear mucosal area mask, folded area mask, clarity classification results, and intestinal segment location identification results for each image frame.
2. The method of claim 1, wherein, Divide the image frames in the video frame sequence into pixel blocks, including: The RGB pixel values of each image frame in the video frame sequence are normalized according to the preset normalization parameters. The normalized image frame is divided into pixel blocks, each pixel block is configured with a relative position code, and the pixel blocks are mapped to a high-dimensional embedding space through a fully connected layer.
3. The method of claim 1, wherein, The shared feature extraction backbone network includes a cascaded first feature extraction layer, a second feature extraction layer, a third feature extraction layer, and a fourth feature extraction layer, as well as a block merging layer distributed between any two feature extraction layers. The first, second, third, and fourth feature extraction layers all use a cascaded Swing Block network to extract features. The block merging layer reduces the resolution of the feature map output by the previous feature extraction layer by half and doubles the number of channels, and then inputs it into the next feature extraction layer.
4. The method according to claim 1, characterized in that, Spatial feature extraction is performed on the stacked features of semantic feature maps at various scales, including: Spatial features are extracted from the stacked features of the semantic feature maps at each scale using a pre-defined spatiotemporal attention submodule. The spatiotemporal attention submodule includes a spatial attention branch, which is used to obtain the maximum value in the stacked features in the temporal dimension to obtain the representative feature map of the video frame sequence. The representative feature map is then used to obtain the channel-dimensional average feature and the channel-dimensional maximum feature. The channel-dimensional average feature and the channel-dimensional maximum feature are stacked in the channel dimension to obtain the channel fusion feature. The channel fusion feature is then subjected to a convolutional transformation to reduce the channel dimension to 1. The convolutional feature is then subjected to batch regularization and activation processing to obtain the spatial feature.
5. The method according to any one of claims 1 to 4, characterized in that, Temporal feature extraction is performed on the stacked features of semantic feature maps at various scales, including: Temporal features are extracted from the stacked features of the semantic feature maps at each scale using a pre-defined spatiotemporal attention submodule. The spatiotemporal attention submodule includes a temporal attention branch, which performs zero-parameter temporal adjustment on stacked features to obtain forward-shifted features, backward-shifted features, and unshifted features. The forward-shifted features, backward-shifted features, and original features are stacked along the channel dimension to obtain temporally adjusted features. Global average pooling is performed on the H×W dimension of the temporally adjusted features to obtain basic temporal features. Two channel reduction transformations are performed on the basic temporal features using preset weight matrices to construct the query and key of attention. The similarity matrix between time steps is calculated based on the query and key of attention. The basic temporal features are weighted using the similarity matrix to obtain weighted features that incorporate global temporal information. The weighted features are restored to the number of channels through 1×1 convolution and the residuals are added to the basic temporal features to obtain the final temporal features. The final temporal features are activated with Sigmoid to obtain the temporal weight matrix.
6. The method of claim 1, wherein, The multi-task head includes a clear mucosal region segmentation head, a fold blind zone coverage assessment head, a clear frame classification head, and an intestinal segment location identification head; The clear mucosal region segmentation head is implemented based on the UNet++ architecture to output a pixel-level mask of the clear mucosal region based on the fused feature map; the wrinkled blind zone coverage evaluation head is implemented based on the lightweight ASPP network to output a pixel-level mask of the wrinkled region based on the fused feature map; the clear frame classification head includes a first global average pooling layer and a first fully connected layer to identify whether the image frame is clear based on the fused feature map; the intestinal segment location identification head includes a second global average pooling layer and a second fully connected layer to identify the intestinal segment location of the image frame pair based on the fused feature map.
7. The method of claim 6, wherein, The clear mucosal region segmentation head includes a first encoder, a second encoder, a third encoder, a second decoder, a first decoder, and an output layer. The first encoder, second encoder, and third encoder are used to extract shallow, medium, and deep features from the fused feature map, respectively. The second decoder is used to upsample the deep features to concatenate the sampled features with the medium features, and then perform two consecutive convolutions on the concatenated feature map to extract the first deep fused feature. The first decoder is used to upsample the first deep fused feature to concatenate the sampled features with the shallow features, and then perform two consecutive convolutions on the concatenated feature map to extract the second deep fused feature. The output layer is used to output a pixel-level mask of the clear mucosal region based on the second deep fused feature.
8. The method of claim 6, wherein, The wrinkle blind zone coverage evaluation head includes a global feature extraction network consisting of parallel 1x1 convolutions, a medium-scale wrinkle feature extraction network with 3x3 dilated convolutions and a dilation rate of 3, a large-scale wrinkle feature extraction network with 3x3 dilated convolutions and a dilation rate of 5, a feature fusion layer, a lightweight decoder, and an output layer. The feature fusion layer is used to concatenate the global, medium-scale, and large-scale wrinkle features extracted from the fused feature map by the global, medium-scale, and large-scale wrinkle feature extraction networks, and then convolve the concatenated feature map to extract fused multi-scale features. The lightweight decoder is used to extract local detail features from the fused multi-scale features, and the output layer is used to output a pixel-level mask of the wrinkle region based on the local detail features.
9. The method of claim 1, wherein, The completeness of colonoscopy observation is quantitatively scored based on the masking of clear mucosal areas, masking of folded areas, clarity classification results, and intestinal segment location identification results for each image frame, including: A pre-defined weighted scoring model was used to quantify the integrity of colonoscopy observations. The weighted scoring model is shown below: Score = σ( w1 S clear + w2 S time + w3 S blind + w4 S dynamic + w5 S segment ) × 100; Score is a quantitative rating, and w1, w2, w3, w4, and w5 are the corresponding sub-items. clear S time S blind S dynamic and S segment The weights are σ, which is the Sigmoid function. clear S is the average percentage of clear mucosal region pixels across all frames in the video frame sequence. time S is the normalized score for the uniformity and concentration of sharp image frames in a video sequence, classified according to their sharpness. blind S is the average percentage of pixels in the wrinkled region across all frames in the video frame sequence. dynamic S represents the percentage of effective image frames in a video frame sequence determined based on temporal features. segment Entropy normalization of the clear observation time of each intestinal segment based on the results of intestinal segment location identification.
10. An artificial intelligence-based colonoscopy observation completeness evaluation system, characterized by, The system includes: The data preprocessing module is used to acquire the video frame sequence of the colonoscopy withdrawal process and divide the image frames in the video frame sequence into pixel blocks; The semantic feature extraction module is used to perform multi-scale feature extraction on each pixel block of the image frame using a shared feature extraction backbone network based on hierarchical vision training, to obtain local texture semantic feature maps, global structure semantic feature maps and global context semantic feature maps with the resolution of the image frame decreasing sequentially and the number of channels increasing sequentially. The spatiotemporal attention module stacks semantic feature maps of the same scale from a video frame sequence along a temporal dimension. Spatial and temporal features are extracted from the stacked features of the semantic feature maps at each scale. The spatial and temporal features of each scale's semantic feature map are then multiplied pixel-by-pixel with the corresponding semantic feature map to obtain enhanced feature maps containing spatiotemporal information, including local texture semantic feature maps, global structure semantic feature maps, and global context semantic feature maps. The enhanced feature map of the global context semantic feature map is upsampled and added pixel-by-pixel to the enhanced feature map of the global structure semantic feature map to obtain a preliminary fused feature map. Finally, the preliminary fused feature map is upsampled and added pixel-by-pixel to the enhanced feature map of the local texture semantic feature map to obtain the final fused feature map. The feature recognition module is used to input the fused feature map into the pre-trained multi-task head and output the clear mucosal region mask, the wrinkled region mask, the clarity classification result, and the intestinal segment location recognition result for each image frame. The scoring module is used to quantify the integrity of colonoscopy observation based on the clear mucosal area mask, wrinkled area mask, clarity classification results, and intestinal segment location identification results of each image frame.
Citation Information
Patent Citations
Endoscopic examination evaluation system and method based on video understanding network
CN118710995A
Enteroscope qualified retreating time evaluation method and system based on artificial intelligence
CN119091206A