Intestinal tract preparation quality intelligent monitoring and precise intervention system based on image recognition
By combining multimodal imaging with intelligent algorithms, the problems of strong scoring subjectivity and poor cross-operator consistency in existing systems have been solved. This has enabled efficient, interpretable, automated assessment and precise intervention of bowel preparation quality, improved the real-time performance and reliability of the system, reduced misjudgments and duplicate examinations, and promoted the intelligentization and standardization of clinical procedures.
Patent Information
- Application Number
- CN202511568856.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-30
- Publication Date
- 2026-02-03
AI Technical Summary
Existing image recognition-based intelligent monitoring and precise intervention systems for bowel preparation quality rely on manual visual assessment, resulting in highly subjective scoring, poor cross-operator consistency, and an inability to meet the needs of large-scale and real-time monitoring. Furthermore, they lack robust automatic segmentation and parameter adaptation capabilities, making them prone to misjudgment and missed judgment, which affects downstream intervention decisions. The scores are also uninterpretable and difficult to accurately map to clinical scoring standards, reducing the accuracy and reliability of intervention recommendations.
It employs a multimodal imaging acquisition module, an image processing and recognition module, an intelligent intestinal monitoring module, an intestinal preparation scoring and intervention module, and a report display and quality management module. Combining image feature extraction, machine learning, and deep learning algorithms, it achieves automated analysis and scoring of intestinal images, outputs interpretable scores and intervention suggestions, enhances recognition accuracy and stability through attention mechanisms and deep temporal coding, and provides real-time monitoring and anomaly alerts.
It improves the accuracy and real-time nature of bowel preparation quality assessment, reduces duplicate examinations, saves medical resources, enhances the traceability and quality control of the diagnostic process, ensures the scientific validity and safety of intervention recommendations, and realizes a closed loop from automatic identification to quantitative assessment to controllable intervention.
Smart Images

Figure CN121458652A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image recognition, machine learning and deep learning, in particular to an intelligent monitoring and precise intervention system for bowel preparation quality based on image recognition. BACKGROUND
[0002] Image recognition technology is a technology for automatically analyzing and recognizing intestinal images through computer vision and image processing methods, aiming to solve the shortcomings of traditional manual visual assessment in speed, accuracy and consistency, and can extract mucus, residue, bloodstains and intestinal wall observable features from the pixel level and regional level to ensure providing structured input for subsequent scoring and decision making. Machine learning technology is a pattern recognition method combining statistical methods and engineering practice, aiming to solve the problems of target recognition and semantic classification under the conditions of multi-source noise, device differences and limited labeled samples in bowel preparation images. Through the engineering combination of image features and shallow learning, distinguishable visual features are extracted and rapid recognition and anomaly detection of residue, mucus and bloody secretions are realized. The combination of image recognition and machine learning in an intelligent monitoring and precise intervention system for bowel preparation quality based on image recognition is a composite method combining visual feature extraction and evolutionary search optimization, aiming to solve the problems of misrecognition and false positives caused by noise, occlusion and diversity in bowel preparation images, so as to ensure that the recognition output can be used for reliable segmentation evaluation and real-time alarm.
[0003] Deep learning technology is an end-to-end representation learning method with multi-layer neural networks as the core, aiming to solve the limitations of traditional models based on manually labeled features in expressing complex morphology and temporal semantic information. It can automatically learn multi-scale and multi-frame correlation cleanliness features through convolution calculation and attention mechanism and output standardized scores and confidence. The bowel preparation scoring based on deep learning in an intelligent monitoring and precise intervention system for bowel preparation quality based on image recognition is an evaluation engine that integrates spatial features and time sequence information and maps to a clinical scoring system, aiming to solve the problems of subjectivity, inconsistency between operators and real-time scoring requirements, and ensure that the system can give an interpretable, quantifiable and usable bowel preparation quality evaluation for driving intervention decisions.
[0004] The existing image recognition-based intestinal preparation quality intelligent monitoring and precise intervention system has the problems of relying on manual visual assessment, leading to strong subjectivity of intestinal preparation quality score, poor consistency across operators, slow interpretation speed, and inability to meet large-scale and real-time monitoring requirements; secondly, the system lacks robust automatic image segmentation and parameter adaptive capabilities, and misjudgment and missed judgment are likely to occur when there are changes in illumination, occlusion and a large number of residues, affecting downstream intervention decisions; and the system lacks a scoring model and confidence output, making the score uninterpretable and difficult to accurately map to clinical scoring standards, thereby reducing the accuracy and reliability of intervention recommendations. SUMMARY
[0005] The present application aims to provide an image recognition-based intestinal preparation quality intelligent monitoring and precise intervention system to solve the problems of the existing image recognition-based intestinal preparation quality intelligent monitoring and precise intervention system in the background art, which relies on manual visual assessment, leading to strong subjectivity of intestinal preparation quality score, poor consistency across operators, slow interpretation speed, and inability to meet large-scale and real-time monitoring requirements; secondly, the system lacks robust automatic segmentation and parameter adaptive capabilities, and misjudgment and missed judgment are likely to occur when there are changes in illumination, occlusion and a large number of residues, affecting downstream intervention decisions; and the system lacks a scoring model and confidence output, making the score uninterpretable and difficult to accurately map to clinical scoring standards, thereby reducing the accuracy and reliability of intervention recommendations.
[0006] To achieve the above objectives, the present invention provides the following technical solution: an intelligent monitoring and precise intervention system for bowel preparation quality based on image recognition, comprising a multimodal imaging acquisition module, an image processing and recognition module, an intelligent bowel monitoring module, a bowel preparation scoring and intervention module, and a report display and quality management module. The bowel preparation imaging acquisition module is used to acquire bowel preparation image data, ensuring the acquisition of high-definition, synchronous, multi-view images with complete metadata for subsequent module analysis. The image processing and recognition module includes an image feature extraction unit and an image recognition unit. The image feature extraction unit is used to preprocess and structurally analyze the acquired images, extracting key quantitative features such as color, texture, shape, and residual coverage from the preprocessed images to ensure a stable and reliable input feature vector for image recognition. The image recognition unit proposes a machine learning-based bowel image recognition algorithm for target detection and segmentation of bowel images. The system includes a classification and assessment module to ensure robust residue identification and site localization under complex visual conditions. The bowel preparation scoring and intervention module comprises a bowel preparation scoring unit and a precise intervention decision unit. The bowel preparation scoring unit proposes a deep learning-based bowel preparation scoring algorithm to intelligently score bowel preparation quality and output confidence levels, ensuring high accuracy, repeatability, and interpretability. The precise intervention decision unit combines the scoring results with patient background, compliance, and clinical rules to generate personalized intervention strategies, ensuring that intervention recommendations are scientific, safe, and easy to implement and record. The intelligent bowel monitoring module monitors the scoring results and intervention strategies in real time and triggers abnormal alarms, ensuring timely detection of low-quality bowel preparation. The report display and quality management module generates clinical reports containing image evidence, scoring, and intervention records, and statistically analyzes bowel preparation quality trends, ensuring traceability, ease of quality control, and system management.
[0007] Preferably, the intestinal preparation imaging acquisition module integrates intestinal preparation images acquired by medical devices and performs timestamp synchronization, resolution standardization, and metadata recording on the acquired data to ensure high-quality, traceable, and structured image input for downstream processing and to guarantee transmission integrity and initial privacy protection.
[0008] Preferably, the image processing and recognition module includes an image feature extraction unit and an image recognition unit. The image feature extraction unit performs preprocessing such as cleaning and denoising on the original image, and calculates local and global multi-scale features, including color histograms, texture, morphology, edge and depth embedding, and time-frame level statistics. This ensures that robust feature vectors that are highly distinguishable between residues and mucosal differences are extracted from complex lighting and pollution conditions, providing robust input for recognition and scoring. The image recognition unit proposes a machine learning-based intestinal image recognition algorithm. By constructing a target detection network under a machine learning model, it classifies and labels image features, ensuring accurate differentiation of intestinal mucosa, residue types and distributions, and outputting quantifiable regional indicators.
[0009] Preferably, the machine learning-based intestinal image recognition algorithm is as follows: First, a linear transformation and nonlinear activation are applied to the feature vectors from the image feature extraction unit to construct an embedding representation adapted to the subsequent machine learning module, thereby unifying the feature scale and improving the representation ability for the intestinal image recognition task. By projecting the original feature vectors to a feature space more suitable for classification, the machine learning's ability to distinguish between different targets such as intestinal residue, liquid, and mucus is enhanced. The specific formula is expressed as follows:
[0010] in, This is represented as an input feature vector, derived from the image feature extraction unit. Represented as the real number field, and Represented as feature dimension, Represented as an embedding weight matrix, it is used to map the original features to... 3D embedding space, It is represented as a bias vector. This is represented as the embedded representation obtained after mapping, with each input frame corresponding to a segment. , This is represented by a hyperbolic tangent nonlinear activation function, used to enhance the nonlinearity of the expression and constrain the output range, which is beneficial for stable training. Then, by applying attention weighting to the embedding vector, the relative importance of each spatial and temporal location in the current recognition task is calculated, ensuring that the machine learning model can automatically focus on key areas containing residue edges, liquid layers, and lesions in the system, thereby improving the sensitivity and specificity of intestinal image recognition. The specific formula is expressed as follows:
[0011] in, This represents the number of embedding vectors collected at the current moment. Represented as the first Embedded vectors, Represented as the index of the number of embedding vectors, Represented as a query vector, it is used to measure the relevance of the embedding to the target being identified. This is expressed as a temperature coefficient, used to adjust the smoothness of the softmax distribution. Represented as the first Each embedded attention weight, The attention-weighted context representation vector serves as input for subsequent classification and segmentation. The attention-weighting strategy, within a machine learning framework, highlights the most critical regions and image frames for assessing bowel preparation quality. This allows the algorithm to achieve higher information density with lower computational cost, providing clearer evidence frames for bowel image recognition to support precise intervention decisions. Secondly, by inputting the attention representation into the discriminant network, the probability distribution of each bowel preparation candidate region belonging to various types of residues and background is calculated, serving as the basis for subsequent scoring and intervention decisions. This transforms continuous image representations into corresponding probability outputs, enabling the machine learning model to quantify the presence of various objects in the intestine, including key image data of residues, liquid layers, mucus, and healthy mucosa. The specific formula is as follows:
[0012] in, This is represented as the weight matrix of the discriminant network. This is represented as the bias vector of the discriminant network. Represented as a category number, including residue, liquid layer, mucus, and healthy mucous membrane. This is expressed as an element-wise exponential normalization function, making the output a probability distribution. This represents the probability vector for each category. The calculated probability vectors provide structured and interpretable category evidence for subsequent image recognition. Then, by fusing the probability vectors obtained from different image frames according to weights based on expert experience, a more robust instantaneous scoring basis is obtained. This ensures that heterogeneous evidence is fused in the system to reduce the impact of single-frame noise on the results of intestinal preparation image recognition. The specific formula is expressed as follows:
[0013] in, This represents the number of observed image frames involved in the fusion. Represented as the first The class probability vector obtained from each observation Represented as the index of the number of observed image frames participating in the fusion. Represented as the first The fusion weights for each observation are set according to a confidence decay strategy. The probability vector after fusion is represented as the instantaneous comprehensive distribution score of the intestinal preparation region. Secondly, by constructing a graph structure based on region adjacency and applying graph Laplacian smoothing to the fused probability vector, isolated misjudgments are suppressed and local consistency is enhanced, ensuring spatial coherence of the intestinal preparation score in the algorithm, thereby reducing the risk of false triggering of unnecessary interventions. The specific formula is expressed as follows:
[0014] in, Represented as an identity matrix, Represented as the smoothing intensity coefficient, Represented as a graph Laplacian matrix, defined as follows: ,in It is an adjacency matrix. For degree matrix, The smoothed probability vector represents the image recognition correction result with spatial consistency. Finally, by applying class determination and confidence calculation decision rules to the smoothed probability vector, a structured output is generated that can be used by subsequent modules, including the class label and confidence score of each intestinal preparation image. The specific formula is as follows:
[0015] in, This represents the category index used for the final determination. Represented as corresponding to The probability vector after category smoothing Represented as the index of the number of categories, This is expressed as the calibrated confidence level value. This is represented as a scaling parameter, used to adjust the impact of the original highest probability on the calibration output. Represented as an offset parameter, used to correct systematic offsets. Represented as the Sigmoid function, it maps the linearly transformed value to the (0,1) interval, making it convenient as the final interpretable confidence level.
[0016] Preferably, the bowel preparation scoring and intervention module includes a bowel preparation scoring unit and a precision intervention decision unit. The bowel preparation scoring unit proposes a deep learning-based bowel preparation scoring algorithm. Through the constructed deep learning model, it maps the identified types and features and outputs segmented scores, total bowel preparation scores, and confidence estimates, ensuring that the scoring has high accuracy, stability, and reliability for clinical decision-making. The precision intervention decision unit generates personalized and actionable intervention suggestions by inputting the scoring results, patient compliance records, past medical history, and clinical pathway rules into a decision rule base, ensuring that the intervention not only conforms to clinical norms but also improves the quality of bowel preparation and patient safety.
[0017] Preferably, the deep learning-based bowel preparation scoring algorithm is as follows: First, the output of each frame's image category and corresponding confidence score from the image recognition unit is converted into a continuous semantic vector. This allows the subsequent deep learning module to understand the semantic information of each frame using a unified and learnable vector representation, thereby supporting accurate bowel preparation scoring. By fusing discrete category evidence and confidence scores into a feature representation that can participate in deep network training, the algorithm's sensitivity and stability to local image evidence are improved. The specific formula is expressed as follows:
[0018] in, This represents the index of the current image frame. This represents the total number of categories determined by the image recognition unit, including residue, liquid layer, mucus, and healthy mucous membrane. This represents the category index determined by the image recognition unit. Represented as an image recognition unit for the first Frame number Confidence value of the category, Represented as the first Learned embedding vectors of categories, Represented as the constructed first Frame semantic vectors serve as the basic input to the deep learning scoring network. Then, the semantic vectors of consecutive frames are sequentially input into a deep temporal encoder to obtain a contextual representation with short- and medium-temporal information. This allows the deep learning gut health scoring system to enhance single-frame judgments using inter-frame dynamics. The specific formula is as follows:
[0019] in, Represented as the timing window length, This represents the vector concatenation operation. Represented as a weight matrix in time-series coding. Represented as the real number field, Represented as the number of dimensions, the concatenated... Projecting a dimensional vector to Vygote indicates that, It is represented as a bias vector. Represented as a nonlinear activation, it provides a smooth, bounded representation. Represented as the first The frame temporal context encoding representation serves as the intermediate semantic input to the scoring network. By capturing dynamic cues, the algorithm ensures that the final score is based not only on single-frame evidence but also on trends in neighboring frames. This reduces false alarms and improves response accuracy in real-time monitoring and alarm scenarios. Then, a deep fully connected transformation is applied to the temporal context representation to extract combined features highly correlated with bowel preparation scores. This allows the algorithm to learn the complex nonlinear mapping relationship between image sequences and clinical scores, providing a trainable feature base for the final continuous scoring. The specific formula is as follows:
[0020] in, Represented as the linear transformation weights of a deep scoring network. This is represented as the corresponding bias vector. Represented as the rectified linear unit activation function, it is used to introduce sparsity and nonlinearity. Represented as the function that maximizes, The hidden features, after transformation, carry high-order information directly related to the score. Through deep transformation, the algorithm can learn complex discriminative structures in a high-dimensional feature space, thereby enhancing its ability to distinguish between different bowel preparation qualities and laying the foundation for scoring accuracy. Secondly, the hidden features are mapped to continuous bowel preparation score values to directly output numerical scores that can be used for clinical interpretation and system decision-making. The specific formula is as follows:
[0021] in, Represented as a rating mapping vector, This is represented as a scalar bias. Represented as the Sigmoid function, Represented as an exponential function, Represented as variables, Represented as the first The continuous bowel preparation score of each frame, where 1 represents the worst score and 5 represents the best score, is directly output as a clinically usable continuous score through end-to-end deep mapping. This facilitates threshold judgment, trend monitoring, and visualization within the system. Then, based on the category distribution provided by the image recognition unit, the frame-level information entropy is calculated to estimate the reliability of the frame score, providing a quantitative basis for subsequent weighted aggregation and intervention decisions. The specific formula is as follows:
[0022] in, This is represented as the first step in step one. Frame number Confidence of the category Represented as a numerically stable term, Represented as information entropy, it represents a measure of uncertainty in frame category. This is represented by the maximum value of entropy. This is represented as frame-level confidence, with values closer to 1 indicating a more certain category distribution. Confidence estimation provides a quantification of uncertainty for deep learning-based bowel preparation scoring, enabling more rigorous reliability testing for automated interventions within the system. Low-confidence frames can be transferred to manual review to ensure clinical safety. Finally, the continuous scores of each frame within the examination sequence are weighted and averaged based on frame-level confidence to form the final bowel preparation score for the entire case. This provides clear numerical evidence for the precision intervention decision-making unit, while retaining evidence from each frame for auditing and manual review. The specific formula is as follows:
[0023] in, This represents the total number of frames used for scoring throughout the entire inspection. The weighted average bowel preparation score, representing the entire case, serves as the main quantitative indicator of the system output. By employing a confidence-weighted aggregation strategy, the influence of high-confidence frames is preferentially retained, while the interference of noisy frames on the final score is reduced. This allows the deep learning bowel preparation score to more reliably drive clinical decision-making within the system.
[0024] Preferably, the intestinal intelligent monitoring module performs real-time trend analysis, anomaly detection, and multi-channel alarms on the scoring time series, identification confidence level, and intervention suggestions to ensure that the system can identify abnormal and low-confidence scenarios, promptly detect abnormal intestinal preparation, and trigger corresponding alarm notifications.
[0025] Preferably, the report display and quality management module automatically generates a visual report that includes key areas of image recognition, segmented scores, total scores, intervention records, and confidence levels, and provides an interactive review interface and statistics on intestinal preparation quality trends, ensuring that users can easily manage and control the quality.
[0026] Compared with the prior art, the beneficial effects of the present invention are: 1. The image recognition unit proposes a machine learning-based intestinal image recognition algorithm. First, the algorithm maps the original feature vectors from the image feature extraction unit into highly adaptable embedding representations, unifying feature scales and enhancing semantic representation capabilities. This significantly improves the machine learning's ability to distinguish different targets in the intestine, such as residue, liquid, and mucus, overcoming the shortcomings of traditional manual interpretation, which is characterized by strong subjectivity and poor stability. Second, the algorithm introduces a local attention weight mechanism, which can automatically identify and amplify image regions containing key image evidence in the spatial and temporal domains. This enables the system to efficiently locate evidence frames that influence the determination of intestinal preparation instructions among numerous image frames. Then, based on the probability distribution of the discriminant network output, the algorithm transforms image representations into structured categorical evidence, providing interpretable input for the scoring and intervention decision-making units. This facilitates the automatic mapping of recognition conclusions into segmented and quantified intestinal preparation quality indicators. Simultaneously, through... By weighted fusion of probabilities from multiple image frames, the algorithm effectively suppresses the influence of single-frame noise and optical changes, improving the robustness of instantaneous judgment and ensuring stable operation of the system under different medical equipment and imaging conditions. Simultaneously, the algorithm employs spatial consistency constraints to smooth local judgments, reducing isolated misjudgments and minimizing false alarms and unnecessary interventions caused by noise. This enhances the safety and reliability of the system in real-world clinical environments. In summary, the machine learning-based intestinal image recognition algorithm not only improves the accuracy and real-time performance of image recognition but also strengthens auditing and manual review capabilities through interpretable confidence outputs and image evidence frame recording. It constructs a closed loop for intelligent monitoring and precise intervention of intestinal preparation quality, from automatic identification to quantitative assessment and controllable intervention. Ultimately, this improves the pass rate of intestinal preparation quality checks, reduces redundant examinations, saves medical resources, and enhances the consistency and traceability of clinical decisions.
[0027] The bowel preparation scoring unit proposes a deep learning-based bowel preparation scoring algorithm. First, the algorithm fuses the discrete categories and confidence scores output by the image recognition unit into a continuous semantic representation. This allows the deep learning model to understand the semantic information of each frame with a unified and learnable vector, significantly improving the ability to distinguish and detect targets such as residue, liquid layer, mucus, and healthy mucosa, thus overcoming the subjectivity and inconsistency issues of human interpretation. Second, by incorporating the semantic vectors of continuous frames into temporal coding, the algorithm can capture dynamic changes between frames and post-rinse cleaning trends, making the scoring independent of the occasional noise of isolated frames. This reduces false alarm rates, improves response accuracy, and supports intelligent suggestions for immediate actions in real-time monitoring scenarios. Third, deep feature transformation and end-to-end continuous scoring mapping enable the algorithm to learn the complex nonlinear mapping between image sequences and clinical scores, directly outputting clinically readable continuous scores. This facilitates integration with precision intervention decision-making units and quality management systems, achieving a comprehensive approach from image recognition to quantitative analysis. This algorithm establishes a closed loop from scoring to intervention instructions. Simultaneously, it provides confidence estimates for each frame score using an uncertainty measure based on category distribution, enabling the system to have a reliable confidence verification mechanism during automated decision-making. High-confidence frames support automatic issuance of intervention suggestions, while low-confidence frames automatically enter the manual review channel, improving the system's safety and controllability in clinical settings. Furthermore, a deep learning-based confidence-weighted sequence aggregation strategy prioritizes the retention of highly credible evidence, suppressing the impact of noisy frames on the final case score, thereby improving the robustness and interpretability of the final score. Overall, this algorithm not only improves the accuracy and real-time performance of bowel preparation assessment, reduces redundant examinations and saves medical resources, but also enhances the traceability and quality control capabilities of the diagnostic process through evidence frame preservation, confidence labeling, and visual reporting. It provides a technically feasible and clinically acceptable core capability for an image recognition-based intelligent monitoring and precise intervention system for bowel preparation quality, contributing to the intelligentization and standardization of clinical processes. Attached Figure Description
[0028] Figure 1 This is a schematic diagram of the structure of the present invention. Detailed Implementation
[0029] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0030] Please see Figure 1This invention provides an intelligent monitoring and precise intervention system for bowel preparation quality based on image recognition. It includes a multimodal imaging acquisition module, an image processing and recognition module, an intelligent bowel monitoring module, a bowel preparation scoring and intervention module, and a report display and quality management module. The bowel preparation imaging acquisition module collects bowel preparation image data, ensuring the acquisition of high-definition, synchronous, multi-view images with complete metadata for subsequent module analysis. The image processing and recognition module includes an image feature extraction unit and an image recognition unit. The image feature extraction unit preprocesses and performs structured analysis on the acquired images, extracting key quantitative features such as color, texture, shape, and residual coverage from the preprocessed images to ensure a stable and reliable input feature vector for image recognition. The image recognition unit proposes a machine learning-based bowel image recognition algorithm for target detection, segmentation, and classification of bowel images, ensuring... The system achieves robust residue identification and site localization under complex visual conditions. The bowel preparation scoring and intervention module includes a bowel preparation scoring unit and a precise intervention decision unit. The bowel preparation scoring unit proposes a deep learning-based bowel preparation scoring algorithm to intelligently score bowel preparation quality and output confidence scores, ensuring high accuracy, repeatability, and interpretability. The precise intervention decision unit combines the scoring results with patient background, compliance, and clinical rules to generate personalized intervention strategies, ensuring that intervention recommendations are scientific, safe, and easy to implement and record. The intelligent bowel monitoring module monitors scoring results and intervention strategies in real time and triggers abnormal alarms, ensuring timely detection of low-quality bowel preparation. The report display and quality management module generates clinical reports containing image evidence, scoring, and intervention records, and statistically analyzes bowel preparation quality trends, ensuring traceability, ease of quality control, and system management.
[0031] See Figure 1 Furthermore, the intestinal preparation imaging acquisition module integrates intestinal preparation images acquired by medical devices and performs timestamp synchronization, resolution standardization, and metadata recording on the acquired data to ensure high-quality, traceable, and structured image input for downstream processing and to guarantee transmission integrity and initial privacy protection.
[0032] See Figure 1Furthermore, the image processing and recognition module includes an image feature extraction unit and an image recognition unit. The image feature extraction unit performs preprocessing such as cleaning and denoising of the original image, as well as calculation of local and global multi-scale features, including color histograms, texture, morphology, edge and depth embedding, and time-frame level statistics, to ensure the extraction of highly distinguishable and robust feature vectors that differentiate between residues and mucosa under complex lighting and pollution conditions, providing robust input for recognition and scoring. The image recognition unit proposes a machine learning-based intestinal image recognition algorithm, which classifies and labels image features by constructing a target detection network under a machine learning model, ensuring accurate differentiation of intestinal mucosa, residue types and distribution, and outputting quantifiable regional indicators.
[0033] See Figure 1 Furthermore, the machine learning-based intestinal image recognition algorithm is as follows: First, a linear transformation and nonlinear activation are applied to the feature vectors from the image feature extraction unit to construct an embedding representation adapted to the subsequent machine learning module, thereby unifying the feature scale and improving the representation ability for the intestinal image recognition task. By projecting the original feature vectors to a feature space more suitable for classification, the machine learning's ability to distinguish between different targets such as intestinal residue, liquid, and mucus is enhanced. The specific formula is expressed as follows:
[0034] in, This is represented as an input feature vector, derived from the image feature extraction unit. Represented as the real number field, and Represented as feature dimension, Represented as an embedding weight matrix, it is used to map the original features to... 3D embedding space, It is represented as a bias vector. This is represented as the embedded representation obtained after mapping, with each input frame corresponding to a segment. , This is represented by a hyperbolic tangent nonlinear activation function, used to enhance the nonlinearity of the expression and constrain the output range, which is beneficial for stable training. Then, by applying attention weighting to the embedding vector, the relative importance of each spatial and temporal location in the current recognition task is calculated, ensuring that the machine learning model can automatically focus on key areas containing residue edges, liquid layers, and lesions in the system, thereby improving the sensitivity and specificity of intestinal image recognition. The specific formula is expressed as follows:
[0035] in, This represents the number of embedding vectors collected at the current moment. Represented as the first Embedded vectors, Represented as the index of the number of embedding vectors, Represented as a query vector, it is used to measure the relevance of the embedding to the target being identified. This is expressed as a temperature coefficient, used to adjust the smoothness of the softmax distribution. Represented as the first Each embedded attention weight, The attention-weighted context representation vector serves as input for subsequent classification and segmentation. The attention-weighting strategy, within a machine learning framework, highlights the most critical regions and image frames for assessing bowel preparation quality. This allows the algorithm to achieve higher information density with lower computational cost, providing clearer evidence frames for bowel image recognition to support precise intervention decisions. Secondly, by inputting the attention representation into the discriminant network, the probability distribution of each bowel preparation candidate region belonging to various types of residues and background is calculated, serving as the basis for subsequent scoring and intervention decisions. This transforms continuous image representations into corresponding probability outputs, enabling the machine learning model to quantify the presence of various objects in the intestine, including key image data of residues, liquid layers, mucus, and healthy mucosa. The specific formula is as follows:
[0036] in, This is represented as the weight matrix of the discriminant network. This is represented as the bias vector of the discriminant network. Represented as a category number, including residue, liquid layer, mucus, and healthy mucous membrane. This is expressed as an element-wise exponential normalization function, making the output a probability distribution. This represents the probability vector for each category. The calculated probability vectors provide structured and interpretable category evidence for subsequent image recognition. Then, by fusing the probability vectors obtained from different image frames according to weights based on expert experience, a more robust instantaneous scoring basis is obtained. This ensures that heterogeneous evidence is fused in the system to reduce the impact of single-frame noise on the results of intestinal preparation image recognition. The specific formula is expressed as follows:
[0037] in, This represents the number of observed image frames involved in the fusion. Represented as the first The class probability vector obtained from each observation Represented as the index of the number of observed image frames participating in the fusion. Represented as the first The fusion weights for each observation are set according to a confidence decay strategy. The probability vector after fusion is represented as the instantaneous comprehensive distribution score of the intestinal preparation region. Secondly, by constructing a graph structure based on region adjacency and applying graph Laplacian smoothing to the fused probability vector, isolated misjudgments are suppressed and local consistency is enhanced, ensuring spatial coherence of the intestinal preparation score in the algorithm, thereby reducing the risk of false triggering of unnecessary interventions. The specific formula is expressed as follows:
[0038] in, Represented as an identity matrix, Represented as the smoothing intensity coefficient, Represented as a graph Laplacian matrix, defined as follows: ,in It is an adjacency matrix. For degree matrix, The smoothed probability vector represents the image recognition correction result with spatial consistency. Finally, by applying class determination and confidence calculation decision rules to the smoothed probability vector, a structured output is generated that can be used by subsequent modules, including the class label and confidence score of each intestinal preparation image. The specific formula is as follows:
[0039] in, This represents the category index used for the final determination. Represented as corresponding to The probability vector after category smoothing Represented as the index of the number of categories, This is expressed as the calibrated confidence level value. This is represented as a scaling parameter, used to adjust the impact of the original highest probability on the calibration output. Represented as an offset parameter, used to correct systematic offsets. Represented as the Sigmoid function, it maps the linearly transformed value to the (0,1) interval, making it convenient as the final interpretable confidence level.
[0040] See Figure 1 Furthermore, the bowel preparation scoring and intervention module includes a bowel preparation scoring unit and a precision intervention decision unit. The bowel preparation scoring unit proposes a deep learning-based bowel preparation scoring algorithm. Through the constructed deep learning model, it maps the identified types and features and outputs segmented scores, total bowel preparation scores, and confidence estimates, ensuring that the scoring has high accuracy, stability, and reliability for clinical decision-making. The precision intervention decision unit generates personalized and actionable intervention suggestions by inputting the scoring results, patient compliance records, past medical history, and clinical pathway rules into a decision rule base, ensuring that the intervention not only conforms to clinical norms but also improves the quality of bowel preparation and patient safety.
[0041] See Figure 1 Furthermore, the deep learning-based gut preparation scoring algorithm is as follows: First, the output of each frame's image category and corresponding confidence score from the image recognition unit is converted into a continuous semantic vector. This allows the subsequent deep learning module to understand the semantic information of each frame using a unified and learnable vector representation, thereby supporting accurate gut preparation scoring. By fusing discrete category evidence and confidence scores into feature representations that can participate in deep network training, the algorithm's sensitivity and stability to local image evidence are improved. The specific formula is expressed as follows:
[0042] in, This represents the index of the current image frame. This represents the total number of categories determined by the image recognition unit, including residue, liquid layer, mucus, and healthy mucous membrane. This represents the category index determined by the image recognition unit. Represented as an image recognition unit for the first Frame number Confidence value of the category, Represented as the first Learned embedding vectors of categories, Represented as the constructed first Frame semantic vectors serve as the basic input to the deep learning scoring network. Then, the semantic vectors of consecutive frames are sequentially input into a deep temporal encoder to obtain a contextual representation with short- and medium-temporal information. This allows the deep learning gut health scoring system to enhance single-frame judgments using inter-frame dynamics. The specific formula is as follows:
[0043] in, Represented as the timing window length, This represents the vector concatenation operation. Represented as a weight matrix in time-series coding. Represented as the real number field, Represented as the number of dimensions, the concatenated... Projecting a dimensional vector to Vygote indicates that, It is represented as a bias vector. Represented as a nonlinear activation, it provides a smooth, bounded representation. Represented as the first The frame temporal context encoding representation serves as the intermediate semantic input to the scoring network. By capturing dynamic cues, the algorithm ensures that the final score is based not only on single-frame evidence but also on trends in neighboring frames. This reduces false alarms and improves response accuracy in real-time monitoring and alarm scenarios. Then, a deep fully connected transformation is applied to the temporal context representation to extract combined features highly correlated with bowel preparation scores. This allows the algorithm to learn the complex nonlinear mapping relationship between image sequences and clinical scores, providing a trainable feature base for the final continuous scoring. The specific formula is as follows:
[0044] in, Represented as the linear transformation weights of a deep scoring network. This is represented as the corresponding bias vector. Represented as the rectified linear unit activation function, it is used to introduce sparsity and nonlinearity. Represented as the function that maximizes, The hidden features, after transformation, carry high-order information directly related to the score. Through deep transformation, the algorithm can learn complex discriminative structures in a high-dimensional feature space, thereby enhancing its ability to distinguish between different bowel preparation qualities and laying the foundation for scoring accuracy. Secondly, the hidden features are mapped to continuous bowel preparation score values to directly output numerical scores that can be used for clinical interpretation and system decision-making. The specific formula is as follows:
[0045] in, Represented as a rating mapping vector, This is represented as a scalar bias. Represented as the Sigmoid function, Represented as an exponential function, Represented as variables, Represented as the first The continuous bowel preparation score of each frame, where 1 represents the worst score and 5 represents the best score, is directly output as a clinically usable continuous score through end-to-end deep mapping. This facilitates threshold judgment, trend monitoring, and visualization within the system. Then, based on the category distribution provided by the image recognition unit, the frame-level information entropy is calculated to estimate the reliability of the frame score, providing a quantitative basis for subsequent weighted aggregation and intervention decisions. The specific formula is as follows:
[0046] in, This is represented as the first step in step one. Frame number Confidence of the category Represented as a numerically stable term, Represented as information entropy, it represents a measure of uncertainty in frame category. This is represented by the maximum value of entropy. This is represented as frame-level confidence, with values closer to 1 indicating a more certain category distribution. Confidence estimation provides a quantification of uncertainty for deep learning-based bowel preparation scoring, enabling more rigorous reliability testing for automated interventions within the system. Low-confidence frames can be transferred to manual review to ensure clinical safety. Finally, the continuous scores of each frame within the examination sequence are weighted and averaged based on frame-level confidence to form the final bowel preparation score for the entire case. This provides clear numerical evidence for the precision intervention decision-making unit, while retaining evidence from each frame for auditing and manual review. The specific formula is as follows:
[0047] in, This represents the total number of frames used for scoring throughout the entire inspection. The weighted average bowel preparation score, representing the entire case, serves as the main quantitative indicator of the system output. By employing a confidence-weighted aggregation strategy, the influence of high-confidence frames is preferentially retained, while the interference of noisy frames on the final score is reduced. This allows the deep learning bowel preparation score to more reliably drive clinical decision-making within the system.
[0048] See Figure 1 Furthermore, the intestinal intelligent monitoring module performs real-time trend analysis, anomaly detection, and multi-channel alarms on the scoring time series, identification confidence level, and intervention suggestions to ensure that the system can identify abnormal and low-confidence scenarios, promptly detect abnormal intestinal preparation, and trigger corresponding alarm notifications.
[0049] See Figure 1 Furthermore, the report display and quality management module automatically generates visual reports that include key areas of image recognition, segmented scores, total scores, intervention records, and confidence levels. It also provides an interactive review interface and statistics on intestinal preparation quality trends, ensuring that users can easily manage and control the quality.
[0050] In practical use, firstly, the bowel preparation imaging acquisition module is used to acquire bowel preparation image data, ensuring the acquisition of high-definition, synchronous, multi-view images with complete metadata for subsequent module analysis; secondly, the image processing and recognition module includes an image feature extraction unit and an image recognition unit. The image feature extraction unit is used to preprocess and structure the acquired images, extracting key quantitative features such as color, texture, shape, and residue coverage from the preprocessed images to ensure a stable and reliable input feature vector for image recognition. The image recognition unit proposes a machine learning-based bowel image recognition algorithm to perform target detection, segmentation, and classification of bowel images, ensuring robust residue identification and site localization even under complex field of view; then, the bowel preparation scoring and intervention module includes a bowel preparation assessment... The system comprises several sub-units, including a bowel preparation scoring unit, a deep learning-based bowel preparation scoring algorithm, and a confidence level assessment unit. The latter uses a deep learning-based algorithm to intelligently score bowel preparation quality and output confidence levels, ensuring high accuracy, repeatability, and interpretability. The precise intervention decision-making unit combines the scoring results with patient background, compliance, and clinical rules to generate personalized intervention strategies, ensuring that intervention recommendations are scientific, safe, and easy to implement and record. Secondly, a bowel intelligent monitoring module monitors the scoring results and intervention strategies in real time and triggers abnormal alerts, ensuring timely detection of low-quality bowel preparation. Finally, a report display and quality management module generates clinical reports containing image evidence, scoring, and intervention records, and statistically analyzes bowel preparation quality trends, ensuring traceability, ease of quality control, and system management.
[0051] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A system for intelligent monitoring and precise intervention of bowel preparation quality based on image recognition, comprising a multimodal imaging acquisition module, an image processing and recognition module, a bowel intelligent monitoring module, a bowel preparation scoring and intervention module, and a report display and quality management module. The bowel preparation imaging acquisition module is used to acquire bowel preparation image data, ensuring the acquisition of high-definition, synchronous, multi-view images with complete metadata for subsequent module analysis. The image processing and recognition module includes an image feature extraction unit and an image recognition unit. The image feature extraction unit is used to preprocess and structure the acquired images, extracting key quantitative features such as color, texture, shape, and residual coverage from the preprocessed images. The image recognition unit extracts... A machine learning-based intestinal image recognition algorithm is used for target detection, segmentation, and classification of intestinal images. The intestinal preparation scoring and intervention module includes an intestinal preparation scoring unit and a precise intervention decision unit. The intestinal preparation scoring unit proposes a deep learning-based intestinal preparation scoring algorithm to intelligently score the quality of intestinal preparation and output confidence scores. The precise intervention decision unit is used to combine the scoring results with patient background, compliance, and clinical rules to generate personalized intervention strategies. The intelligent intestinal monitoring module is used to monitor the scoring results and intervention strategies in real time and trigger abnormal alarms. The report display and quality management module is used to generate clinical reports containing image evidence, scoring, and intervention records and to statistically analyze trends in intestinal preparation quality.
2. The intelligent monitoring and precise intervention system for bowel preparation quality based on image recognition according to claim 1, characterized in that: The intestinal preparation imaging acquisition module integrates intestinal preparation images acquired by medical devices and performs timestamp synchronization, resolution standardization, and metadata recording on the acquired data.
3. The intelligent monitoring and precise intervention system for bowel preparation quality based on image recognition according to claim 1, characterized in that: The image processing and recognition module includes an image feature extraction unit and an image recognition unit. The image feature extraction unit performs preprocessing such as cleaning and denoising on the original image, and calculates local and global multi-scale features, including color histograms, texture, morphology, edge and depth embedding, and time-frame level statistics. This ensures that robust feature vectors that are highly distinguishable between residues and mucosal differences are extracted from complex lighting and pollution conditions, providing robust input for recognition and scoring. The image recognition unit proposes a machine learning-based intestinal image recognition algorithm. By constructing a target detection network under a machine learning model, it classifies and labels image features, distinguishes the intestinal mucosa, residue types and distribution, and outputs quantifiable regional indicators.
4. The intelligent monitoring and precise intervention system for bowel preparation quality based on image recognition according to claim 3, characterized in that, Linear transformation and nonlinear activation are applied to the feature vectors from the image feature extraction unit to construct an embedding representation adapted to the subsequent machine learning module, so as to unify the feature scale and improve the representation ability of intestinal image recognition task. By projecting the original feature vectors onto the feature space of the predetermined classification, the machine learning ability to distinguish different targets such as residue, liquid and mucus in the intestine is enhanced.
5. The intelligent monitoring and precise intervention system for bowel preparation quality based on image recognition according to claim 4, characterized in that, A linear transformation and nonlinear activation are applied to the feature vectors from the image feature extraction unit, as expressed by the following formula: in, This is represented as an input feature vector, derived from the image feature extraction unit. Represented as the real number field, and Represented as feature dimension, Represented as an embedding weight matrix, it is used to map the original features to... 3D embedding space, It is represented as a bias vector. This is represented as the embedded representation obtained after mapping, with each input frame corresponding to a segment. , Represented as a hyperbolic tangent nonlinear activation function; By applying attention weights to the embedding vectors, the relative importance of each spatial and temporal location in the current identification task is calculated, ensuring that the machine learning model can automatically focus on key areas containing residue edges, liquid layers, and lesions in the system. The specific formula is expressed as follows: in, This represents the number of embedding vectors collected at the current moment. Represented as the first Embedded vectors, Represented as the index of the number of embedding vectors, Represented as a query vector, it is used to measure the relevance of the embedding to the target being identified. This is expressed as a temperature coefficient, used to adjust the smoothness of the softmax distribution. Represented as the first Each embedded attention weight, It is represented as an attention-weighted context representation vector, which serves as the input for subsequent classification and segmentation. The attention-weighted strategy highlights the most critical regions and image frames for judging the quality of intestinal preparation in the image within the machine learning framework. By inputting the attention representation into the discriminative network, the probability distribution of each intestinal preparation candidate region belonging to various types of residues and background is calculated. This distribution serves as the basis for subsequent scoring and intervention decisions. Continuous image representations are converted into corresponding probability outputs, enabling the machine learning model to quantify the presence of various objects in the gut, including key image data of residues, liquid layers, mucus, and healthy mucosa. The specific formula is expressed as follows: in, This is represented as the weight matrix of the discriminant network. This is represented as the bias vector of the discriminant network. Represented as a category number, including residue, liquid layer, mucus, and healthy mucous membrane. This is expressed as an element-wise exponential normalization function, making the output a probability distribution. This represents the probability vector for each category; By fusing probability vectors obtained from different image frames according to weights based on expert experience, a more robust instantaneous scoring basis is obtained. The specific formula is expressed as follows: in, This represents the number of observed image frames that participated in the fusion. Represented as the first The class probability vector obtained from each observation Represented as the index of the number of observed image frames participating in the fusion. Represented as the first The fusion weights for each observation are set according to a confidence decay strategy. It is represented as a fused probability vector, serving as the instantaneous comprehensive distribution score of the intestinal preparation region; By constructing a graph structure based on region adjacency and applying graph Laplacian smoothing to the fusion probability vector, isolated misjudgments are suppressed and local consistency is enhanced, ensuring spatial coherence of the gut preparation score in the algorithm. The specific formula is expressed as follows: in, Represented as an identity matrix, Represented as the smoothing intensity coefficient, Represented as a graph Laplacian matrix, defined as follows: ,in It is an adjacency matrix. For degree matrix, The smoothed probability vector represents the image recognition correction result with spatial consistency. Finally, by applying class determination and confidence calculation decision rules to the smoothed probability vector, a structured output is generated that can be used by subsequent modules, including the class label and confidence score of each intestinal preparation image. The specific formula is as follows: in, This represents the category index used for the final determination. Represented as corresponding to The probability vector after category smoothing Represented as the index of the number of categories, This is expressed as the calibrated confidence level value. This is represented as a scaling parameter, used to adjust the impact of the original highest probability on the calibration output. Represented as an offset parameter, used to correct systematic offsets. Represented as the Sigmoid function, it maps the linearly transformed value to the (0,1) interval, making it convenient as the final interpretable confidence level.
6. The intelligent monitoring and precise intervention system for bowel preparation quality based on image recognition according to claim 1, characterized in that: The bowel preparation scoring and intervention module includes a bowel preparation scoring unit and a precision intervention decision unit. The bowel preparation scoring unit proposes a deep learning-based bowel preparation scoring algorithm. Through the constructed deep learning model, it maps the identified types and features and outputs segmented scores, total bowel preparation scores, and confidence estimates, ensuring that the scoring has high accuracy, stability, and reliability for clinical decision-making. The precision intervention decision unit generates personalized and actionable intervention suggestions by inputting the scoring results, patient compliance records, past medical history, and clinical pathway rules into a decision rule base, ensuring that the intervention not only conforms to clinical norms but also improves the quality of bowel preparation and patient safety.
7. The intelligent monitoring and precise intervention system for intestinal preparation quality based on image recognition according to claim 6, characterized in that, The output of each frame image category and corresponding confidence score from the image recognition unit is converted into a continuous semantic vector so that the subsequent deep learning module can understand the semantic information of each frame with a unified and learnable vector representation, thereby supporting accurate gut preparation scoring. By fusing discrete category evidence and confidence scores into feature representations that can participate in deep network training, the algorithm's sensitivity and stability to local image evidence are improved.
8. The intelligent monitoring and precise intervention system for bowel preparation quality based on image recognition according to claim 7, characterized in that, The output of each image category and corresponding confidence level from the image recognition unit is converted into a continuous semantic vector, as expressed by the following formula: in, This represents the index of the current image frame. This represents the total number of categories determined by the image recognition unit, including residue, liquid layer, mucus, and healthy mucous membrane. This represents the category index determined by the image recognition unit. Represented as an image recognition unit for the first Frame number Confidence value of the category, Represented as the first Learned embedding vectors of categories, Represented as the constructed first Frame semantic vectors serve as the basic input to deep learning scoring networks; The semantic vectors of consecutive frames are sequentially input into a deep temporal encoder to obtain a contextual representation with short- and medium-temporal information. This allows the deep learning gut health score to enhance single-frame judgment using inter-frame dynamics. The specific formula is as follows: in, Represented as the timing window length, This represents the vector concatenation operation. Represented as a weight matrix in time-series coding. Represented as the real number field, Represented as the number of dimensions, the concatenated... Projecting a dimensional vector to Vygote indicates that, It is represented as a bias vector. Represented as a nonlinear activation, it provides a smooth, bounded representation. Represented as the first Frame timing context encoding representation; A deep fully connected transformation is applied to the temporal context representation to extract combinatorial features highly correlated with bowel preparation scores. This enables the algorithm to learn the complex nonlinear mapping relationship between image sequences and clinical scores, providing a trainable feature base for the final continuous scoring. The specific formula is as follows: in, Represented as the linear transformation weights of a deep scoring network. This is represented as the corresponding bias vector. Represented as the rectified linear unit activation function, it is used to introduce sparsity and nonlinearity. Represented as the function that maximizes, This is represented as the transformed hidden feature representation; The hidden features are mapped to continuous bowel preparation score values to directly output numerical scores that can be used for clinical interpretation and system decision-making. The specific formula is as follows: in, Represented as a rating mapping vector, This is represented as a scalar bias. Represented as the Sigmoid function, Represented as an exponential function, Represented as variables, Represented as the first The continuous bowel preparation score of the frame, where 1 represents the worst score and 5 represents the best score; Frame-level information entropy is calculated based on the category distribution given by the image recognition unit to estimate the credibility of the frame's score, providing a quantitative basis for subsequent weighted aggregation and intervention decisions. The specific formula is as follows: in, This is represented as the first step in step one. Frame number Confidence of the category Represented as a numerically stable term, Represented as information entropy, it represents a measure of uncertainty in frame category. This is represented by the maximum value of entropy. This is expressed as frame-level confidence; the closer the value is to 1, the more certain the category distribution. The continuous scores of each frame within the examination sequence are weighted and averaged based on frame-level confidence to form the final bowel preparation score at the whole case level. This provides clear numerical evidence for the precision intervention decision-making unit, while retaining evidence from each frame for auditing and manual review. The specific formula is as follows: in, This represents the total number of frames used for scoring throughout the entire inspection. This is represented as the weighted average bowel preparation score for the entire case.
9. The intelligent monitoring and precise intervention system for bowel preparation quality based on image recognition according to claim 1, characterized in that: The intestinal intelligent monitoring module performs real-time trend analysis, anomaly detection, and multi-channel alarms on the scoring time series, identification confidence level, and intervention suggestions to ensure that the system can identify abnormal and low-confidence scenarios, promptly detect abnormal intestinal preparation, and trigger corresponding alarm notifications.
10. The intelligent monitoring and precise intervention system for intestinal preparation quality based on image recognition according to claim 1, characterized in that: The report display and quality management module automatically generates visual reports that include key areas identified in the image, segmented scores, total scores, intervention records, and confidence levels, and provides an interactive review interface and statistics on bowel preparation quality trends.
Citation Information
Cited By
Screening and intervention system and method for intestinal preparation before outpatient colonoscopy
CN122177346A
An enteroscope video detection method and device based on quality stability joint evaluation
CN122347584A