Automated detection and differentiation of small bowel lesions during capsule endoscopy
By combining deep learning and transfer learning, a multi-layer convolutional neural network is used to detect and classify lesions in capsule endoscopy images, solving the problems of time consumption and misdiagnosis in existing technologies. This achieves rapid and accurate detection and classification of small intestinal lesions, thus optimizing the diagnostic efficiency of capsule endoscopy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DIGESTAID ARTIFICIAL INTELLIGENCE DEV LDA
- Filing Date
- 2021-11-18
- Publication Date
- 2026-05-26
AI Technical Summary
Existing capsule endoscopy is time-consuming and error-prone in detecting and classifying small bowel lesions, and existing machine learning technologies lack sufficient accuracy in clinical practice, leading to inappropriate treatment.
Deep learning methods, combined with transfer learning and semi-active learning, were employed to extract features and classify capsule endoscopy images using a multi-layer convolutional neural network. The bleeding potential of lesions was assessed using the Saurin classification system, and the training process was optimized to improve accuracy.
It enables rapid and accurate detection and classification of small bowel lesions, improves the accuracy of clinical diagnosis and treatment, and optimizes the diagnostic efficiency of capsule endoscopy.
Smart Images

Figure CN116830148B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to lesion detection and classification in medical image data. More specifically, it relates to the automated identification of small bowel lesions in capsule endoscopy images to assess lesion severity and subsequent medical treatment. Background Technology
[0002] Capsule endoscopy has become the primary endoscopic method for small bowel examination. By carefully examining video frames from the capsule, physicians can detect, identify, and characterize lesions in the gastrointestinal tract lining. However, this video capsule endoscopy is very time-consuming for gastroenterologists and is prone to human error and oversight. In contrast, with capsule endoscopy, the recording of such images is readily available and digitally stored for retrospective examination and comparison. In this context, image data provides a robust foundation for computer-aided diagnosis using machine learning systems for lesion characterization and subsequent decision-making. The goal of lesion detection and classification in the small bowel is to generate a more accurate and thorough automated characterization of lesion severity, contributing to medical diagnosis and treatment.
[0003] Valério, Maria Teresa et al., in "Lesions Multiclass Classification in Endoscopic Capsule Frames." Procedia Computer Science 164(2019):637-645, improved medical experts' understanding of the time-consuming and error-prone identification of small bowel lesions. Furthermore, the authors proposed an automated method for identifying these lesions using a deep learning network based on medically annotated wireless capsule endoscopy images.
[0004] Li, Xiuli, et al., in "Exploring transfer learning for gastrointestinal bleeding detection on small-size imbalanced endoscopy images," 2017, 39th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC). IEEE, 2017, explored transfer learning for gastrointestinal bleeding detection on small-size imbalanced endoscopy images using neural networks on large image representations.
[0005] Reference CN 110705440 filters images into three channels and feeds them into a convolutional network model trained on the Kvasir dataset.
[0006] The paper US 2020 / 065961 A1 uses convolutional networks to classify data into different particle scales to detect lesions in biological tissues. The data requirement is to identify defect / damage regions in images.
[0007] Document CN 111127412A discloses a pathological image recognition device based on generative game theory networks, aiming to address the problems of existing pathological recognition methods, such as reliance on experience, time-consuming manual annotation costs, low recognition efficiency, and poor accuracy. This method is used to assess Crohn's disease lesions, but does not distinguish the type of lesion presented in each image.
[0008] Document CN 107730489A discloses a wireless capsule endoscope with a small bowel lesion detector for effective and accurate detection of small bowel lesions. The classification and localization of small bowel lesions are achieved through an image segmentation algorithm. This document does not use transfer learning, does not distinguish lesion types, and does not actively update the training data with new data for the next training iteration.
[0009] Document CN 111739007A discloses a bubble region recognizer trained by a convolutional neural network, which does not take into account the classification of small intestinal lesions.
[0010] Gastrointestinal diseases, such as lesions within the small intestine, are epidemiologically common. Often, when these lesions are not diagnosed and treated, they can lead to unfavorable clinical outcomes. Gastrointestinal endoscopy plays a crucial role in the diagnosis and treatment of digestive cancers, inflammatory bowel disease (IBD), and other gastrointestinal conditions. For example, if colorectal polyps are detected early, they can be safely removed, preventing colorectal cancer. Therefore, all polyps must be detected. Another example of the critical role of gastrointestinal endoscopy is the endoscopic assessment of IBD activity by detecting ulcers and erosions in the small and colonic intestines. Assessing IBD activity is essential for the diagnosis, treatment, and management of IBD patients. However, gastrointestinal endoscopy is a time-consuming, labor-intensive, and challenging task. Sometimes, endoscopists may exhibit signs of fatigue or attention deficit, failing to accurately identify all relevant endoscopic findings. More specifically, capsule endoscopy is a minimally invasive option for assessing the gastrointestinal tract, particularly for small bowel endoscopy. In fact, capsule endoscopy is the gold standard and first-line examination for assessing the small intestine (i.e., unexplained gastrointestinal bleeding and small bowel pathology).
[0011] Endoscopic image acquisition is a state-of-the-art technique for understanding a patient's intestines. Typically, endoscopic components are equipped with portable image recording devices and a mechanism to convert these captures into a digital representation and store them on a personal computer.
[0012] Due to the nature of endoscopic image acquisition, light or other photographic conditions are often lacking to allow for direct small bowel sorting. Against this backdrop, machine learning techniques have been proposed to automate such tasks; however, recent machine learning techniques have failed to deliver the overall accuracy or false negative rate usable in clinical practice, thus leading to inappropriate treatment. Summary of the Invention
[0013] This invention provides a method for accurately distinguishing common lesions of the small intestine in endoscopic images based on deep learning. This automated identification, classification, and bleeding risk estimation of small intestinal lesions can be used in clinical practice for diagnosis and treatment planning.
[0014] Initially, by using ImageNet 1 Pre-trained convolutional layers of a given architecture were used to further test the capsules using a subset of endoscopic images, demonstrating the potential for lesion detection. The clinically disruptive nature of this invention is supported by the capabilities of an artificial intelligence system, enabling it to not only detect but also accurately differentiate all relevant endoscopic findings / lesions. In fact, the ability of neural networks to distinguish subtle polymorphic lesions is paramount in clinical practice, allowing for complete capsule endoscopy diagnosis.
[0015] Furthermore, the stratification of bleeding potential for each lesion is a relevant novelty introduced by this invention into the prior art. One of the most important and common indications for capsule endoscopy is unexplained gastrointestinal bleeding, and accurate assessment of the bleeding potential of endoscopic findings is crucial for clinical follow-up management. Therefore, by accurately predicting the bleeding potential of capsule endoscopic findings / lesions, this invention helps clinical teams better define patient diagnosis and treatment management, which can translate into optimized clinical outcomes.
[0016] This technology can be applied to capsule endoscopy software to help gastroenterologists detect small bowel lesions. Furthermore, capsule endoscopy is an expensive procedure, but it is gaining increasing relevance and application in clinical practice.
[0017] The following is considered to be related to highlighting the problem addressed by the present invention from methods known in the art for detecting and differentiating small bowel lesions in capsule endoscopy.
[0018] A preferred approach is to classify images using machine learning techniques. Deep learning uses algorithms to model high-level abstractions in data using deep graphs with multiple processing steps. Using a multi-layered architecture, machines employing deep learning techniques process the raw data to find highly correlated values or groups that distinguish topics.
[0019] This method detects relevant small bowel lesions in capsule endoscopy images and distinguishes them based on their bleeding potential using the Saurin classification system. The Saurin classification system measures the bleeding potential of small bowel lesions. It is a useful tool for patient assessment and treatment strategies. Its use has a direct impact on clinicians' diagnosis and decision-making. Such embodiments of the invention utilize transfer learning and semi-active learning. Transfer learning allows for feature extraction and high-accuracy classification using reasonable dataset sizes. Semi-active implementation allows for continuous improvement of the classification system. A preferred embodiment of the invention preferably utilizes transfer learning for feature extraction on capsule endoscopy images, which can be based on either the Saurin classification system used for capsule endoscopy images or a semi-active learning strategy.
[0020] The method then divides the dataset into multiple stratified folds, preferably in which images for a given patient are included in only one fold. Additionally, or alternatively, the data are trained and validated by grouping patients into random folds, i.e., images from any patient belong to either the training or validation set.
[0021] A preferred approach is to further train a combination of network architectures using selected training and validation sets, particularly including feature extraction and classification components. The range of convolutional neural networks to be trained includes, but is not limited to, VGG16, InceptionV3, Xception, EfficientNetB5, EfficientNetB7, ResNet50, and ResNet125. Preferably, their weights are frozen, except for batch normalization layers, and coupled to the classification component. The classification component comprises at least two densely connected layers, preferably 2048 and 1024 in size, and at least one dropout layer, preferably 0.1 inches in size, between them.
[0022] Alternatively, but not preferably, the sorting component can be used with more densely connected layers or with densely connected layers of different sizes. Alternatively, but not preferably, the sorting component can also be used without discard layers.
[0023] Furthermore, the optimal execution architecture is selected based on overall accuracy and sensitivity. Performance metrics include, but are not limited to, the F1 score. Additionally, the method preferably uses the optimal execution architecture to train a series of classification component combinations, which preferably include, but are not limited to, two to four densely connected layers in sequence starting at 4096 and halving down to 512. A dropout layer with a dropout rate of 0.1 is placed between the last two layers.
[0024] Finally, the optimal execution solution is trained using the entire dataset with patient groupings. Other embodiments of the invention may include similar classification networks, training weights, and hyperparameters.
[0025] These can include the use of any image classification network, whether new or yet to be designed.
[0026] Typically, this method comprises two modules that provide the necessary data to the remaining modules: a prediction collector and an output collector. The prediction collector reads the video and selects images with discoveries. The output collector passes these images with discoveries for processing.
[0027] Examples of the advantages of this invention include: training parameters using machine learning results from a daily increasing dataset based on the cloud; automatically predicting endoscopic images using deep learning methods to identify and differentiate small bowel lesions from capsule endoscopy image inputs based on the Saurin classification system; and improving image classification speed and corresponding accuracy through the use of transfer learning. Attached Figure Description
[0028] Figure 1 A method for classifying small bowel lesions during capsule endoscopy according to an embodiment of the present invention is shown.
[0029] Figure 2 A method for automated detection and differentiation of small bowel lesions during capsule endoscopy is shown.
[0030] Figure 3 The main process for automatically detecting and differentiating small bowel lesions in capsule endoscopy is shown.
[0031] Figure 4 Explain the structure of the classification network that distinguishes based on bleeding potential.
[0032] Figure 5 One implementation of a classification network based on bleeding potential is described.
[0033] Figure 6A preferred embodiment of the invention is illustrated, showing accuracy curves trained on a small subset of image and labeled data. An example of the results from iterations of method 8000 is shown.
[0034] Figure 7 Exemplary accuracy curves during training on a small subset of image and label data, according to an embodiment of the present invention, are shown. Examples are derived from the results of iterations of method 8000.
[0035] Figure 8 Exemplary ROC curves and AUC values obtained after training on a small subset of image and label data according to an embodiment of the present invention are shown. Results are used for model selection. An example of the results from iterations of method 8000 is shown, along with scaling of the ROC curves.
[0036] Figure 9 An exemplary confusion matrix is shown after training on a small subset of image and label data according to an embodiment of the present invention. The result is used for model selection. The number of images in the small subset of data and the proportion of the corresponding classes are shown in parentheses.
[0037] Figure 10 An example of lesion classification according to an embodiment of the present invention is shown.
[0038] Figure 11 The results of performing deep learning-based lesion classification on data volumes 240 and 250 according to an embodiment of the present invention are shown.
[0039] Figure 12 Examples of classified lesions awaiting expert verification are shown. Detailed Implementation
[0040] This invention discloses a novel method and system for detecting and differentiating lesions in images acquired during capsule endoscopy.
[0041] Some preferred embodiments will be described in more detail with reference to the accompanying drawings, in which embodiments of the invention are illustrated. However, the invention can be implemented in various ways and should not be construed as limited to the embodiments disclosed herein.
[0042] It should be understood that while this invention includes a detailed description of cloud computing, implementations of the teachings cited herein are not limited to cloud computing environments. Rather, embodiments of the invention can be implemented in conjunction with any other type of computing environment now known or developed in the future.
[0043] The term "deep learning" is a machine learning technique that uses multiple data processing layers to classify datasets with high accuracy. It can be a trained network (model or device) that learns from multiple inputs and outputs. A deep learning network can be a deployed network (model or device) generated from the trained network and provides output responses to inputs.
[0044] The term "supervised learning" is a deep learning training method in which the machine is provided with data that has already been classified from human sources. In supervised learning, features are learned from labeled input.
[0045] The term "convolutional neural network" or "CNN" is a network used in interconnected deep learning to identify objects and regions in a dataset. CNNs evaluate the raw data in a series of stages to assess the learned features.
[0046] The term "transfer learning" refers to a machine that stores information learned when attempting to solve a problem in order to solve another problem of a similar nature to the first one.
[0047] We use the term "semi-active learning" to describe the process of machine learning. Before performing the next learning step, the network is trained by attaching a set of labeled data to a training dataset from trusted external entities. For example, as the machine collects more samples from dedicated worker steps, it is less likely to ignore images with similar characteristics.
[0048] The term "computer-aided diagnosis" refers to a machine that analyzes medical images to suggest possible diagnoses.
[0049] The term "lymphangiectasia" refers to the obstruction and dilation of capillaries. It can be functional (unrelated to pathology), primary, or secondary to other diseases. Based on previously published material, we define lymphangiectasia as scattered leukoplakia of the intestinal mucosa. These mucosal changes can be diffuse or patchy.
[0050] The term "xanthoma" refers to the accumulation of cholesterol-rich material in intestinal mucosal macrophages. Endoscopic findings during capsule endoscopy are defined as patchy lesions with a white / yellow appearance.
[0051] The terms “ulcer” and “erosion” refer to mucosal ruptures in the small intestinal mucosa. These lesions are distinguished based on the estimated size and depth of the penetration. An “ulcer” is defined as a depression in the epithelial lining, with a whitish base and surrounding swollen mucosa >5 mm in diameter. Conversely, a mucosal “erosion” is defined as a minimal loss of the epithelial layer surrounded by normal mucosa.
[0052] The term "vascular lesions" in the small intestine encompasses a variety of individual lesions, particularly erythema, vasodilation, varicose veins, and venous dilatation. Erythema is defined as flat, punctate lesions (<1 mm) with bright red areas within the mucosal layer, lacking a vascular appearance. Vasodilation is defined as a clearly defined bright red lesion within the mucosal layer consisting of tortuous and clustered capillary dilatations. Varicose veins are defined as the presence of serous venous dilatations. Venous dilatation is identified if bluish venous dilatations are detected within the normal submucosal region.
[0053] The term "protruding lesion" refers to lesions that bulge into the lumen of the small intestine. These lesions may have different causes, including polyps, epithelial tumors, subepithelial lesions, and nodules.
[0054] The term "blood" is used to describe the presence of bright blood occupying part or all of the intestinal lumen. It represents active or recent bleeding.
[0055] "Blood residue" refers to fragments or whole blood clots that appear as dark red or brown residue in the small intestine or adhere to the mucosa. These residues, once separated, represent previous bleeding.
[0056] This invention relates to a method for classifying small bowel lesions based on deep learning in capsule endoscopy images according to the bleeding potential of the small bowel lesions. Figure 1 Typically, embodiments of the present invention provide a visual understanding of deep learning-based small bowel lesion classification methods. Automated lesion classification of small bowel images during capsule endoscopy is a challenging task because lesions with different bleeding potentials have similar shapes and contrasts. Large variations in the small bowel preparation prior to capsule ingestion further complicate automated small bowel lesion classification. Although the automated training and classification time is fast (averaging 10 seconds for a test dataset of 2000 images), the output is not satisfactory for rapid diagnosis by experts.
[0057] The method includes an image acquisition module; a storage module; a training input module; a processing module; an examination input module; a training module; a prediction module; and an output collector module.
[0058] Image acquisition module 1000 receives examination input volumes from a capsule endoscopy provider. By way of example only, providers can be, but are not limited to, OMOM, Given, Mirocam, and Fujifilm. Images and corresponding labels are loaded onto storage module 2000. Storage module 2000 includes multiple classification network architectures 100, a trained convolutional network architecture 110, and hyperparameters for self-training. Storage module 2000 can be a local server or a cloud server. The storage module contains training input label data from capsule endoscopy images and metadata required to run processing module 3000, training module 4000, prediction module 5000, second prediction module 6000, and output collector module 7000. Input label data includes, but is not limited to, images and corresponding lesion classifications. Metadata includes, but is not limited to... Figure 4 Examples of classification network architectures 100, training convolutional neural network architectures 110, training hyperparameters, training metrics, fully trained models, and selected fully trained models are shown in the figure.
[0059] Before running optimized training in training module 4000, the image and label data in storage module 2000 are processed in processing module 3000. The processing module normalizes the images according to the deep model architecture for training in 3000 or evaluation in 4000. Upon manual or pre-arranged request, the processing module normalizes the image data in storage module 2000 according to the deep model to be run in training module 4000. Optionally, upon manual or pre-arranged request, the processing module generates data pointers to storage module 2000 to form some or all of the image and ground-truth labels required to run training module 3000. To prepare for each training session, the dataset is divided into multiple folds, with patient-specific images exclusively entering one fold and only one fold for both training and testing. The training set is split for model training to generate data pointers for all the image and ground-truth labels required to run training process 9000. k-folds are optionally applied to stratified grouping of patients in the training set to generate data pointers for partial images and ground truth labels required for the model validation process 8000 to run the training module 4000. Segmentation ratios and number of shares are available in the metadata of the storage module. Operators include, but are not limited to, users, convolutional neural networks trained to optimize k-folds, or computational routines simply suited to perform the task. As an example only, the dataset is split by patients, dividing it into 90% for training and 10% for testing. Optionally, the images selected for training can be divided into 80% for training and 20% for validation during training. A 5-fold (5-share, 5-fold) stratified grouping by patient is applied to the images selected for training. The processing module normalizes the examination volume data 5000 according to the deep model architecture for operation in the prediction module 5000, either manually or upon pre-request.
[0060] like Figure 2As shown, the training module 4000 includes a model validation process 8000, a model selection step 400, and a model training step 9000. The model validation section iteratively selects a combination of a classification architecture 100 and a convolutional network 110 to train a deep model for small bowel lesion classification. The classification network 100 has densely connected layers and dropout layers to classify small bowel lesions based on their bleeding potential. A neural convolutional network 110 trained on a large dataset is coupled to the classification network 100 to train the deep model 300. The deep model 300 is trained using partial training images 200 and ground truth labels 210. Performance metrics of the trained deep model 120 are calculated using multiple partial training images 220 and ground truth labels 230. The model selection step 400 is based on the calculated performance metrics. In process 310, the model training section 9000 trains the selected deep model architecture 130 using the entire dataset of training images 240 and ground truth labels 250. In prediction module 6000, the trained deep model 140 outputs a small bowel lesion classification 270 from a given evaluation image 260. Examination volume 5000, including images from capsule endoscopy videos, is the input to prediction module 6000. Prediction module 6000 classifies the image volume of examination volume 5000 using a best-performing trained deep model from 4000 (see [link to relevant documentation]). Figure 3 The output collector module 7000 receives the classified quantities and loads them into the storage module after verification is performed by a neural network or any other computing system suitable for performing the verification task, or optionally by a physician expert in gastrointestinal imaging.
[0061] Embodiments of the present invention provide a deep learning-based method for automatic lesion classification of small bowel images during capsule endoscopy. Compared to existing computer-based methods described above for lesion classification of capsule-replicated images, the method described herein achieves improved recognition quality and diagnostic usefulness. Existing methods for small bowel lesion classification do not utilize transfer learning or semi-active training. This improved classification model employs an optimized training process across multiple deep model architectures.
[0062] By way of example only, the present invention includes a server containing training results for an architecture in which training results from large cloud-based datasets such as, but not limited to, ImageNet, ILSVRC, and JFT are available. Architecture variations include, but are not limited to, VGG, ResNet, Inception, Xception, or Mobile, EfficientNets, etc. All data and metadata can be stored in a cloud-based solution or on a local computer. Embodiments of the present invention also provide various methods for faster deep model selection. Figure 2 A method for classifying small intestinal lesions using deep learning according to an embodiment of the present invention is shown. Figure 2 The method includes a pre-training phase 8000 and a training phase 9000. Training phase 8000 is performed, where it stops early on a small subset of data to select the optimal deep neural network for small bowel lesion classification from multiple combinations of convolutional and classification parts. For example, a classification network with two densely connected layers of size 512 is coupled with an Xception model to train on a random set generated by k-fold cross-validation and patient grouping. Another random set is selected as the test set.
[0063] In the optimization loop for (i) classification and transfer learning of the deep neural network; and (ii) training the combination of hyperparameters, the training process with early stopping and testing is repeated 8000 times on a random subset. The image feature extraction component of the deep neural network is not an architectural variant of the top layer accessible from the storage module. The layers of the feature extraction component remain frozen but are accessible through the aforementioned storage module during training. The batch normalization layer of the feature extraction component is unfrozen, thus allowing efficient training of the system with capsule endoscope imagers from cloud images that present different features. The classification component has at least two blocks, each containing a densely connected layer followed by a dropout layer. The last block of the classification component has a batch normalization layer followed by a densely connected layer whose depth dimension is equal to the number of lesion types to be classified.
[0064] The fitness of the optimization procedure is calculated to (i) guarantee the minimum precision and sensitivity for all classes defined by the threshold; (ii) minimize the difference between training, validation, and testing losses; and (iii) maximize the learning on the last convolutional layer. For example, if training shows evidence of overfitting, a combination of models with lower depths is selected for evaluation.
[0065] The training phase 9000 is applied to the best-performing deep neural network using the entire dataset.
[0066] A fully trained deep model 140 can be deployed to prediction module 6000. Subsequently, each evaluation image 260 is classified to output a lesion classification 270. The output collection module has means for communicating with other systems to perform expert verification and validation on the new prediction data volumes arriving at 270. Such communication means include a display module for user input, a fully trained neural network for decision-making, or any computationally programmable process for performing this task. The validated classifications are loaded into storage modules manually or upon scheduling requests to become part of the datasets required for running pipelines 8000 and 9000.
[0067] like Figure 5 As shown, the implementation of the classification network 100 can classify bleeding potential into N: normal, POL: lymphangiectasia, POX: xanthoma, P1E: erosion, P1PE: bleeding point, P1U e P2U: ulcer, P1PR and P2PR: protrusions, P2V: vessels, and P3: blood, which are shown and grouped accordingly. In a given iteration of method 8000 ( Figure 7 , Figure 8 and Figure 9 The optimization process described here uses accuracy curves, ROC curves, and AUC values, as well as a confusion matrix, from training on small subsets of image and label data.
[0068] Figure 8 Exemplary ROC curves and AUC values obtained after training on small subsets of image and labeled data are shown, where 10 (P3-AUC: 1.00), 12 (P0X-AUC: 0.99), 14 (P2U-AUC: 0.99), 16 (P1PE-AUC: 0.99), 18 (P1E-AUC: 0.99), 20 (N-AUC: 0.99), 22 (P1U-AUC: 0.99), 24 (P2V-AUC: 1.00), 26 (P1PR-AUC: 0.99), 28 (P2PR-AUC: 1.00), 30 (P0L-AUC: 0.99), and 32 represent random guesses.
[0069] Figure 9 This shows an exemplary confusion matrix after training on a small subset of image and label data. The results are used for model selection. The number of images in the small subset of data and the proportion of the corresponding classes are shown in parentheses.
[0070] Figure 10 An example of lesion classification according to an embodiment of the present invention is shown, wherein bleeding potential is present in 500; lymphangiectasia is present in 510; low bleeding potential is present in 520; high bleeding potential is present in 520; ulcer is present in 530; and low bleeding potential is present in 540; ulcer is present in 540.
[0071] Figure 11 The results of performing deep learning-based lesion classification on data volumes 240 and 250 according to an embodiment of the present invention are shown. The results of small intestine classification using the training method 8000 of the present invention are significantly improved compared to the results using existing methods (without method 8000).
[0072] Figure 12 An example of a classified lesion to be verified by the output collector module 7000 is shown. As an example only, a physician specializing in gastrointestinal imaging identifies small bowel lesions by analyzing labeled images classified by the depth model 140.
[0073] Figure 5 The document describes options for reclassifying images on the last layer of the classification network 100. Optionally, a confirmation or reclassification is sent to the storage module.
[0074] The foregoing detailed description should be understood as illustrative and exemplary in all respects, and not restrictive, and the scope of the invention disclosed herein is determined not by the detailed description, but by the claims as interpreted according to the full scope permitted by patent law. It should be understood that the embodiments shown and described herein are merely illustrative of the principles of the invention, and various modifications can be made by those skilled in the art within the scope of the appended claims.
Claims
1. A computer-implemented method for automatically identifying and characterizing small bowel lesions in medical images obtained from capsule endoscopy by classifying pixels as lesions or non-lesions, wherein, The method includes: - Select multiple subsets of all capsule endoscopy images, with each subset considering only images from the same patient; - Select another subset as the verification set, wherein the subset does not overlay the selected images on the previously selected subset; - One of several combinations of convolutional neural network image feature extraction components, followed by a subsequent classification neural network component for classifying pixels as small bowel lesions, is pre-trained (8000) on each of the selected subsets, wherein the pre-training: o Stop early; o Evaluate the performance of each of the combinations; o Repeat on a new, different subset, along with another network combination and training hyperparameters, where if the f1 metric is low, this new combination considers a higher number of densely connected layers, and if the f1 metric suggests overfitting, it considers fewer densely connected layers. - Select (400) the best combination of architectures to perform during pre-training; - The selected architecture combination is fully trained and validated during training (9000) using a whole set of colon capsule endoscopy images to obtain an optimized architecture combination; - Predict (6000) small bowel lesions using the optimized architecture combined for classification; - The classification output (270) of the prediction (6000) is received by an output collection module having a communication device to a third party, the third party being able to perform verification by interpreting the accuracy of the classification output and being able to correct erroneous predictions, wherein the third party includes at least one of the following: another neural network, any other computing system suitable for performing the verification task, or optionally a physician expert in gastrointestinal imaging. - Store the corrected predictions in the storage component.
2. The method according to claim 1, wherein, The classification neural network component includes at least two blocks, each block having a densely connected layer followed by a dropout layer.
3. The method according to claim 1 or 2, wherein, The last block of the classification neural network component includes a batch normalization layer followed by a densely connected layer, wherein the depth dimension is equal to the number of lesion types to be classified.
4. The method according to claim 1, wherein, The optimal combination of architectures to execute during pre-training is the best among the following: VGG16, IncpetionV3, Xception, EfficientNetB5, EfficientNetB7, ResNet50, and ResNet125.
5. The method according to claim 1 or 4, wherein, The optimal execution combination is selected based on overall accuracy and the f1 metric.
6. The method according to claim 1 or 4, wherein, The training of the optimal execution combination sequentially includes two to four densely connected layers starting at 4096 and halved up to 512.
7. The method according to claim 1, 4 or 6, wherein, There is a drop layer with a drop rate of 0.1 between the last two layers of the optimal execution combination.
8. The method according to claim 1, 4 or 6, wherein, The training of the subset includes a training-to-validation ratio of 10% to 90%.
9. The method according to claim 1, 4 or 6, wherein, The verification is performed through user input.
10. The method according to claim 1 or 9, wherein, The subset includes images predicted by sequentially executing the steps of the method in the storage component.
11. A portable endoscope device comprising instructions that, when executed by a processor, cause a computer to perform the steps of the method according to any one of claims 1 to 10.