System and method for analyzing a stream of images

By automatically analyzing capsule endoscopy images using deep learning neural networks and machine learning systems, the problem of long image review time in capsule endoscopy examinations has been solved, enabling rapid identification of GIT segment transitions and improving analysis efficiency and accuracy.

CN115485719BActive Publication Date: 2026-02-06GIVEN IMAGING LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180032425.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-05-01
Filing Date
2021-04-27
Publication Date
2026-02-06
Estimated Expiration
2041-04-27

AI Technical Summary

Technical Problem

When using capsule endoscopy to examine the gastrointestinal tract, the reader needs to manually review a large number of images during the image analysis process, resulting in long reading time and low efficiency, and making it difficult to quickly identify transitions between specific GIT segments.

Method used

Deep learning neural networks and machine learning systems are used to analyze images captured by capsule endoscopy, automatically identify GIT segments and refine transition points, and improve the accuracy of transition detection through classification and smoothing operations.

Benefits of technology

It reduces the number of images readers need to review, improves image analysis efficiency, shortens report generation time, and enhances the ability to identify GIT segment transitions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115485719B_ABST
    Figure CN115485719B_ABST
Patent Text Reader

Abstract

According to aspects of the present disclosure, a system includes at least one processor and at least one memory storing instructions that, when executed by the processor(s), cause the system to: obtain images of a portion of a gastrointestinal tract (GIT) captured by a capsule endoscopy device; for each of the images, provide, by a deep learning neural network, a score classifying the image into each of successive segments of the GIT; classify each image of a subset of the images for which the score meets a confidence criterion into one of the successive segments of the GIT; refine the classification of the images in the subset by processing signals corresponding to the classification of the images in the subset over time; and estimate, among the images in the subset, a transition between two adjacent segments of the GIT based on the refined classification of the images in the subset (1010).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims the benefit and priority of U.S. Provisional Application No. 63 / 018,890, filed May 1, 2020, which is incorporated herein by reference in its entirety. Technical Field

[0003] This disclosure relates to image analysis methods and systems, and more specifically, to systems and methods for analyzing image streams based on scene changes. Background Technology

[0004] Capsule endoscopy (CE) allows for the examination of the entire gastrointestinal tract (GIT) through an endoscope. Several capsule endoscopy systems and methods are designed to examine specific parts of the GIT, such as the small intestine (SB) or colon. CE is a non-invasive procedure that does not require hospitalization, and patients can continue most of their daily activities while the capsule is inside their body.

[0005] In a typical CE procedure, the patient is referred by a physician. The patient then arrives at a medical facility (e.g., a clinic or hospital) for the procedure. Under the supervision of a healthcare professional (e.g., a nurse or physician), the patient swallows a capsule approximately the size of a multivitamin and is provided with a wearable device, such as a sensor strap and recorder placed in a bag and strap to be worn on the patient's shoulder. The wearable device typically includes a storage device. The patient can then receive instructions and / or guidance before being discharged to resume their daily activities.

[0006] The capsule captures images as it passes naturally through the GIT. The images and additional data (e.g., metadata) are then transmitted to a recorder worn by the patient. The capsule is typically disposable and is expelled naturally with defecation. Surgical data (e.g., captured images or portions thereof, along with additional metadata) is stored on the wearable device's storage.

[0007] Wearable devices are typically returned to the medical facility by the patient along with the surgical data stored on them. The surgical data is then downloaded to a computing device, usually located at the medical facility, which stores engine software. The engine then processes the received surgical data into a compiled research report (or "research report"). A typical research report includes thousands of images (approximately 6,000). The number of images to be processed is usually in the tens of thousands, averaging around 90,000.

[0008] A reader (possibly a supervising physician, a specialist physician, or a referring physician) can access the study report via a reader application. The reader then reviews the study report, assesses the surgery, and provides his input via the reader application. Since the reader needs to review thousands of images, the reading time for a study report can typically average half an hour to an hour, and the reading task can be tedious. The reader application then generates a report based on the compiled study report and the reader's input. On average, it takes an hour to generate a report. The report can include, for example, images of interest, e.g., images identified as including pathology, selected by the reader; an assessment or diagnosis of the patient's medical condition based on the surgery data (i.e., the study report) and / or the reader-provided follow-up and / or treatment recommendations. The report can then be forwarded to the referring physician. The referring physician can decide on the required follow-up or treatment based on the report. SUMMARY

[0009] The present disclosure relates to systems and methods for analyzing a stream of images of the gastrointestinal tract (GIT). More specifically, the present disclosure relates to determining points in the stream of images that correspond to transitions between particular segments of the GIT, such as the transition between the stomach and the small intestine, or the transition between the small intestine and the colon. Although examples are shown and described with respect to images captured in vivo by a capsule endoscopy device, the disclosed techniques can be applied to images captured by other devices or mechanisms, including, for example, anatomical images captured by MRI.

[0010] According to aspects of the present disclosure, a system for analyzing images includes at least one processor and at least one memory having instructions stored therein. The instructions, when executed by the at least one processor, cause the system to: obtain a plurality of images of at least a portion of a gastrointestinal tract (GIT) captured by a capsule endoscopy device; for each image of the plurality of images, provide, by a deep learning neural network, a score classifying the image into each of a plurality of consecutive segments of the GIT; classify each image of a subset of the plurality of images for which the score satisfies a confidence criterion as being in one of the consecutive segments of the GIT; refine the classification of the images in the subset by processing signals that vary over time corresponding to the classification of the images in the subset; and estimate, among the images in the subset, a transition between two adjacent segments of the consecutive segments of the GIT based on the refined classification of the images in the subset.

[0011] In various embodiments, the instructions, when executed by the at least one processor, further cause the system to provide the subset of the plurality of images as the images in the plurality of images for which the score, when normalized, is above an upper threshold or below a lower threshold but not between the upper threshold and the lower threshold.

[0012] In various embodiments, in refining the classification of the images in the subset, the instructions, when executed by the at least one processor, cause the system to apply a smoothing operation to the classification of the images in the subset to provide a refined classification of the images in the subset.

[0013] In various embodiments, in applying the smoothing operation, the instructions, when executed by the at least one processor, cause the system to, for each image in the subset: obtain image classifications within a window surrounding the image, and select a median of the image classifications within the window as the refined classification of the image.

[0014] In various embodiments, the transition between the two adjacent segments is a transition between an earlier segment of the GIT and a later segment of the GIT.

[0015] In various embodiments, the two adjacent segments are the stomach and the small intestine, and the instructions, when executed by the at least one processor, further cause the system to determine whether a gastric retention condition exists based on comparing the scores of the plurality of images to a threshold number of small intestine classifications.

[0016] In various embodiments, the two adjacent segments include a first segment of the GIT and a second segment of the GIT, and the instructions, when executed by the at least one processor, further cause the system to: for each image in the plurality of images, provide, by a second deep learning neural network, scores classifying the image as the first segment of the GIT, the second segment of the GIT, and an anatomical feature adjacent to a transition point between the first segment and the second segment; and based on the scores provided by the second deep learning neural network, refine the transition between the first and second segments of the GIT to an earlier point before an initially estimated transition.

[0017] In various embodiments, in refining the transition, the instructions, when executed by the at least one processor, cause the system to: (i) for each image in the plurality of images: calculate a difference between a score classifying the image as the first segment of the GIT and a score classifying the image as the anatomical feature adjacent to the transition point between the first and second segments of the GIT, and calculate a sum of the calculated differences from a first image in the plurality of images up to the image; and (ii) determine the refined transition as the image corresponding to a global minimum or maximum in the calculated sum before the initially estimated transition.

[0018] In various embodiments, the two adjacent segments include a first segment of the GIT and a second segment of the GIT, and the instructions, when executed by the at least one processor, further cause the system to refine the transition between the first segment and the second segment to a later point after the estimated transition based on at least one of: a burst of classifications to the first segment of the GIT after the estimated transition, or a fluctuation between classifications to the first segment of the GIT and to the second segment of the GIT after the estimated transition that exceeds a fluctuation limit.

[0019] According to aspects of the present disclosure, a system for analyzing images includes at least one processor and at least one memory having instructions stored therein. The instructions, when executed by the at least one processor, cause the system to: obtain a plurality of images of at least a portion of a gastrointestinal tract (GIT) captured by a capsule endoscopy device; estimate, among the plurality of images, a transition between a first segment of the GIT and a second segment of the GIT based on classification scores from a first deep learning neural network that classifies images into at least two classifications, wherein the at least two classifications include a segment of the GIT and a second segment of the GIT; and refine the transition between the first segment of the GIT and the second segment of the GIT to an earlier point before the estimated transition based on classification scores of the plurality of images from a second deep learning neural network that classifies images into at least three classifications, wherein the at least three classifications include the first segment of the GIT, the second segment of the GIT, and an anatomical feature adjacent to a transition point between the first segment of the GIT and the second segment of the GIT.

[0020] In various embodiments of the system, in refining the transition, the instructions, when executed by the at least one processor, cause the system to: (i) for each image of the plurality of images: calculate a difference between a score classifying the image as the first segment of the GIT and a score classifying the image as the anatomical feature adjacent to the transition point between the first segment of the GIT and the second segment of the GIT, and calculate a sum of the calculated differences from a first image of the plurality of images up to the image; and (ii) determine the refined transition as the image corresponding to a global minimum or maximum in the calculated sum before the initially estimated transition.

[0021] In various embodiments of the system, the first segment of the GIT is pre- small intestine, the second segment of the GIT is small intestine, and the anatomical feature adjacent to the transition point between the first segment of the GIT and the second segment of the GIT is a bulbous anatomical structure (i.e., the duodenal bulb). In various embodiments of the system, the anatomical feature adjacent to the transition point between the first segment of the GIT and the second segment of the GIT is the pyloric valve.

[0022] According to aspects of the present disclosure, a system for analyzing images includes at least one processor and at least one memory having instructions stored therein. The instructions, when executed by the at least one processor, cause the system to: obtain a plurality of images of at least a portion of a gastrointestinal tract (GIT) captured by a capsule endoscopy device; estimate, among the plurality of images, a first transition between a first segment and a second segment of the GIT based on classification scores from a first deep learning neural network that classifies images into at least two classifications, wherein the at least two classifications include a segment of the GIT and a second segment of the GIT; and refine the transition between the first segment and the second segment of the GIT to a later point after the estimated transition based on at least one of: a burst of classification to the first segment of the GIT after the estimated transition, or a fluctuation between classification to the first segment of the GIT and classification to the second segment of the GIT after the estimated transition that exceeds a fluctuation tolerance.

[0023] According to aspects of the present disclosure, a system for analyzing images includes at least one processor and at least one memory having instructions stored therein. The instructions, when executed by the at least one processor, cause the system to: obtain a plurality of images of at least a portion of a gastrointestinal tract (GIT) captured by a capsule endoscopy device; for each image of the plurality of images, provide, by a machine learning system, a score classifying the image into each of at least two consecutive segments of the GIT; perform noise filtering on the classification scores to provide remaining classification scores, wherein the remaining classification scores correspond to a subset of the plurality of images; classify each image of the subset based on the remaining classification scores to provide a signal corresponding to the classification; and estimate a transition from an earlier segment to a later segment of the at least two consecutive segments of the GIT based on the classification signal.

[0024] In various embodiments of the system, the machine learning system is one of: a deep learning neural network or a classical machine learning system.

[0025] In various embodiments of the system, the transition is a transition from pre-ileum to ileum, and the instructions are executed in an offline configuration after the capsule endoscopy device exits the patient. In various embodiments of the system, the instructions, when executed by the at least one processor, cause the system to refine the transition to an earlier transition point corresponding to a global minimum of a function, wherein the function is based on a cumulative sum of differences between classification scores corresponding to the pre-ileum and classification scores corresponding to an anatomical feature adjacent to a transition point between the pre-ileum and the ileum.

[0026] In various embodiments of the system, the transition is a transition from small intestine to colon, where the instructions are executed in an offline configuration after the capsule endoscopy device exits the patient. In various embodiments of the system, the instructions, when executed by the at least one processor, cause the system to refine the transition to a later transition point based on at least one of: a classification burst to the small intestine after the transition, or a fluctuation between a classification to the small intestine and a classification to the colon after the transition that exceeds a fluctuation limit.

[0027] In various embodiments of the system, the instructions, when executed by the at least one processor, cause the system to remove irrelevant images, including at least one of: images occurring before the transition or images occurring after the transition.

[0028] In various embodiments of the system, the instructions, when executed by the at least one processor, cause the system to provide localization information to a user, where the localization information includes at least one of: information indicating that an image before the transition is classified as an image of an earlier segment of the GIT, or information indicating that an image after the transition is classified as an image of a later segment of the GIT.

[0029] In various embodiments of the system, the machine learning system provides a score classifying an image into each of at least three consecutive segments of the GIT, where the transition is a transition from a first segment to a second segment of the at least three consecutive segments of the GIT. The instructions, when executed by the at least one processor, further cause the system to estimate a second transition from the second segment to a third segment of the at least three consecutive segments of the GIT based on the classification signal.

[0030] According to aspects of the present disclosure, a system for analyzing images includes a capsule endoscopy device configured to capture, over time, a plurality of images of at least a portion of a gastrointestinal tract (GIT) of a person, a receiving device configured to be affixed to the person, communicatively coupled with the capsule endoscopy device, and to receive the plurality of images, and a computing system configured to be communicatively coupled with the receiving device and to receive the plurality of images, wherein the computing system includes at least one processor and at least one memory having instructions stored therein. The instructions, when executed by the at least one processor, cause the computing system to: classify, by a machine learning system, each image of the plurality of images as one of at least two consecutive segments of the GIT to provide a classification of each image; based on the classifications of the plurality of images, estimate an occurrence of a transition from an earlier segment to a later segment of the at least two consecutive segments of the GIT, the classifications of the plurality of images changing from the earlier segment to the later segment and remaining in the later segment for at least one of a threshold duration or a threshold number of images; and based on the estimated occurrence of the transition, provide a message indicating that the receiving device can be removed from the person.

[0031] In various embodiments of the system, the earlier segment is a small intestine and the later segment is a colon, and the message indicates that, based on the estimated occurrence of the transition from the small intestine to the colon, the receiving device can be removed.

[0032] According to aspects of the present disclosure, a system for analyzing images includes at least one processor and at least one memory having instructions stored therein. The instructions, when executed by the at least one processor, cause the system to: during a capsule endoscopy procedure, obtain a plurality of images captured by a capsule endoscopy device of a portion of a gastrointestinal tract (GIT) traversed by the capsule endoscopy device; for each image of the plurality of images, and during the capsule endoscopy procedure: provide, by a machine learning system, a score classifying the image as each of at least two consecutive segments of the GIT, and perform an online classification of the image as one of the at least two consecutive segments of the GIT to provide an online classification of the image; during the capsule endoscopy procedure, provide an online estimated transition from an earlier segment to a later segment of the at least two consecutive segments of the GIT based on a causal processing of online classifications of the plurality of images changing from the earlier segment to the later segment; after the capsule endoscopy procedure ends, perform an offline classification of each image of a subset of the plurality of images to provide offline classifications; and after the capsule endoscopy procedure ends, provide an offline estimated transition from the earlier segment to the later segment of the at least two consecutive segments of the GIT based on an acausal processing of the offline classifications.

[0033] In various embodiments of the system, the machine learning system provides a score classifying each of the images into each of at least three consecutive segments of the GIT, wherein the offline estimated transition is a transition from a first segment to a second segment of the at least three consecutive segments of the GIT. The instructions, when executed by the at least one processor, further cause the system to estimate a transition from the second segment to a third segment of the at least three consecutive segments of the GIT.

[0034] In various embodiments of the system, the instructions, when executed by the at least one processor, cause the system to determine, based on the online estimated transition, that the capsule endoscopy procedure has ended.

[0035] According to aspects of the present disclosure, a non-transitory machine-readable medium stores instructions that, when executed by a processor, cause the processor to perform a method comprising: classifying, by a machine learning system, each of a plurality of images into one of at least two consecutive segments of a gastrointestinal tract (GIT) to provide a classification of each of the images, wherein the plurality of images are captured by a capsule endoscopy device within a human body over time; based on the classification of the plurality of images, estimating an occurrence of a transition from an earlier segment to a later segment of the at least two consecutive segments of the GIT, the classification of the plurality of images changing from the earlier segment to the later segment and remaining in the later segment for at least one of a threshold duration or a threshold number of images; and based on the estimated occurrence of the transition, providing a message indicating that a receiving device secured to the human body can be removed from the human body, wherein the receiving device is configured to be communicatively coupled with the capsule endoscopy device and to receive the plurality of images.

[0036] Further details and aspects of exemplary embodiments of the present disclosure are described in greater detail below, with reference to the drawings. BRIEF DESCRIPTION OF DRAWINGS

[0037] The above and other aspects and features of the present disclosure will become more apparent from the following detailed description, taken in conjunction with the accompanying drawings, in which like reference characters refer to like elements throughout.

[0038] Figure 1 is a diagram illustrating a gastrointestinal tract (GIT) in accordance with aspects of the present disclosure;

[0039] Figure 2 is a block diagram of an exemplary system for analyzing images captured by a capsule endoscopy device in vivo in accordance with aspects of the present disclosure;

[0040] Figure 3 is a block diagram of an exemplary computing device in accordance with aspects of the present disclosure;

[0041] Figure 4is a block diagram of an exemplary deep learning neural network according to aspects of the present disclosure;

[0042] Figure 5 is an exemplary classification score and predicted class based on the deep learning neural network of Figure 4

[0043] Figure 6 is a portion of the plot of Figure 5

[0044] Figure 7 is an exemplary normalized score of the classification score of Figure 6

[0045] Figure 8 is a plot of applying an exemplary confidence criterion to the normalized score of Figure 7

[0046] Figure 9 is a plot of an exemplary smoothing operation according to aspects of the present disclosure;

[0047] Figure 10 is a plot of an exemplary estimated transition between two adjacent segments of the gastrointestinal tract according to aspects of the present disclosure;

[0048] Figure 11 is a flowchart of an exemplary operation to estimate a transition between two adjacent segments of the gastrointestinal tract according to aspects of the present disclosure;

[0049] Figure 12 is a block diagram of another exemplary deep learning neural network according to aspects of the present disclosure;

[0050] Figure 13 is a plot of exemplary classifications provided by the deep learning neural network of Figure 4 and Figure 12

[0051] Figure 14 is a plot of exemplary values computed based on classification scores provided by the deep learning neural network of Figure 12

[0052] Figure 15 is a plot of exemplary classification scores exhibiting a burst of classification after an estimated transition according to aspects of the present disclosure. DETAILED DESCRIPTION

[0053] ​​​​​​The present disclosure relates to systems and methods for analyzing medical images, and more particularly, to systems and methods for analyzing image streams of the gastrointestinal tract.

[0054] The present disclosure provides systems and methods for transition detection in image streams captured during CE procedures. The transitions detected in the image streams can be transitions from images of one anatomical region of the gastrointestinal tract (GIT) to images of another anatomical region, or can be transitions from images of a segment of the GIT where pathology exists to images of another segment of the GIT where pathology does not exist, or can be transitions from images of a sick / diseased segment to images of a healthy segment, and / or combinations thereof. Thus, as used herein, a "segment" of the GIT includes, but is not limited to, an anatomical portion having a given name. Rather, the term "segment" also includes a portion of the GIT having a particular characteristic, such as sick / diseased, healthy, presence of pathology, and / or absence of pathology, etc. Once one or more transitions are detected, the image stream can be partitioned according to the segments of the GIT to which they correspond. Although examples are shown and described with respect to images captured in vivo by a capsule endoscopy device, the disclosed techniques can be applied to images captured by other devices or mechanisms, including, for example, anatomical images captured by MRI or images captured with infrared.

[0055] In various aspects, the disclosed transition detection utilizes a combination of transition detection operations. Certain transition detection operations are used to detect transitions from pre-small intestine images to small intestine images (e.g., Figure 4 , Figure 5 , Figures 7 to 10 , Figures 12 to 14 ), and certain transition detection operations are used to detect transitions from small intestine images to colon images (e.g., Figure 4 , Figure 5 , Figures 7 to 10 , Figure 15 ). In addition, some transition detection operations are suitable for "online" use, i.e., for use while the capsule is advancing through the GIT (e.g., Figure 4 , Figure 5 ), and some transition detection operations are suitable for "offline" use, i.e., for use after the capsule has exited the patient's body (e.g., Figure 4 , Figure 5 , Figures 7 to 10 , Figures 12 to 14 ).

[0056] In short, Figure 4 and Figure 5 transition detection operations can be used for both online transition detection and offline transition detection. Figures 7 to 10 and Figures 12 to 15The operations of FIGS. 1-3 are primarily considered for offline transition detection, but some or all aspects of these figures can also be applicable to online transition detection depending on the availability of computing resources. Figure 4 , Figure 5 , Figures 7 to 10 The detection operations of FIGS. 1-3 can be used to detect transitions between multiple segments of a GIT, such as detecting a transition from an anterio-intestinal image to an intestinal image, and detecting a transition from an intestinal image to a colonic image. Figures 12 to 14 The operations of FIG. 1 are primarily considered for detecting a transition from an anterio-intestinal image to an intestinal image, while Figure 15 The operations of FIG. 3 are primarily considered for detecting a transition from an intestinal image to a colonic image. However, Figures 12 to 15 The operations of FIG. 1 can also be applicable to detecting other transitions.

[0057] As another way of summarizing aspects of the present disclosure, online detection of a transition between an anterio-intestinal image and an intestinal image, and / or a transition between an intestinal image and a colonic image can apply the operations of Figure 4 and Figure 5 Online detection of a transition between an anterio-intestinal image and an intestinal image can apply some or all of the operations of Figure 4 , Figure 5 , Figures 7 to 10 and Figures 12 to 14 Online detection of a transition between an intestinal image and a colonic image can apply some or all of the operations of Figure 4 , Figure 5 , Figures 7 to 10 and Figure 15 Online detection of a transition between an intestinal image and a colonic image can apply some or all of the operations of Figures 7 to 10 and Figures 12 to 15 depending on the availability of computing resources.

[0058] According to the present disclosure, by detecting transitions in a stream of images in vivo, a portion of the image stream can be identified to provide localization information to a reader of a study report (typically a physician) and / or to remove irrelevant images. According to some aspects of the present disclosure, a user (e.g., a physician) can build their understanding of a case by reviewing a study report that includes a display of images (e.g., captured by CE imaging device 212) that are, for example, automatically selected as likely images of interest. Since a study report typically includes thousands of images, its review can be a tedious task. Reducing the number of images included in a study report can simplify the review process for a user, reduce the reading time per case, and can result in better diagnoses. For example, in a small bowel procedure, once a transition to the colon is identified, all images captured after the transition point can be removed. This can help generate a study report with a reduced number of images, and thus can reduce study report reading time. Furthermore, this can save processing time and resources.

[0059] In various embodiments, transition detection can be utilized online to indicate the end of a procedure, which would let the patient "check out" and allow the patient to disengage the device as the capsule progresses through non-relevant portions of the GIT. In various embodiments, transition detection can be used to define an anatomical region of interest (e.g., the small intestine) and / or to segment the GIT (or a portion thereof) into different anatomical regions. Various combinations of one or more such embodiments are considered to be within the scope of the present disclosure.

[0060] In the following detailed description, specific details are set forth in order to provide a thorough understanding of the present disclosure. However, persons of ordinary skill in the art will understand that the present disclosure can be practiced without these specific details. In other instances, well-known methods, procedures and components have not been described in detail so as not to obscure the present disclosure. Some features or elements described with respect to one system can be combined with features or elements described with respect to other systems. For the sake of clarity, the discussion of identical or similar features or elements can not be repeated.

[0061] Although the present disclosure is not limited in this regard, discussions utilizing terms such as "processing," "computing," "calculating," "determining," "establishing," "analyzing," "checking," or the like, can refer to the action and / or processes of a computer, computing platform, computing system, or other electronic computing device, that manipulates and / or transforms data represented as physical (e.g., electronic) quantities within the computer's registers and / or memories into other data similarly represented as physical quantities within the computer's registers and / or memories or other information non- transitory storage medium that can store instructions to perform operations and / or processes. Although the present disclosure is not limited in this regard, the terms "plurality" and "a plurality" as used herein can include, for example, "multiple" or "two or more". The term "plurality" or "a plurality" can be used throughout the specification to describe two or more components, devices, elements, units, parameters, or the like. The term set as used herein can include one or more items. Unless explicitly stated, the methods described herein are not limited to a particular order or sequence of steps. Also, some of the described methods or elements thereof can occur or be executed concurrently, at the same time, or in disordered order.

[0062] The term "location" and its derivatives, as referred to herein with respect to images, can refer to an estimated location of the capsule along the GIT at the time the image was captured, or to an estimated location of the portion of the GIT shown in the image along the GIT.

[0063] The type of CE procedure can be determined based on, among other things, the portion of the GIT of interest and to be imaged (e.g., the colon or small bowel (“SB”)), or based on a specific use (e.g., for checking the status of a GI disease such as Crohn’s disease, or for colon cancer screening).

[0064] Unless specifically indicated otherwise, the terms “surrounding” or “adjacent” as used herein in relation to an image (e.g., an image surrounding or adjacent to another image(s)) can relate to spatial and / or temporal properties. For example, an image surrounding or adjacent to another image(s) can be an image estimated to be located in the vicinity of the other image(s) along the GIT, and / or an image captured in the vicinity of the capture time of the other image, within a certain threshold (e.g., within one or two centimeters, or within one, five, or ten seconds).

[0065] The terms “GIT” and “portion of the GIT” can refer to or include the other, respectively, depending on their context. Thus, the term “portion of the GIT” can also refer to the entire GIT, and the term “GIT” can also refer to only a portion of the GIT.

[0066] The terms “image” and “frame” can refer to or include the other, respectively, and can be used interchangeably in the present disclosure to refer to a single capture of an imaging device. For convenience, the term “image” can be used more frequently in the present disclosure, but it is to be understood that references to images should also apply to frames.

[0067] Throughout the present specification, the term(s) “classification score(s)” or “score(s)” can be used to indicate a value or vector of values applicable to one or a group of classes of images / frames. In various implementations, the value or vector of values of one or more classification scores can be or can reflect a probability. In various embodiments, a model can output a classification score that can be a probability. In various embodiments, a model can output a classification score that can not be a probability.

[0068] As used herein, “machine learning system” means and includes any computing system implementing any type of machine learning. As used herein, “deep learning neural network” means and includes a neural network having several hidden layers and which does not require feature selection or feature engineering. In contrast, a “classic” machine learning system is a machine learning system that requires feature selection or feature engineering.

[0069] Reference is made to Figure 1The diagram illustrates the GIT 100. The GIT 100 is an organ system in humans and other animals. The GIT 100 typically includes a mouth 102 for ingesting food, salivary glands 104 for producing saliva, an esophagus 106 through which food passes with the aid of contractions, a stomach 108 for secreting acids to aid digestion, a liver 110, a gallbladder 112, a pancreas 114, a small intestine 116 (e.g., SB) for absorbing nutrients, and a colon 400 (e.g., large intestine) for absorbing water and storing waste as feces before defecation. Food ingested through the mouth is digested by the GIT to absorb nutrients, and the remaining waste is expelled as feces through the anus 430.

[0070] Study reports for different parts of the GIT 100 (e.g., SB), colon 400, esophagus 106, and / or stomach 108 can be presented via a suitable user interface. As used herein, the terms "study" and "studies" refer to and include reports from images produced by a CE imaging device (e.g., Figure 2 The image selected from the captured images (212) may also optionally include information other than the images. The type of surgery performed can determine which part of the GIT 100 is of interest. Examples of types of surgery performed include, but are not limited to, SB surgery, colon surgery, SB and colon surgery, surgery designed specifically to expose or examine the SB, surgery designed specifically to expose or examine the colon, surgery designed specifically to expose or examine the colon and SB, or surgery for exposing or examining the entire GIT (i.e., esophagus, stomach, SB, and colon).

[0071] Figure 2 A block diagram of a system for analyzing medical images captured in vivo via CE surgery is shown. The system typically includes a capsule system 210 configured to capture images of the GIT (Gastrointestinal Intraocular Threat) and a computing system 300 (e.g., a local system and / or a cloud system) configured to process the captured images.

[0072] The capsule system 210 may include a swallowable CE imaging device 212 (e.g., a capsule) configured to capture images of the GIT as the CE imaging device 212 passes through it. The images may be stored on the CE imaging device 212 and / or transmitted to a receiving device 214, which typically includes an antenna. In some capsule systems 210, the receiving device 214 may be located on the patient who has swallowed the CE imaging device 212 and may take the form, for example, a strap worn by the patient or a patch attached to the patient.

[0073] The capsule system 210 can be communicatively coupled with the computing system 300 and can transmit the captured images to the computing system 300. The computing system 300 can process the received images using image processing techniques, machine learning techniques, and / or signal processing techniques, among other techniques. The computing system 300 can include a local computing device local to the patient and / or the patient’s treatment facility, a cloud computing platform provided by a cloud service, or a combination of a local computing device and a cloud computing platform.

[0074] In cases where the computing system 300 includes a cloud computing platform, the images captured by the capsule system 210 can be transmitted online to the cloud computing platform. In various embodiments, the images can be transmitted via a receiving device 214 worn or carried by the patient. In various embodiments, the images can be transmitted via the patient’s smartphone, or via any other device that is connected to the internet and can be coupled with the CE imaging device 212 or the receiving device 214.

[0075] Figure 3 A high-level block diagram of an exemplary computing system 300 that can be used with the image analysis system of the present disclosure is shown. The computing system 300 can include a processor or controller 305 (which can be or include, for example, one or more central processing unit processors (CPUs), one or more graphics processing units (GPUs or GPGPUs), a chip, or any suitable computing device), an operating system 215, a memory 320, a storage device 330, an input device 335, and an output device 340. Modules or devices for collecting or receiving medical images collected by the CE imaging device 212 Figure 2 ) worn on the patient or for displaying or selecting for display these medical images (e.g., a workstation) can be or include Figure 3 The computing system 300 shown, or that can be executed by, the computing system 300. The communication component 322 of the computing system 300 can allow for communication with remote or external devices, e.g., via the internet or another network, via radio, or via a suitable network protocol such as a file transfer protocol (FTP).

[0076] The computing system 300 includes an operating system 315, which can be or include any segment of code designed and / or configured to perform tasks involving coordinating, scheduling, arbitrating, supervising, controlling, or otherwise managing the operation of the computing system 300, e.g., scheduling the execution of programs. The memory 320 can be or include, for example, a random access memory (RAM), read only memory (ROM), dynamic RAM (DRAM), synchronous DRAM (SD-RAM), double data rate (DDR) memory chips, flash memory, volatile memory, non-volatile memory, cache memory, a buffer, a short term memory unit, a long term memory unit, or other suitable memory units or storage units. The memory 320 can be or include a plurality of possibly different memory units. The memory 320 can store, for example, instructions for performing methods (e.g., executable code 325), and / or data such as user responses, interrupts, and the like.

[0077] The executable code 325 can be any executable code, e.g., an application, a program, a process, a task, or a script. The executable code 325 can be executed by the controller 305, possibly under the control of the operating system 315. For example, execution of the executable code 325 can result in displaying or selecting for display medical images as described herein, and / or can implement any of the operations described herein. In some systems, more than one computing system 300 or component of a computing system 300 can be used for various functions described herein. For various modules and functions described herein, one or more computing systems 300 or components of a computing system 300 can be used. Devices including components similar to or different from those included in the computing system 300 can be used, and can be connected to a network and used as a system. The one or more processors 305 can be configured to perform the methods of the present disclosure by, for example, executing software or code. The storage device 330 can be or include, for example, a hard drive, a floppy drive, a compact disc (CD) drive, a compact disc - recordable (CD-R) drive, a universal serial bus (USB) device, or other suitable removable and / or fixed memory units. Data such as instructions, code, medical images, image streams, and the like can be stored in the storage device 330 and can be loaded into the memory 320 from the storage device 330, where it can be processed by the controller 305. In some embodiments, the storage device 330 can be omitted. Figure 3 Some of the components shown can be omitted.

[0078] Input device 335 may include, for example, a mouse, keyboard, touchscreen, or touchpad, or any suitable input device. It will be appreciated that any suitable number of input devices may be operatively coupled to computing system 300. Output device 340 may include one or more monitors, screens, displays, speakers, and / or any other suitable output devices. It will be appreciated that any suitable number of output devices may be operatively coupled to computing system 300, as shown in box 340. Any suitable input / output (I / O) device may be operatively coupled to computing system 300; for example, a wired or wireless network interface card (NIC), modem, printer or fax machine, universal serial bus (USB) device, or external hard drive may be included in input device 335 and / or output device 340.

[0079] include Figure 3 Multiple computer systems 300, including some or all of the components shown, can be used with the described systems and methods. For example, the CE imaging device 212, receiver, cloud-based system, and / or workstation or portable computing device for displaying images may include... Figure 3 Some or all of the components of a computer system. Including, for example... Figure 3 The cloud platform (e.g., a remote server) containing components such as the computing system 300 can receive surgical data, such as images and metadata, process and generate research reports, and display the generated research reports for physician review (e.g., on a web browser running on a workstation or portable computer). The "on-premises" option allows the use of workstations or local servers within the healthcare facility to store, process, and display images and / or research reports.

[0080] Now for reference Figure 4 The diagram illustrates a block diagram of an exemplary deep learning neural network 400 for image classification. The deep learning neural network 400 can be constructed from... Figure 2 and Figure 3 The computational system 300 is used to implement and execute the deep learning neural network 400. Typically, as those skilled in the art will understand, the deep learning neural network 400 includes an input layer, multiple hidden layers, and an output layer. The input layer, hidden layers, and output layer all include neurons or nodes. Certain neurons between layers are interconnected via weights, and such neurons in the deep learning neural network 400 compute output values ​​by applying a specific function to input values ​​from neurons or nodes in the previous layer. The function applied to the input values ​​is determined by a weight vector and biases. In a deep learning neural network, learning is performed by iteratively adjusting these biases and weights. Some layers may not be weighted layers but may be functional layers, such as max-pooling layers.

[0081] In various embodiments, the deep learning neural network 400 comprises a convolutional neural network (CNN). In machine learning, a CNN is a class of artificial neural networks most commonly applied to the analysis of visual images. As will be appreciated by those skilled in the art, the convolutional aspect of a CNN involves applying matrix processing operations to local portions of an image, and the results of these operations are a set of features used to train the neural network. The deep learning neural network 400 comprising a CNN typically includes convolutional layers, activation function layers, and pooling (typically max pooling) layers that serve to reduce dimensions without losing too many features. Additional information can be included in the operations that generate these features. Providing unique information that gives rise to the features that give the neural network information can be used to ultimately provide a way of aggregating to distinguish between different data input to the neural network.

[0082] In the illustrated embodiment, the deep learning neural network 400 can utilize one or more CNNs to classify one or more images 422 taken by the CE imaging device 212 (see Figure 2 ) as part of the GIT. In the illustrated embodiment, the GIT parts classified by the deep learning neural network 400 include pre-small intestine 412 (which includes the stomach), small intestine 414, and colon 416. The deep learning neural network 400 can be executed on the computer system 300 Figure 3 ). Those skilled in the art will appreciate the deep learning neural network 400 and how to implement it.

[0083] The deep learning neural network 400 can be trained based on labels 424 for the training images 422 and / or objects in the training images. For example, the images 422 can be labeled as part of the GIT (e.g., pre-small intestine, small intestine, or colon). In various embodiments, the training can include supervised learning. The training can further include augmenting the training images 422 to include adding noise, changing colors, hiding portions of the training images, scaling the training images, rotating the training images, and / or stretching the training images. Those skilled in the art will appreciate the training of the deep learning neural network 400 and how to implement the training.

[0084] In various embodiments, the deep learning neural network 400 can be used to classify images 422 captured by the capsule endoscopy imaging device 212 (see Figure 2 ). The classification of the images 422 can include classifying each image as a segment of the GIT. For example, the image classification can include pre-small intestine 412, small intestine 414, and colon 416. The deep learning neural network 400 provides a classification score for each segment of the GIT. Figure 4 The classifications 412-416 are exemplary, and other classifications of other portions, segments, or contiguous segments of the GIT are considered to be within the scope of the present disclosure.

[0085] The above description pertains to classifying images acquired by a capsule endoscopy device into segments of a GIT (Genomic Intra-In ...

[0086] Figure 5 It shows the results from machine learning systems (such as...) Figure 4 Examples of classification scores from a Deep Learning Neural Network 400 for images of the gastrointestinal tract (GIT) acquired by a capsule endoscopy device. Each index value refers to a separate image, therefore Figure 5 The charts in the chart correspond to 50,000 individual GIT images acquired by the capsule endoscope device. Figure 5 The bottom chart illustrates exemplary scores provided by a machine learning system (e.g., a deep learning neural network 400) classifying each image into the pre-small intestine portion (e.g., stomach and in vitro frames before swallowing) 412, the small intestine portion 414, and the colon portion 416 of the GIT. The values ​​and ranges of the scores are exemplary and can vary depending on the specific implementation of the machine learning system (e.g., the deep learning neural network 400). In the illustrated embodiment, a high score (e.g., above a threshold) indicates a high degree of confidence that the label of the frame is the label indicated by that score. For example, a high colon score indicates a high degree of confidence that the frame is a colon image. In cases where the score is "middle" as defined by a boundary threshold, such a middle score would indicate moderate confidence and uncertainty. Although Figure 5 The predicted categories and classification scores shown are presented across image frames, but they can also be viewed as predicted categories and classification scores over time. Therefore, the classification scores and / or predicted categories can be considered signals that vary over time. As described in more detail below, treating the classification scores or predicted categories as time-varying signals allows signal processing techniques to be applied to such signals.

[0087] Figure 5 The top chart illustrates methods for classifying images based on classification scores, where the predicted classification corresponds to the highest classification score. Figure 4 Taking the deep learning neural network 400 as an example, in the top chart 510, predicted category "0" corresponds to the pre-small intestine 412, predicted category "1" corresponds to the small intestine 414, and predicted category "2" corresponds to the colon 416. Figure 5In the graphs 510, 520, it can be seen that the transition from the small intestine to the small intestine occurs approximately halfway between frames 0 and 5000, while the transition from the small intestine to the colon occurs approximately halfway between frames 30000 and 35000. The top graph 510 indicates the true transition point 532 between the small intestine and the colon identified and marked by a medical expert examining the GIT images, and indicates the suggested transition point 534 based on the techniques of the present disclosure, which will be described below. As Figure 5 indicated, the suggested transition point 534 is approximately 26.3 minutes different from the true transition point 532. The reason for the difference is that the suggested transition point 534 can be a more conservative suggestion, with a lower likelihood of false positives. Because the classification scores are typically more noisy and less clear than the example noise shown in Figure 5 the middle graph 520, the more conservative suggestion increases the likelihood that the images after the transition point are correctly classified.

[0088] According to aspects of the present disclosure, Figure 5 a transition detection operation is presented whereby a suggested transition point 534 is identified when the predicted class changes in the correct direction, and a threshold number of frames or threshold duration is maintained. The correct direction of the predicted class change depends on the particular application. In the example of Figure 5 a capsule endoscopy application, the correct direction of the predicted class change can be a change from "0" to "1" or from "1" to "2". The threshold number of frames or threshold duration that the predicted class is maintained can vary depending on the particular application. In the example of Figure 5 a capsule endoscopy application, the threshold number of frames or threshold duration can correspond to a duration of approximately 26.2 minutes, although another threshold number of frames or duration can be used. According to aspects of the present disclosure, in view of the computational simplicity of identifying the suggested transition, Figure 5 the operation for suggesting a transition point of the present disclosure can be referred to as "simple" transition detection. In addition, this computational simplicity allows Figure 5 the operation of the present disclosure to be suitable for both online transition detection and offline transition detection.

[0089] In various embodiments, when the "simple" transition detection is used for online applications, the detection can estimate that the transition occurred during the procedure and / or within a relatively short time after the transition occurred (e.g., at most one hour, one and a half hours, or at most two hours after the transition occurred). Based on the estimated transition occurrence, a message can be provided to the patient, for example, to indicate that the procedure has ended and that the receiving device (e.g., the 214 of the present disclosure) can be unhooked or removed from the patient. Such a message would allow the patient to fully resume their activities without having to wait for the capsule to exit the patient's body. Figure 2

[0090] Figure 6 The graphs of the present disclosure are Figure 5 ​enlarged portion of the plot, and shows a portion of the pre-small intestine score and small intestine score for images indexed 0 to 10000. Between about indices 2800 to 4200, the pre-small intestine and small intestine classification scores are closer together, and fluctuate back and forth corresponding to the highest scoring classification. Thus, the transition point in the 2800 to 4200 index range can not have high confidence. The following description provides techniques for refining and increasing the confidence of the transition point.

[0091] Figure 7 is an example of a technique for removing classification scores that do not meet a confidence criterion and retaining classification scores that meet the confidence criterion. In the illustrated embodiment, the classification scores are normalized to be between 0 and 1. In various embodiments, different normalizations can be used, and the normalization can be linear or non-linear. Figure 7 The confidence criterion in includes an upper threshold 712 and a lower threshold 714. Normalized scores between the upper threshold 712 and the lower threshold 714 will not meet the confidence criterion and are removed. Normalized scores above the upper threshold 712 or below the lower threshold 714 will meet the confidence criterion and are retained, as shown in Figure 8 . Figure 8 The result in is a subset of images for which the classification scores meet the confidence criterion. The classification scores for the subset of images can be considered a signal that varies over time. Figure 7 and Figure 8 The illustrated confidence criterion is exemplary, and other confidence criteria are considered to be within the scope of the present disclosure. For example, in addition to or instead of a confidence criterion, signal processing techniques for removing noise can be applied.

[0092] In various embodiments, Figure 8 The result in can be further processed, or can not be further processed. Figure 9 is an example of further processing that can be performed, and illustrates a smoothing operation. Figure 9 is one example of a process for refining the classification by processing the signal that varies over time corresponding to the image classification using signal processing techniques. Figure 9 The left chart 910 reflects that a subset of images for which the classification fluctuates between the two classes. In various embodiments, the smoothing process operates to modify the classification to a more general classification around these classifications.

[0093] According to aspects of the present disclosure, the smoothing operation can be applied to each image, such as image 920. In various embodiments, for each image, the smoothing operation involves taking the image classification within a window 930 around the image 920, and then refining the classification of the image 920 to be the median of the image classifications within the window 930. Figure 9The exemplary window 930 is shown centered on a particular image. The image centered on the window 930 has a classification of "1." However, the median classification within the window 930 is a classification of "0." Thus, the classification of the image 920 centered on the window 930 is refined / corrected from a classification of "1" to a classification of "0," as shown by the right chart 940 in Figure 9 It can be seen that the median classification within the window 930 is a classification of "0." Thus, the classification of the image 920 centered on the window 930 is refined / corrected from a classification of "1" to a classification of "0," as shown by the right chart 940 in Figure 9 .

[0094] The effect of the smoothing operation shown is exemplary, and the results of applying the smoothing operation will vary depending on the distribution of the classifications and the size of the window. The size of the window can vary depending on the particular application being implemented. In various embodiments, the window need not be centered on the image, and can be a window that extends entirely forward or a window that extends entirely backward, or can be positioned at another location around the image. Such variations and other variations of the window are considered to be within the scope of the present disclosure. Additionally, smoothing operations other than median calculation are also contemplated. For example, in various embodiments, the smoothing operation can involve a majority selection or an absolute majority selection, or another operation. Such variations of the smoothing operation are considered to be within the scope of the present disclosure. Other processing for refining the classifications is contemplated, including other signal processing techniques for processing the signal corresponding to the image classifications.

[0095] In various embodiments, the smoothing operation (e.g., as shown in Figure 7 ) can be used in conjunction with a confidence threshold filter (e.g., as shown in Figure 8 and Figure 9 ). In various embodiments, the smoothing operation (e.g., as shown in Figure 8 ) can be used without the confidence threshold filter. In the embodiment shown, the image classifications corresponding to Figure 10 are shown in Figure 10 . Applying a smoothing operation with a relatively narrow window to the classifications of Figure 10 does not produce a noticeable or significant difference. Using the classifications in Figure 5 , the directional transition point 1010 from class "0" to class "1" is identified without applying a threshold delay (e.g., the 26.2 minute delay of Figure 10 ). If class "0" corresponds to classifying the image as pre-small intestine and class "1" corresponds to classifying the image as small intestine, then the transition 1010 would indicate a transition of the image from a pre-small intestine image (such as a stomach) to a small intestine image. As another example, if class "0" corresponds to classifying the image as small intestine and class "1" corresponds to classifying the image as colon, then the transition 1010 would indicate a transition of the image from a small intestine image to a colon image. Although no threshold delay is applied in Figure 5 , it is contemplated that a delay (e.g., 26.2 minutes) as shown in Figure 10 may also be applied.Figure 7 operation. However, due to the refinements provided by the denoising operations of Figure 8 and Figure 9 and the smoothing operation of Figure 7 such a delay can not be needed. Such variations are considered to be within the scope of the present disclosure.

[0096] In various embodiments, and again with reference to Figure 7 the transition can be represented as a range (not shown). In the example of Figures 7 to 10 the transition can be indicated as occurring between approximately index 2800 to index 4200. A range-based transition can be beneficial in various situations. For example, using category “0” as pre-small intestine and category “1” as small intestine, a range-based transition would indicate that images before index 2800 are most likely pre-small intestine images, while images after index 4200 are most likely small intestine images. If small intestine images are needed, images before index 2800 can be ignored, while if pre-small intestine images are needed, images after index 4200 can be ignored. However, in both cases, images between index 2800 and index 4200 can be included as part of both image subsets.

[0097] Aspects and embodiments of Figure 9 described above are exemplary, and various variations are considered to be within the scope of the present disclosure. For example, Figure 10 and Figure 9 show that certain frames around index 3500 are removed for not satisfying the confidence criteria, leaving a gap in the classification around index 3500. In various embodiments, the remaining images are re-indexed to be consecutive (not shown), so that there is no gap in Figure 5 In this case, the X-axis represents index number, rather than original frame number. Such variations and other variations are considered to be within the scope of the present application.

[0098] In contrast to the “simple” transition detection disclosed in connection with Figures 7 to 10 the transition detection operation of Figure 7 may be referred to as “baseline” transition detection, which can include the confidence criteria / denoising operations of Figure 8 and Figure 9 the smoothing operation of Figure 10 and the non-delayed transition detection of Figure 11

[0099] Referring now to Figures 4 to 10 shows a flowchart of a method of detecting a transition between pre-small intestine and small intestine images according to Figures 4 to 11 ​a flowchart of operations of the method 1100. At block 1110, the operations obtain a plurality of images of at least a portion of a gastrointestinal tract (GIT) captured by a capsule endoscopy device. At block 1120, for each of the plurality of images, the operations provide, by a deep learning neural network, a score classifying the image into each of a plurality of consecutive segments of the GIT. At block 1130, the operations classify each of a subset of the plurality of images for which the score satisfies a confidence criterion into one of the consecutive segments of the GIT. At block 1140, the operations refine the classification of the images in the subset by processing signals over time corresponding to the classification of the images in the subset. And at block 1150, the operations estimate, among the images in the subset, a transition between two adjacent segments of the consecutive segments of the GIT based on the refined classification of the images in the subset.

[0100] Figure 4 Embodiments of the method 1100 are exemplary, and the disclosed systems and methods can be applied to segments of the gastrointestinal tract (GIT) other than the pre- small intestine, small intestine, and colon. Additionally, in various embodiments, another machine learning system can be applied instead of the deep learning neural network shown in Figure 12 “classic” machine learning system is a machine learning system that involves feature engineering. Such applications are considered to be within the scope of the present disclosure.

[0101] The description above describes a method for estimating a transition between two adjacent segments of the GIT. The description below describes a technique for refining the estimated transition to a more anterior point.

[0102] Figure 12 is a block diagram of a deep learning neural network 1200 for classifying images into portions of the gastrointestinal tract (GIT). In the illustrated embodiment, the portions of the GIT classified by the deep learning neural network 1200 include pre-small intestine 1212 (which includes the stomach), small intestine 1214, and anatomical features 1216 adjacent to a transition point between the pre-small intestine 1212 and the small intestine 1214. Generally, Figures 4 to 11 The deep learning neural network 1200 of the method 1100 can be used to classify images into a first segment of the GIT, a second consecutive segment of the GIT, and anatomical features adjacent to a transition point between the first segment and the second segment of the GIT. In the case of a transition from pre-small intestine to small intestine, the anatomical features adjacent to the transition point include the pyloric valve and the bulb. As explained below, such anatomical features can be used to identify the transition point between the first segment and the second segment of the GIT in a stream of images.

[0103] By training the deep learning neural network 1200 to classify anatomical features adjacent to the transition point between two consecutive segments of the GIT, the deep learning neural network 1200 as demonstrated can be applied to refine the estimated baseline transition between the first segment and the second segment estimated by the method of Figure 3 The deep learning neural network 1200 can be trained based on the labels 1224 for the training images 1222 and / or the objects in the training images. For example, the images 1222 can be labelled as pre- small intestine, anatomical feature, or small intestine. The deep learning neural network 1200 can be executed on the computer system 300 Figure 13 The skilled person will understand the deep learning neural network 1200 and how to implement it.

[0104] Figure 4 is an example of comparing the classification based on the deep learning neural network 400 of Figure 12 with the classification based on the deep learning neural network 1200 of Figure 13 Figures 4 to 11 The top chart 1310 shows the classification of images as pre-small intestine (class “0”), small intestine (class “1”), and colon (class “2”) based on the method described in conjunction with Figure 12 The middle chart 1320 shows the classification of the same images as pre-small intestine (class “0”), small intestine entry (class “1”), and small intestine (class “2”) based on the classification scores provided by the deep learning neural network 1200 of Figure 12 The estimated transition 1340 between pre-small intestine and small intestine is placed in the same location in the top chart 1310 and the middle chart 1320. The small intestine entry includes the bulbous structure (i.e. the duodenal bulb), which can be an anatomical feature adjacent to the transition point between the stomach and the small intestine. As the duodenal bulb is at the beginning of the SB, it can be a better landmark for identifying the transition than other anatomical features such as the pyloric valve or the pylorus, which are part of the stomach. Furthermore, the duodenal bulb has a unique structure, which aids its detection. Its mucosa looks different from other types of mucosa in the GIT. It does not include villi or folds, and the bulb is open. For convenience, the term “small intestine entry” can be used to describe this anatomical feature. However, it should be understood that the actual classification is an anatomical feature adjacent to the transition point between two GIT segments.

[0105] ​Bottom chart 1330 is an enlarged version of the portion of the middle chart 1320 surrounding the estimated transition 1340 between the pre- and post-small intestine. As shown in bottom chart 1330, the classification of the small intestinal entrance provides more context about where the estimated transition 1340 between the pre- and post-small intestine is located relative to the image classified as the small intestinal entrance. As shown in bottom chart 1330, the transition from the pre- to post-small intestine may precede the estimated transition 1340 in the top chart 1310 because images prior to the estimated transition 1340 point are classified as small intestinal entrances.

[0106] Based on all aspects of this disclosure, by Figure 14 The classification scores provided by the deep learning neural network 1200 can be used to refine the estimated baseline shift to a previous point before that estimated baseline shift. Figure 12 The image shown is based on the estimated transition around 1340. Figure 13 A graph showing the calculation of the pre-intestinal classification score (1212) and the intestinal entrance score (1216). The value at each index in the graph is the cumulative sum. Specifically, for index k, the cumulative sum S... k Provided by the following formula:

[0107]

[0108] in, It is the category "1" score of index i, and It is the category "0" score for index i.

[0109] Because the initial image's category "1" score will be less than the category "0" score, the initial image's S k The value will be negative. For example, when category "0" corresponds to the pre-small intestine and category "1" corresponds to the small intestine inlet, the initial image will have a larger pre-small intestine score because the initial image is a pre-small intestine image. S will be higher as long as the image is a pre-small intestine image. k The value will decrease, for example, in region 1410. However, once the image is a small intestine inlet image, the small intestine inlet score becomes larger, and S k The value will increase, for example, in region 1420. However, as... Figure 14 As shown, the classification between the pre-small intestine and the small intestinal tract may fluctuate, due to... Figure 14 The graph shows the fluctuations in regions 1410 and 1420. However, once the image is an image of the small intestine inlet, the classification of the small intestine inlet will be more frequent, and S k The value of S will increase more frequently than it will decrease. Therefore, generally speaking, S k The decreasing trend of the value can indicate the classification before the small intestine, while S k An increasing trend in the value can indicate the classification of the small intestine inlet. Therefore, the shift would be determined by S occurring before the initially estimated baseline shift of 1340.k A true or global minimum value 1430 in the values is indicated. Those skilled in the art will recognize various techniques for determining a true minimum value within a range of values.

[0110] The cumulative sum S k described above is exemplary and various variations can be contemplated. For example, in various embodiments, the cumulative sum S k may increase before the transition point and decrease after the transition point, such that Figure 14 the graph of S Figures 12 to 14 The techniques of S

[0111] Accordingly, Figure 15 operations are described that refine an estimated baseline transition to an earlier point before the estimated baseline transition. In conjunction with Figure 15 operations are described that refine an estimated baseline transition to a later point after the estimated baseline transition.

[0112] Figures 4 to 11 A graph of classification scores for a first segment of two adjacent segments of a GIT is shown, and an estimated baseline transition 1510 from the first segment to a second segment of the two adjacent segments of the GIT is indicated. For example, the first segment can be the small intestine, and the second segment can be the colon. The estimated baseline transition 1510 can be determined based on the techniques described in conjunction with Figure 15 .

[0113] According to various aspects of this disclosure, if, after estimated transition 1510, the classification score of the first segment increases to or above threshold 1520 for at least a burst duration of 1530, then a classification burst of the first segment can be used to refine estimated transition 1510 to a later point after the burst. Threshold 1520 and burst duration 1530 can vary based on the specific application and the specific part of the GIT of interest. As an example, estimated baseline transition 1510 could be a transition from the small intestine to the colon. If, after estimated transition 1510 to the colon, the classification score of the small intestine increases to or above threshold 1520 for at least a burst duration of 1530, then such a classification burst of the small intestine can be used to refine estimated transition 1510 to a later point after the burst. Such a classification burst of the small intestine can indicate, for example, that fecal matter from the colon is introduced into the final part of the small intestine, causing the final part of the small intestine to be initially misclassified as the colon. The burst can be used to potentially identify this situation and move estimated baseline transition 1510 to a later point that may be closer to the true transition from the small intestine to the colon.

[0114] It can be imagined Figure 4 Various variations of the technology. For example, in various embodiments, instead of an outbreak, if the estimated transformation classification fluctuates between two classifications in a manner exceeding the fluctuation tolerance, the estimated transformation can be refined to a later point after the estimated transformation. Other variations are considered to be within the scope of this disclosure.

[0115] Therefore, the systems and methods described above are used to analyze in vivo image streams of the gastrointestinal tract captured by a capsule endoscopy device to perform transition detection within the in vivo image stream. Some transition detection operations can be used in offline applications, while others can be used in online applications.

[0116] As a first example of an offline application, a classification machine learning system with at least two categories is used, such as two categories corresponding to the anterior segment and the posterior segment of the GIT (e.g., pre-intestinal and small intestine). This machine learning system can be a classical machine learning system or a deep learning neural network (e.g., Figure 7 The classification scores provided by the machine learning system are treated as signals that change over time. Noise filtering signal processing operations can be performed (e.g., Figure 8 and Figure 9 Alternatively or additionally, smoothing signal processing operations can be performed (e.g., Figure 4 The changes in the obtained signal from the earlier to the later segments of the GIT can be identified as estimated transitions, such as estimated transitions from the pre-intestinal to the small intestine, or from the small intestine to the colon.

[0117] As a second example of offline application, a first classification neural network with at least two classes (e.g., pre-small intestine and small intestine) can be used, and a second classification neural network that also classifies anatomical structures or regions adjacent to transitions between different anatomical segments (e.g., the pyloric flap at the end of the stomach or the bulb at the entrance of the small intestine) can be used. Examples of anatomical structures include the pyloric flap at the end of the stomach or the bulb at the entrance of the small intestine. In various embodiments, other machine learning systems can be used instead of classification neural networks, such as classical machine learning systems.

[0118] As a third example of offline application, a classification machine learning system with at least three classes is used, such as three classes corresponding to three consecutive segments of the GIT (e.g., pre-small intestine, small intestine, and colon). The first transition from pre-small intestine to small intestine can be detected in the manner described herein (e.g., Figures 7 to 10 , Figures 12 to 14 ), and the second transition from small intestine to colon can be detected in the manner described herein (e.g., Figures 7 to 10 , Figure 15 ). Figure 5

[0119] As an example of online application, Figure 2 a“simple” transition detection operation can be used to estimate the occurrence of a transition from an earlier segment of the GIT to a later segment (e.g., from stomach to small intestine, or from small intestine to colon). A message can be provided based on the detected occurrence to inform the patient that the procedure has ended and that the receiving device can be removed (e.g., 214 of Figure 5 Such online application can shorten and simplify the procedure for the patient, and allow the patient to fully resume his / her daily activities before the capsule leaves the body.

[0120] Applications can involve both online and offline aspects simultaneously. As an example, a classification neural network with at least two classes (e.g., Figure 4 : small intestine and colon) can be applied in an online manner (e.g., Figures 7 to 10 ) to determine a transition point (e.g., from small intestine to colon). In addition, signal processing techniques (e.g., Figures 12 to 15 , Figure 5 ) can be applied in an offline manner to refine the online transition point.

[0121] Another example of application with both online and offline aspects includes GIT segmentation with online transition detection (e.g., Figures 7 to 10 ) and offline refinement of the online transition point (e.g., Figures 12 to 15 , ​ ).

[0122] ​While examples are shown and described with respect to images captured in vivo by a capsule endoscopy device, the disclosed technology can be applied to images captured by other devices or mechanisms, including, for example, anatomical images captured by MRI or images captured with infrared light (instead of or in addition to visible light). As another example, the disclosed technology can be applied to online detection of transitions during various procedures, such as online detection of small bowel to small bowel transition during endoscopy or colon to small bowel transition during colonoscopy.

[0123] The embodiments disclosed herein are examples of the present disclosure and can be embodied in various forms. For example, although certain embodiments herein are described as separate embodiments, each embodiment herein can be combined with one or more of the other embodiments herein. The specific structural and functional details disclosed herein are not to be interpreted as limiting, but are to be interpreted as representative of the representative base of the present disclosure and as teaching a skilled person to employ the present disclosure in various ways in practically any suitably detailed structure. Like reference numerals can refer to similar or identical elements throughout the description of the figures.

[0124] The phrases “in one embodiment”, “in an embodiment”, “in various embodiments”, “in some embodiments”, or “in other embodiments” can each refer to one or more of the same or different embodiments according to the present disclosure. The phrase “A or B” means “(A), (B), or (A and B)”. The phrase “at least one of A, B, or C” means “(A); (B); (C); (A and B); (A and C); (B and C); or (A, B and C)”.

[0125] Any operation, method, program, algorithm, or code described herein can be transformed into or expressed as a programming language or computer program embodied on a computer or machine-readable medium. As used herein, the terms “programming language” and “computer program” each include any language used to specify instructions to a computer, and include (but are not limited to) the following languages ​​and their derivatives: Assembler, Basic, Batch files, BCPL, C, C+, C++, Delphi, Fortran, Java, JavaScript, machine code, operating system command languages, Pascal, Perl, PL1, Python, scripting languages, Visual Basic, meta-languages ​​that specify the program itself, and all first-, second-, third-, fourth-, fifth-, or next-generation computer languages. Databases and other data schemas are also included, as well as any other meta-languages. No distinction is made between interpreted, compiled, or both interpreted and compiled languages. No distinction is made between compiled and source versions of a program. Therefore, references to a program (in which a programming language may exist in multiple states, such as source, compiled, object, or linked) are references to any and all of these states. References to the procedure may encompass the actual instructions and / or the intent of those instructions.

[0126] It should be understood that the foregoing description is for illustrative purposes only. To the extent consistent with this disclosure, any or all aspects detailed herein may be used in conjunction with any or all other aspects detailed herein. Various alternatives and modifications can be devised by those skilled in the art without departing from this disclosure. Therefore, this disclosure is intended to cover all such alternatives, modifications, and variations. The embodiments described with reference to the accompanying drawings are merely illustrative of certain examples of this disclosure. Other elements, steps, methods, and techniques that are not substantially different from those described in the foregoing and / or appended claims are also intended to fall within the scope of this disclosure.

[0127] While several embodiments of this disclosure are shown in the accompanying drawings, they are not intended to limit the disclosure thereto, as the aim is to make the scope of the disclosure as broad as permitted by the art, and this specification should be read in the same manner. Therefore, the above description should not be construed as restrictive, but merely as illustrative of particular embodiments. Other modifications within the scope and spirit of the appended claims will be contemplated by those skilled in the art.

Claims

1. A system for analyzing images, the system comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the system to: obtain a plurality of images of at least a portion of a gastrointestinal tract (GIT) captured by a capsule endoscopy device; for each image of the plurality of images, provide, by a deep learning neural network, a score classifying the image into each of a plurality of consecutive segments of the GIT; classify each image of a subset of the plurality of images for which the score meets a confidence criterion as one of the consecutive segments of the GIT; refine the classification of the images of the subset by processing signals over time corresponding to the classification of the images of the subset; and estimate, among the images of the subset, a transition between two adjacent segments of the consecutive segments of the GIT based on the refined classification of the images of the subset; wherein the instructions, when executed by the at least one processor, further cause the system to provide the subset of the plurality of images as the images of the plurality of images for which the score, when normalized, is above an upper threshold or below a lower threshold but not between the upper threshold and the lower threshold; wherein, in refining the classification of the images of the subset, the instructions, when executed by the at least one processor, cause the system to apply a smoothing operation to the classification of the images of the subset to provide the refined classification of the images of the subset; wherein, in applying the smoothing operation, the instructions, when executed by the at least one processor, cause the system to, for each image of the subset: obtain the classifications of the images within a window around the image; and select a median of the classifications of the images within the window as the refined classification of the image.

2. The system of claim 1, wherein, the transition between the two adjacent segments is a transition between a more anterior segment of the GIT and a more posterior segment of the GIT.

3. The system of claim 2, wherein, the two adjacent segments are a stomach and a small intestine, wherein the instructions, when executed by the at least one processor, further cause the system to determine whether a gastric retention condition exists based on comparing the scores of the plurality of images to a threshold number of small intestine classifications.

4. The system of claim 1, wherein, the two adjacent segments include a first segment of the GIT and a second segment of the GIT, wherein the instructions, when executed by the at least one processor, further cause the system to: for each image of the plurality of images, provide, by a second deep learning neural network, a score classifying the image into the first segment of the GIT, the second segment of the GIT, and an anatomical feature adjacent to a transition point between the first segment and the second segment; and refine the transition between the first and second segments of the GIT to a more anterior point before the estimated transition based on the scores provided by the second deep learning neural network.

5. The system of claim 4, wherein, in refining the transition, the instructions, when executed by the at least one processor, cause the system to: for each image of the plurality of images: computing a difference between a score of classifying the image as the first segment of the GIT and a score of classifying the image as an anatomical feature adjacent to a transition point between the first segment and the second segment of the GIT, and computing a sum of the computed differences from a first image in the plurality of images up to the image; and determining the refined transition as an image corresponding to a global minimum or maximum in the computed sum prior to the initially estimated transition.

6. The system of claim 1, wherein, the two adjacent segments include a first segment of the GIT and a second segment of the GIT, wherein the instructions, when executed by the at least one processor, further cause the system to refine the transition between the first segment and the second segment to a later point after the estimated transition based on at least one of: a burst of classification to the first segment of the GIT after the estimated transition, or a fluctuation between classification to the first segment of the GIT and classification to the second segment of the GIT after the estimated transition exceeds a fluctuation tolerance.

7. A system for analyzing images, the system comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the system to: obtain a plurality of images of at least a portion of a gastrointestinal tract (GIT) captured by a capsule endoscopy device; estimate, among the plurality of images, a transition between a first segment of the GIT and a second segment of the GIT based on classification scores from a first deep learning neural network that classifies images into at least two classifications, the at least two classifications including a segment of the GIT and a second segment of the GIT; and refine the transition between the first segment of the GIT and the second segment of the GIT to an earlier point prior to the estimated transition based on classification scores of the plurality of images from a second deep learning neural network that classifies images into at least three classifications, the at least three classifications including the first segment of the GIT, the second segment of the GIT, and an anatomical feature adjacent to a transition point between the first segment and the second segment of the GIT, wherein in refining the transition, the instructions, when executed by the at least one processor, cause the system to: for each image in the plurality of images: compute a difference between a score of classifying the image as the first segment of the GIT and a score of classifying the image as the anatomical feature adjacent to the transition point between the first segment and the second segment of the GIT, and compute a sum of the computed differences from a first image in the plurality of images up to the image; and determine the refined transition as an image corresponding to a global minimum or maximum in the computed sum prior to the initially estimated transition.

8. The system of claim 7, wherein, the first segment of the GIT is pre-small intestine, the second segment of the GIT is small intestine, and the anatomical feature adjacent to the transition point between the first segment and the second segment of the GIT is a bulbous anatomical structure.

9. The system of claim 7, wherein, An anatomical feature adjacent to a transition point between the first segment and the second segment of the GIT is a pyloric valve.

10. A system for analyzing images, the system comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the system to: obtain a plurality of images of at least a portion of a gastrointestinal tract (GIT) captured by a capsule endoscopy device; for each image of the plurality of images, provide, by a machine learning system, a score classifying the image into each of at least two consecutive segments of the GIT; perform noise filtering on the classification scores to provide remaining classification scores corresponding to a subset of the plurality of images; classify each image of the subset based on the remaining classification scores to provide a signal corresponding to the classification; and estimate a transition from an earlier segment to a later segment of the at least two consecutive segments of the GIT based on the classification signal, wherein the transition is a transition from pre-ileum to ileum, wherein the instructions are executed in an offline configuration after the capsule endoscopy device exits a patient, wherein the instructions, when executed by the at least one processor, further cause the system to refine the transition to an earlier transition point corresponding to a global minimum of a function based on a cumulative sum of differences between classification scores corresponding to the pre-ileum and classification scores corresponding to an anatomical feature adjacent to a transition point between the pre-ileum and the ileum.

11. The system of claim 10, wherein, the machine learning system is one of: a deep learning neural network or a classical machine learning system.

12. The system of claim 10, wherein, the transition is a transition from ileum to colon, wherein the instructions are executed in an offline configuration after the capsule endoscopy device exits a patient.

13. The system of claim 12, wherein, the instructions, when executed by the at least one processor, further cause the system to refine the transition to a later transition point based on at least one of: a burst of classification to the ileum after the transition or a fluctuation between classification to the ileum and classification to the colon after the transition exceeds a fluctuation limit.

14. The system of claim 10, wherein, the instructions, when executed by the at least one processor, further cause the system to remove irrelevant images, the irrelevant images comprising at least one of: images occurring before the transition or images occurring after the transition.

15. The system of claim 10, wherein, the instructions, when executed by the at least one processor, further cause the system to provide localization information to a user, the localization information comprising at least one of: information indicating images before the transition are classified as images of an earlier segment of the GIT or information indicating images after the transition are classified as images of a later segment of the GIT.

16. The system of claim 10, wherein, the machine learning system provides a score classifying an image into each of at least three consecutive segments of the GIT, wherein the transition is a transition from a first segment to a second segment of the at least three consecutive segments of the GIT, wherein the instructions, when executed by the at least one processor, further cause the system to estimate a second transition from the second segment to a third segment of the at least three consecutive segments of the GIT based on the classification signal.

Citation Information

Patent Citations

  • Learning method, image recognition device, and computer-readable storage medium

    US20190034800A1

  • System and method for detection of transitions in an image stream of the gastrointestinal tract

    US9324145B1