Image scoring for intestinal pathology
An image scoring system using machine learning algorithms addresses the challenges of inconsistent intestinal disease diagnosis by standardizing image analysis, enhancing diagnostic reliability and treatment outcomes.
Patent Information
- Application Number
- JP2024004236
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-10-19
- Filing Date
- 2024-01-16
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2039-10-18
AI Technical Summary
Diagnosing and monitoring intestinal diseases such as Crohn's disease, celiac disease, and ulcerative colitis are challenging due to inconsistent image quality and variability in symptom expression, leading to unreliable diagnoses and inconsistent treatment outcomes.
An image scoring system using machine learning algorithms to analyze digestive tract images, assigning informativeness and severity scores, and sorting images for standardized diagnosis and monitoring.
Provides a standardized and reliable method for diagnosing and monitoring intestinal diseases by reducing variability in image interpretation and improving treatment effectiveness.
Smart Images

Figure 0007719221000002 
Figure 0007719221000003 
Figure 0007719221000004
Abstract
Description
[Background technology]
[0001] For a given intestinal disorder or disease (bowel disease), diagnosing and monitoring the given bowel disease poses various challenges depending on how and where the symptoms or characteristics of the bowel disease manifest in a given patient. Certain bowel diseases have proven difficult to diagnose reliably and consistently, even for experts. Some examples include Crohn's disease, celiac disease, irritable bowel syndrome, and ulcerative colitis. These difficulties in diagnosis and monitoring can lead to further problems not only for individual patients but also for the medical field; without accurate and reliable methods for diagnosing and monitoring specific aspects of a given bowel disease, determining the effectiveness of treatment for the given bowel disease is even more difficult, requiring extensive trial and error with subjective measurements, lacking a standard of care, and resulting in inconsistent results that can be left to chance.
[0002] Advances in endoscopic technology have enabled the retrieval of photographic images of organs typically not visible without more invasive procedures, such as surgery or laparoscopy. However, even with the use of imaging equipment to capture more images previously unavailable to researchers and practitioners, many obstacles remain to more accurate diagnosis and monitoring. For example, new methods of endoscopy can produce some images that contain artifacts or low-quality images that may not contain enough information to reliably base a diagnosis. Additionally, symptoms or characteristics of bowel disease may be expressed in only a few ways depending on the location of a given image relative to the affected organ, leading to other discrepancies in diagnosis and monitoring. When enterologists (gastroenterologists) or other medical professionals review bowel images, different reviewers often reach different conclusions despite reviewing the same images. Even the same reviewer may reach different conclusions when reviewing the same images in different circumstances or at different times. Summary of the Invention [Means for solving the problem]
[0003] Provided herein are example embodiments of systems, devices, articles of manufacture, methods, and / or computer program products (non-transitory computer readable media), and / or combinations and subcombinations thereof, for achieving image scoring for bowel pathologies as further described herein.
[0004] An example embodiment may include a computer-implemented method including: receiving, via at least one processor, a set of image frames, at least a subset of the set of image frames depicting an interior surface of a digestive tract of a given patient; automatically assigning, via the at least one processor, a first score to each image frame of the set of image frames, the first score corresponding to an informativeness of the image frame regarding a likelihood of the image frame depicting the presence or absence of a given feature on the interior surface of the digestive tract; automatically assigning, via the at least one processor, a second score to each image frame having the first score within a specified range, the second score corresponding to a severity of the given feature; and automatically sorting, via the at least one processor, the set of image frames according to at least one of the first score or the second score.
[0005] An example embodiment may include the method described above, wherein the arranging includes ranking at least some of the image frames of the set of image frames in descending order of score with respect to at least one of the first score and the second score; The method includes at least one of ranking the image frames in ascending order of score, sorting at least some of the image frames of the set of image frames in a random order of score, or positioning at least some of the image frames of the set of image frames in the array based on adjacent image frames in the array having different scores, where the number of adjacent image frames in the array having different score values meets or exceeds a predetermined threshold.
[0006] An example embodiment may include the method described above, further including automatically presenting, by at least one processor, a first image frame of the set of image frames to the reviewer, and automatically presenting, by at least one processor, a next image frame of the set of image frames to the reviewer in the order in which the image frames were automatically arranged.
[0007] An example embodiment may include the method described above, further including analyzing metadata from each image frame of the set of image frames, the metadata including at least one of a timestamp, a location, a relative location within the digestive tract, and at least one image property, and automatically selecting, by at least one processor, at least one of the image frames from the set of image frames based at least in part on the metadata.
[0008] An example embodiment may include the above method, wherein the selecting further includes automatically determining, by the at least one processor, for a given image frame, at least one of whether the first score or the second score does not meet a predetermined threshold; whether a visual artifact is determined to be present in the image frame; whether timestamp metadata, position metadata, or relative position metadata indicates that the image frame represents a sample for which the image-based indicator is likely to be unreliable; or whether the relative proximity of the given image frame to another image frame is below a predetermined proximity threshold based on differences in time, position, or relative position within the digestive tract; and automatically removing, by the at least one processor, the given image frame from the set of image frames based on the determination.
[0009] An example embodiment may include the method described above, wherein the removing includes relabeling the image frame so that it is not treated as a member of the set of image frames.
[0010] An example embodiment may include the above method, wherein a given image frame is selected by at least one of random sampling, sampling by time, sampling by position, sampling by frequency, sampling by relative position, sampling by relative position proximity, and sampling by change in value of at least one score.
[0011] An example embodiment may include the method described above, wherein the digestive tract includes at least one of the esophagus, stomach, small intestine, and large intestine.
[0012] An example embodiment may include the above method, wherein the set of image frames originates from at least one imaging device capable of image generation and includes at least one of photographic image frames, ultrasound image frames, radiographic image frames, confocal image frames, or tomographic image frames.
[0013] An example embodiment may include the method described above, wherein the at least one imaging device is an ingestible camera, an endoscope, a therioscope, a laparoscope, an ultrasound The device includes at least one device selected from the group consisting of a wave scanner, an X-ray detector, a computed tomography scanner, a positron emission tomography scanner, a magnetic resonance tomography scanner, an optical coherence tomography scanner, and a confocal microscope.
[0014] An example embodiment may include the method described above, further including automatically generating, by the at least one processor, a diagnosis of a disease related to the digestive tract of the given patient based on the at least one score.
[0015] An example embodiment may include the method described above, wherein a diagnosis is generated in response to observing a given patient's response to a stimulus or in response to a treatment administered to a given patient for a disease.
[0016] An example embodiment may include the above method performed using an imaging system (e.g., an endoscope, a video capsule endoscope, etc.) to monitor disease progression or disease improvement or to quantify the effectiveness of treatment (e.g., by tracking a severity score).
[0017] An example embodiment may include quantifying, by at least one processor, the severity of at least one small intestinal enteropathy, e.g., celiac disease, tropical sprue, drug-induced enteropathy, protein-losing enteropathy, or dysbiotic enteropathy.
[0018] An example embodiment may include the method described above, wherein the disease includes at least one of Crohn's disease, celiac disease, irritable bowel syndrome, ulcerative colitis, or a combination thereof.
[0019] An example embodiment may include the method described above, wherein the set of image frames includes a plurality of homogenous images.
[0020] An example embodiment may include the method described above, wherein the assignment of at least the first score and the second score is performed according to an algorithm corresponding to at least one of a pattern recognition algorithm, a classification algorithm, a machine learning algorithm, or at least one artificial neural network.
[0021] An example embodiment may include the above method, further including receiving, via at least one processor, information regarding a given patient's medical condition, medical history, patient demographics (e.g., age, height, weight, sex, etc.), nutritional habits, and / or clinical tests, and passing, via the at least one processor, the information as input to the algorithm.
[0022] An example embodiment may include the method described above, wherein the at least one artificial neural network includes at least one of a feedforward neural network, a recurrent neural network, a modular neural network, or a memory network.
[0023] An example embodiment may include the method described above, wherein the feedforward neural network corresponds to at least one of a convolutional neural network, a probabilistic neural network, a time-delay neural network, a perceptron neural network, or an autoencoder.
[0024] In one example embodiment, a computer-implemented method may include receiving, via at least one processor, a set of image frames, wherein at least a subset of the set of image frames depict an interior surface of a digestive tract of a given patient; The method includes automatically extracting, by at least one processor, at least one region of interest within an image frame of the set of image frames, wherein the at least one region of interest is defined by determining, by the at least one processor, that an edge value exceeds a predetermined threshold; automatically orienting, by the at least one processor, the at least one region of interest based at least in part on the edge value; automatically inputting, via the at least one processor, the at least one region of interest into an artificial neural network to determine at least one score corresponding to the image frame, wherein the at least one score represents at least one of information in the image frame to indicate a digestive tract or a given feature affecting severity of the given feature; and automatically assigning, via the at least one processor, the at least one score to the image frame based on output of the artificial neural network in response to inputting the at least one region of interest.
[0025] An example embodiment may include the above method, further including at least one of automatically removing, by the at least one processor, at least some color data from the at least one region of interest, or automatically converting, by the at least one processor, the at least one region of interest to grayscale.
[0026] An example embodiment may include the method described above, further including automatically normalizing, by at least one processor, at least one image attribute of the at least one region of interest.
[0027] An example embodiment may include the method described above, wherein the at least one image attribute includes intensity.
[0028] An example embodiment may include the above method, wherein the region of interest within the image frame does not exceed 25 percent of the total area of the image frame.
[0029] An example embodiment may include the above method, wherein each region of interest within an image frame has the same aspect ratio as any other region of interest within the image frame.
[0030] An example embodiment may include the method described above, wherein the at least one artificial neural network includes at least one of a feedforward neural network, a recurrent neural network, a modular neural network, and a memory network.
[0031] An example embodiment may include the method described above, wherein the feedforward neural network corresponds to at least one of a convolutional neural network, a probabilistic neural network, a time-delay neural network, a perceptron neural network, and an autoencoder.
[0032] In one example embodiment, the method may include the above, further including automatically identifying, by at least one processor, at least one additional region of interest in an additional image frame of the set of image frames, wherein the additional image frame and the image frame each represent a homogenous image from the set of image frames; automatically inputting, via the at least one processor, the at least one additional region of interest into an artificial neural network; and automatically assigning, via the at least one processor, at least one additional score to the additional image frame of the set of image frames based on an output of the artificial neural network in response to inputting the at least one region of interest.
[0033] An example embodiment may include the method described above, wherein the at least one score is determined by at least one decision tree.
[0034] An example embodiment may include the method described above, wherein the at least one decision tree is part of at least one of a random decision forest or a regression random forest.
[0035] An example embodiment may include the method described above, wherein the at least one decision tree includes at least one classifier.
[0036] An example embodiment may include the method described above, further including generating, via at least one processor, at least one regression based at least in part on the at least one decision tree.
[0037] An example embodiment may include the method described above, wherein the at least one regression includes at least one of a linear model, a cutoff, a weighting curve, a Gaussian function, a Gaussian field, a Bayesian prediction, or a combination thereof.
[0038] An example embodiment may include a non-transitory computer-implemented method including: receiving, via at least one processor, a first set of image frames and a corresponding first set of metadata for input to a first machine learning (ML) algorithm; inputting, via the at least one processor, the first set of image frames and the corresponding first set of metadata to the first ML algorithm; generating, via the at least one processor, a second set of image frames and a corresponding second set of metadata as output from the first ML algorithm, wherein the second set of image frames corresponds to at least one of a tagged version of the first set of image frames and a modified version of the first set of image frames, and the second metadata set includes at least one score corresponding to at least one image frame in the second set of image frames; and inputting, via the at least one processor, the second set of image frames and the second metadata set to the second ML algorithm, thereby generating a refined set of image frames or at least one of refined metadata corresponding to the refined set of image frames.
[0039] An example embodiment may include the method described above, wherein the first ML algorithm and the second ML algorithm belong to different categories of ML algorithms.
[0040] An example embodiment may include the method described above, wherein the first ML algorithm and the second ML algorithm are separate algorithms of a single category of ML algorithms.
[0041] An example embodiment may include the method described above, wherein the first ML algorithm and the second ML algorithm comprise different iterations or epochs of the same ML algorithm.
[0042] An example embodiment may include the method described above, wherein the first set of image frames includes a plurality of image frames depicting the digestive tract of a given patient and at least one additional image frame depicting a corresponding digestive tract of at least one other patient.
[0043] An example embodiment may include the method described above, wherein the at least one processor automatically selects image frames depicting the interior surface of the digestive tract of the given patient separately from any other image frames, or the at least one processor automatically selects a subset of the set of image frames depicting the interior surface of the digestive tract of the given patient together with at least one additional image frame corresponding to the digestive tract of at least one other patient. and dynamically selecting image frames depicting the digestive tract of a given patient, such that the selected image frames depicting the digestive tract of at least one other patient are interspersed with at least one image frame depicting a corresponding digestive tract of at least one other patient.
[0044] An example embodiment may include the method described above, wherein the second ML algorithm is configured to generate an output that is used as an input to at least one of the first ML algorithm and the third ML algorithm.
[0045] An example embodiment may include a computer-implemented method including: receiving, via at least one processor, output of an imaging device, the imaging device output including a plurality of image frames forming at least a subset of a set of image frames depicting an interior surface of a digestive tract of a given patient; automatically decomposing, via the at least one processor and at least one machine learning (ML) algorithm, at least one image frame of the plurality of image frames into a plurality of regions of interest, wherein at least one region of interest is defined by determining, by the at least one processor, that an edge value exceeds a predetermined threshold; automatically assigning, via the at least one processor, a first score based at least in part on the edge value of each region of interest; automatically shuffling, via the at least one processor, the set of image frames; and outputting, via the at least one processor, an organized presentation of the shuffled set of image frames.
[0046] An example embodiment may include the method described above, wherein the decomposition further includes training, via at least one processor, at least one ML algorithm based at least in part on the output.
[0047] An example embodiment may include the method described above, wherein the decomposing further includes assigning, via the at least one processor, a second score to the organized representation of the set of image frames.
[0048] An example embodiment may include a system including a memory and at least one processor, the processor configured to perform operations including receiving a set of image frames, where at least a subset of the set of image frames depict an inner surface of a digestive tract of a given patient; automatically assigning a first score to each image frame of the set of image frames, the first score corresponding to an informativeness of the image frame regarding the likelihood of the image frame depicting the presence or absence of a given feature on the inner surface of the digestive tract; automatically assigning a second score to each image frame having the first score within a specified range, the second score corresponding to a severity of the given feature; and automatically sorting the set of image frames according to at least one of the first score or the second score.
[0049] An example embodiment may include the system described above, wherein the arranging includes at least one of: ranking at least some image frames of the set of image frames in descending order of score with respect to at least one of the first score and the second score; ranking at least some image frames of the set of image frames in ascending order of score; sorting at least some image frames of the set of image frames in a random order of score; or positioning at least some image frames of the set of image frames in the array based on adjacent image frames in the array having different scores, wherein the number of adjacent image frames in the array having different score values meets or exceeds a predetermined threshold.
[0050] An example embodiment may include the system described above, wherein the at least one processor is further configured to automatically present a first image frame of the set of image frames to the reviewer, and automatically present a next image frame of the set of image frames to the reviewer in the order in which the image frames were automatically arranged.
[0051] An example embodiment may include the system described above, wherein the at least one processor is further configured to analyze metadata from each image frame of the set of image frames, the metadata including at least one of a timestamp, a location, a relative location within the digestive tract, and at least one image property, and automatically select at least one of the image frames from the set of image frames based at least in part on the metadata.
[0052] An example embodiment may include the above system, wherein the selecting further includes automatically determining at least one of, for a given image frame, whether the first score or the second score does not meet a predetermined threshold; whether a visual artifact is determined to be present in the image frame; whether timestamp metadata, position metadata, or relative position metadata indicates that the image frame represents a sample for which the image-based indicator is likely to be unreliable; or whether the relative proximity of the given image frame to another image frame is below a predetermined proximity threshold based on differences in time, position, or relative position within the digestive tract; and automatically removing the given image frame from the set of image frames based on the determination.
[0053] An example embodiment may include the system described above, wherein the removing includes relabeling the image frame so that it is not treated as a member of the set of image frames.
[0054] An example embodiment may include the system described above, wherein a given image frame is selected by at least one of random sampling, sampling by time, sampling by position, sampling by frequency, sampling by relative position, sampling by relative position proximity, and sampling by change in value of at least one score.
[0055] An example embodiment may include the system described above, wherein the digestive tract includes at least one of an esophagus, a stomach, a small intestine, and a large intestine.
[0056] An example embodiment may include the system described above, wherein the set of image frames originates from at least one imaging device capable of image generation and includes at least one of photographic image frames, ultrasound image frames, radiographic image frames, confocal image frames, or tomographic image frames.
[0057] An example embodiment may include the system described above, wherein the imaging device includes at least one device selected from the group consisting of an oral camera, an endoscope, a therioscope, a laparoscope, an ultrasound scanner, an X-ray detector, a computed tomography scanner, a positron emission tomography scanner, a magnetic resonance imaging scanner, an optical coherence tomography scanner, and a confocal microscope.
[0058] An example embodiment may include the system described above, wherein the at least one processor is further configured to automatically generate a diagnosis of a disease related to the digestive tract of the given patient based on the at least one score.
[0059] An example embodiment may include the system described above, wherein the diagnosis is based on a given patient's response to a stimulus. The antibody may be generated in response to an observed response or in response to a treatment administered to a given patient for a disease.
[0060] An example embodiment may include the system described above, wherein the disease includes at least one of Crohn's disease, celiac disease, irritable bowel syndrome, ulcerative colitis, or a combination thereof.
[0061] An example embodiment may include the system described above, wherein the set of image frames includes a plurality of homogenous images.
[0062] An example embodiment may include the system described above, wherein the assignment of at least the first score and the second score is performed according to an algorithm corresponding to at least one of a pattern recognition algorithm, a classification algorithm, a machine learning algorithm, or at least one artificial neural network.
[0063] An example embodiment may include the system described above, wherein the at least one artificial neural network includes at least one of a feedforward neural network, a recurrent neural network, a modular neural network, or a memory network.
[0064] An example embodiment may include the system described above, wherein the feedforward neural network corresponds to at least one of a convolutional neural network, a probabilistic neural network, a time-delay neural network, a perceptron neural network, or an autoencoder.
[0065] An example embodiment may include a system including a memory and at least one processor, the processor configured to perform operations including: receiving a set of image frames, at least a subset of the set of image frames depicting an interior surface of a digestive tract of a given patient; automatically extracting at least one region of interest within an image frame of the set of image frames, the at least one region of interest being defined by determining that an edge value exceeds a predetermined threshold; automatically orienting the at least one region of interest based at least in part on the edge value; automatically inputting the at least one region of interest into an artificial neural network to determine at least one score corresponding to the image frame, the at least one score representing at least one of information in the image frame for indicating a given characteristic affecting the severity of the digestive tract or a given function; and automatically assigning the at least one score to the image frame based on an output of the artificial neural network in response to inputting the at least one region of interest.
[0066] An example embodiment may include the system described above, wherein the at least one processor is further configured to at least one of automatically removing at least some color data from the at least one region of interest or automatically converting the at least one region of interest to grayscale.
[0067] An example embodiment may include the system described above, wherein the at least one processor is further configured to automatically normalize at least one image attribute of the at least one region of interest.
[0068] An example embodiment may include the system described above, wherein the at least one image attribute includes intensity.
[0069] An example embodiment may include the system described above, wherein the region of interest within the image frame does not exceed 25 percent of the total area of the image frame.
[0070] An example embodiment may include the system described above, wherein each region of interest within an image frame has the same aspect ratio as any other region of interest within the image frame.
[0071] An example embodiment may include the system described above, wherein the at least one artificial neural network includes at least one of a feedforward neural network, a recurrent neural network, a modular neural network, and a memory network.
[0072] An example embodiment may include the system described above, wherein the feedforward neural network corresponds to at least one of a convolutional neural network, a probabilistic neural network, a time-delay neural network, a perceptron neural network, and an autoencoder.
[0073] An example embodiment may include the system described above, wherein the at least one processor is further configured to automatically identify at least one additional region of interest in an additional image frame of the set of image frames, wherein the additional image frame and the image frame each represent a homogenous image from the set of image frames; automatically inputting the at least one additional region of interest into an artificial neural network; and automatically assigning at least one additional score to the additional image frame of the set of image frames based on an output of the artificial neural network in response to inputting the at least one additional region of interest.
[0074] An example embodiment may include the system described above, wherein the at least one score is determined by at least one decision tree.
[0075] An example embodiment may include the system described above, wherein the at least one decision tree is part of at least one of a random decision forest or a regression random forest.
[0076] An example embodiment may include the system described above, wherein the at least one decision tree includes at least one classifier.
[0077] An example embodiment may include the system described above, wherein the at least one processor is further configured to generate at least one regression based at least in part on the at least one decision tree.
[0078] An example embodiment may include the system described above, wherein the at least one regression includes at least one of a linear model, a cutoff, a weighting curve, a Gaussian function, a Gaussian field, or a Bayesian prediction.
[0079] In one example embodiment, a system may include a memory and at least one processor, the processor receiving a first set of image frames and corresponding first set of metadata for input to a first machine learning (ML) algorithm, inputting the first set of image frames and corresponding first set of metadata to the first ML algorithm, and generating a second set of image frames and corresponding second set of metadata as output from the first ML algorithm, the second set of image frames including tagged versions of the first set of image frames and corresponding first set of metadata. and inputting the second set of image frames and the second metadata set to a second ML algorithm, thereby generating a refined set of image frames or at least one of the refined metadata corresponding to the refined set of image frames.
[0080] An example embodiment may include the system described above, wherein the first ML algorithm and the second ML algorithm belong to different categories of ML algorithms.
[0081] An example embodiment may include the system described above, wherein the first ML algorithm and the second ML algorithm are separate algorithms of a single category of ML algorithms.
[0082] An example embodiment may include the system described above, wherein the first ML algorithm and the second ML algorithm comprise different iterations or epochs of the same ML algorithm.
[0083] An example embodiment may include the system described above, wherein the first set of image frames includes a plurality of image frames depicting the digestive tract of a given patient and at least one additional image frame depicting a corresponding digestive tract of at least one other patient.
[0084] An example embodiment may include the system described above, wherein the at least one processor is further configured to automatically select image frames depicting the interior surfaces of the given patient's digestive tract separately from any other image frames, or automatically select image frames, a subset of the image frames depicting the interior surfaces of the given patient's digestive tract, along with at least one additional image frame corresponding to the digestive tract of at least one other patient, such that the selected image frames depicting the given patient's digestive tract are interspersed with at least one image frame depicting the corresponding digestive tract of the at least one other patient.
[0085] An example embodiment may include the system described above, wherein the decomposing further includes assigning a second score to the organized representation of the set of image frames.
[0086] An example embodiment may include a system including a memory and at least one processor, the processor configured to perform operations including receiving output from an imaging device, the imaging device output including a plurality of image frames forming at least a subset of a set of image frames depicting an interior surface of a digestive tract of a given patient; automatically decomposing at least one image frame of the plurality of image frames into a plurality of regions of interest via at least one machine learning (ML) algorithm, at least one region of interest being defined by determining that an edge value exceeds a predetermined threshold; automatically assigning a first score based at least in part on the edge value of each region of interest; automatically shuffling the set of image frames; and outputting an organized presentation of the shuffled set of image frames.
[0087] An example embodiment may include the system described above, wherein the decomposing further includes training at least one ML algorithm based at least in part on the output.
[0088] An example embodiment may include the system described above, wherein the decomposing further includes assigning a second score to the organized representation of the set of image frames.
[0089] An example embodiment may include a non-transitory computer-readable storage medium having stored thereon instructions that, when executed by at least one computing device, cause operations including receiving a set of image frames, at least a subset of the set of image frames depicting the interior surface of a given patient's digestive tract; automatically assigning a first score to each image frame of the set of image frames, the first score corresponding to the image frame's informativeness regarding the image frame's likelihood of depicting the presence or absence of a given feature on the interior surface of the digestive tract; automatically assigning a second score to each image frame having the first score within a specified range, the second score corresponding to the severity of the given feature; and automatically sorting the set of image frames according to at least one of the first score or the second score.
[0090] An example embodiment may include the non-transitory computer-readable storage medium described above, wherein the arranging includes at least one of: ranking at least some image frames of the set of image frames in descending order of score with respect to at least one of the first score and the second score; ranking at least some image frames of the set of image frames in ascending order of score; sorting at least some image frames of the set of image frames in a random order of score; or positioning at least some image frames of the set of image frames in the array based on adjacent image frames in the array having different scores, wherein the number of adjacent image frames in the array having different score values meets or exceeds a predetermined threshold.
[0091] An example embodiment may include the non-transitory computer-readable medium described above, wherein the operating further includes automatically presenting a first image frame of the set of image frames to the reviewer, and automatically presenting a next image frame of the set of image frames to the reviewer in the order in which the image frames were automatically arranged.
[0092] An example embodiment may include the non-transitory computer-readable medium described above, wherein the operating further includes analyzing metadata from each image frame of the set of image frames, the metadata including at least one of a timestamp, a location, a relative location within the digestive tract, and at least one image property, and automatically selecting at least one of the image frames from the set of image frames based at least in part on the metadata.
[0093] An example embodiment may include the non-transitory computer-readable medium described above, wherein the selecting further includes automatically determining at least one of, for a given image frame, whether the first score or the second score does not meet a predetermined threshold; whether a visual artifact is determined to be present in the image frame; whether timestamp metadata, position metadata, or relative position metadata indicates that the image frame represents a sample for which the image-based indicator is likely to be unreliable; or whether the relative proximity of the given image frame to another image frame is below a predetermined proximity threshold based on differences in time, position, or relative position within the digestive tract; and automatically removing the given image frame from the set of image frames based on the determination.
[0094] An example embodiment may include the non-transitory computer-readable medium described above, wherein the deleting includes relabeling the image frame so that the image frame is not treated as a member of the set of image frames.
[0095] An example embodiment may include the non-transitory computer-readable medium described above, wherein a given image frame is selected by at least one of random sampling, sampling by time, sampling by position, sampling by frequency, sampling by relative position, sampling by relative position proximity, and sampling by change in value of at least one score.
[0096] An example embodiment may include the non-transitory computer-readable medium described above, wherein the digestive organ includes at least one of the esophagus, the stomach, the small intestine, and the large intestine.
[0097] An example embodiment may include the non-transitory computer-readable medium described above, wherein the set of image frames originates from at least one imaging device capable of image generation and includes at least one of photographic image frames, ultrasound image frames, radiographic image frames, confocal image frames, or tomographic image frames.
[0098] An example embodiment may include the non-transitory computer-readable medium described above, wherein the imaging device includes at least one device selected from the group consisting of an oral camera, an endoscope, a therioscope, a laparoscope, an ultrasound scanner, an X-ray detector, a computed tomography scanner, a positron emission tomography scanner, a magnetic resonance imaging scanner, an optical coherence tomography scanner, and a confocal microscope.
[0099] An example embodiment may include the non-transitory computer-readable medium described above, wherein the operations further include automatically generating a diagnosis of a disease related to the digestive tract of a given patient based on the at least one score.
[0100] An example embodiment may include the non-transitory computer-readable medium described above, wherein a diagnosis is generated in response to observing a given patient's response to a stimulus or in response to a treatment administered to a given patient for a disease.
[0101] An example embodiment may include the non-transitory computer-readable medium described above, wherein the disease includes at least one of Crohn's disease, celiac disease, irritable bowel syndrome, ulcerative colitis, or a combination thereof.
[0102] An example embodiment may include the non-transitory computer-readable medium described above, wherein the set of image frames includes a plurality of homogenous images.
[0103] An example embodiment may include the non-transitory computer-readable medium described above, wherein the assignment of at least the first score and the second score is performed according to an algorithm corresponding to at least one of a pattern recognition algorithm, a classification algorithm, a machine learning algorithm, or at least one artificial neural network.
[0104] An example embodiment may include the non-transitory computer-readable medium described above, wherein the at least one artificial neural network includes at least one of a feedforward neural network, a recurrent neural network, a modular neural network, or a memory network.
[0105] An example embodiment may include the non-transitory computer-readable medium described above, wherein the feedforward neural network corresponds to a convolutional neural network, a probabilistic neural network, a time-delay neural network, a perceptron neural network, or an autoencoder.
[0106] An example embodiment may include a non-transitory computer-readable storage medium having stored thereon instructions that, when executed by at least one computing device, cause operations including receiving a set of image frames, at least a subset of the set of image frames depicting an interior surface of a given patient's digestive tract; automatically extracting at least one region of interest within an image frame of the set of image frames, the at least one region of interest being defined by determining that an edge value exceeds a predetermined threshold; automatically orienting the at least one region of interest based at least in part on the edge value; automatically inputting the at least one region of interest into an artificial neural network to determine at least one score corresponding to the image frame, the at least one score representing at least one of information in the image frame to be indicative of a given characteristic affecting the severity of the digestive tract or a given function; and automatically assigning the at least one score to the image frame based on an output of the artificial neural network in response to inputting the at least one region of interest.
[0107] An example embodiment may include the non-transitory computer-readable medium described above, wherein operating further includes at least one of automatically removing at least some color data from the at least one region of interest or automatically converting the at least one region of interest to grayscale.
[0108] An example embodiment may include the non-transitory computer-readable medium described above, wherein operating further includes automatically normalizing at least one image attribute of the at least one region of interest.
[0109] An example embodiment may include the non-transitory computer-readable medium described above, wherein the at least one image attribute includes intensity.
[0110] An example embodiment may include the non-transitory computer-readable medium described above, wherein the region of interest within the image frame does not exceed 25 percent of the total area of the image frame.
[0111] An example embodiment may include the non-transitory computer-readable medium described above, wherein each region of interest within an image frame has the same aspect ratio as any other region of interest within the image frame.
[0112] An example embodiment may include the non-transitory computer-readable medium described above, wherein the at least one artificial neural network includes at least one of a feedforward neural network, a recurrent neural network, a modular neural network, and a memory network.
[0113] An example embodiment may include the non-transitory computer-readable medium described above, wherein the feedforward neural network corresponds to at least one of a convolutional neural network, a probabilistic neural network, a time-delay neural network, a perceptron neural network, and an autoencoder.
[0114] An example embodiment may include the non-transitory computer-readable medium described above, wherein the operating includes automatically identifying at least one additional region of interest in an additional image frame of the set of image frames, the additional image frame and the image frame each representing a homogenous image from the set of image frames; automatically inputting the at least one additional region of interest into an artificial neural network; and automatically assigning at least one additional score to the additional image frame of the set of image frames based on an output of the artificial neural network in response to inputting the at least one additional region of interest. and
[0115] An example embodiment may include the non-transitory computer-readable medium described above, wherein the at least one score is determined by at least one decision tree.
[0116] An example embodiment may include the non-transitory computer-readable medium described above, wherein the at least one decision tree is part of at least one of a random decision forest or a regression random forest.
[0117] An example embodiment may include the non-transitory computer-readable medium described above, wherein the at least one decision tree includes at least one classifier.
[0118] An example embodiment may include the non-transitory computer-readable medium described above, wherein operating further includes generating at least one regression based at least in part on the at least one decision tree.
[0119] An example embodiment may include the non-transitory computer-readable medium described above, wherein the at least one regression includes at least one of a linear model, a cutoff, a weighting curve, a Gaussian function, a Gaussian field, or a Bayesian prediction.
[0120] An example embodiment may include a non-transitory computer-readable storage medium having stored thereon instructions that, when executed by at least one computing device, cause operations including receiving a first set of image frames and a corresponding first set of metadata for input to a first machine learning (ML) algorithm; inputting the first set of image frames and the corresponding first set of metadata into the first ML algorithm; generating a second set of image frames and a corresponding second set of metadata as output from the first ML algorithm, wherein the second set of image frames corresponds to at least one of a tagged version of the first set of image frames and a modified version of the first set of image frames, and the second metadata set includes at least one score corresponding to at least one image frame in the second set of image frames; and inputting the second set of image frames and the second metadata set into the second ML algorithm, thereby generating a refined set of image frames or at least one of refined metadata corresponding to the refined set of image frames.
[0121] An example embodiment may include the non-transitory computer-readable medium described above, wherein the first ML algorithm and the second ML algorithm belong to different categories of ML algorithms.
[0122] An example embodiment may include the non-transitory computer-readable medium described above, wherein the first ML algorithm and the second ML algorithm are separate algorithms of a single category of ML algorithms.
[0123] An example embodiment may include the non-transitory computer-readable medium described above, wherein the first ML algorithm and the second ML algorithm comprise different iterations or epochs of the same ML algorithm.
[0124] An example embodiment may include the non-transitory computer-readable medium described above, wherein the first set of image frames includes a plurality of image frames depicting the digestive tract of a given patient and at least one additional image frame depicting a corresponding digestive tract of at least one other patient.
[0125] An example embodiment may include the non-transitory computer-readable medium described above, wherein the operating further includes one of automatically selecting image frames depicting an interior surface of the given patient's digestive tract separately from any other image frames, or automatically selecting image frames, a subset of the image frames depicting the interior surface of the given patient's digestive tract, along with at least one additional image frame corresponding to the digestive tract of at least one other patient, such that the selected image frames depicting the given patient's digestive tract are interspersed with at least one image frame depicting the corresponding digestive tract of the at least one other patient.
[0126] An example embodiment may include the non-transitory computer-readable medium described above, wherein the second ML algorithm is configured to generate an output that is used as an input to at least one of the first ML algorithm and the third ML algorithm.
[0127] An example embodiment may include a non-transitory computer-readable storage medium having stored thereon instructions that, when executed by at least one computing device, cause the at least one computing device to perform operations including: receiving output of an imaging device, the output of the imaging device including a plurality of image frames forming at least a subset of a set of image frames depicting an interior surface of a given patient's digestive tract; automatically decomposing at least one image frame of the plurality of image frames into a plurality of regions of interest via at least one machine learning (ML) algorithm, at least one region of interest being defined by determining that an edge value of each region of interest exceeds a predetermined threshold; automatically assigning a first score based at least in part on the edge value of each region of interest; automatically shuffling the set of image frames; and outputting an organized presentation of the shuffled set of image frames.
[0128] An example embodiment may include the non-transitory computer-readable medium described above, wherein the decomposing further includes training at least one ML algorithm based at least in part on the output.
[0129] An example embodiment may include the non-transitory computer-readable medium described above, wherein the decomposing further includes assigning a second score to the organized representation of the set of image frames.
[0130] One embodiment may include a method including receiving image frames output from a camera ingested by a patient; decomposing the image frames into a plurality of regions of interest; assigning values to the image frames based on the plurality of regions of interest; repeating the receiving, decomposing, and assigning for the plurality of image frames; selecting a plurality of selected image frames from the plurality of image frames, each selected image frame having a respective value that exceeds a given threshold; generating a slide deck from the plurality of selected image frames; and presenting the slide deck to a reader who evaluates the slide deck using a one-dimensional score, the slide deck being shuffled and presented in an organized manner calculated to reduce reader bias; and supervising a machine learning (ML) algorithm configured to perform the decomposition and assignment in a fully automated manner based on at least one of the values or one-dimensional scores of one or more of the plurality of image frames.
[0131] An embodiment may include a system including a memory and at least one processor, the processor being communicatively coupled to the memory and configured to receive image frames output from a camera taken by a patient, decompose the image frames into a plurality of regions of interest, and generate a plurality of regions of interest. assigning values to image frames based on the cardiac regions; repeating the receiving, decomposing, and assigning for a plurality of image frames; selecting a plurality of selected image frames from the plurality of image frames, each selected image frame having a respective value above a given threshold; generating a slide deck from the plurality of selected image frames; and presenting the slide deck to a reader who evaluates the slide deck using a one-dimensional score, the slide deck being shuffled and presented in an organized manner calculated to reduce reader bias; and monitoring a machine learning (ML) algorithm configured to perform the decomposition and assignment in a fully automated manner based on at least one of the values or one-dimensional scores of one or more image frames of the plurality of image frames.
[0132] An example embodiment may include a non-transitory computer-readable storage medium having stored thereon instructions that, when executed by at least one processor, cause the at least one processor to perform operations including receiving image frames output from a camera ingested by a patient; decomposing the image frames into a plurality of regions of interest; assigning values to the image frames based on the plurality of regions of interest; repeating the receiving, decomposing, and assigning for the plurality of image frames; selecting a plurality of selected image frames from the plurality of image frames, wherein each selected image frame has a respective value that exceeds a given threshold; generating a slide deck from the plurality of selected image frames; and presenting the slide deck to a reader who evaluates the slide deck using a one-dimensional score, wherein the slide deck is shuffled and presented in an organized manner calculated to reduce reader bias; and supervising a machine learning (ML) algorithm configured to perform the decomposition and assignment in a fully automated manner based on at least one of the image frames and the values or one-dimensional scores of one or more of the plurality of image frames.
[0133] An example embodiment may include a computer-implemented method including: receiving, via at least one processor, a set of image frames, at least a subset of the set of image frames depicting an interior surface of a digestive tract of a given patient; and automatically assigning, via the at least one processor, a first score to each image frame of the set of image frames, the first score corresponding to an informativeness of the image frame regarding a likelihood of the image frame depicting the presence or absence of a given feature on the interior surface of the digestive tract, and the given feature being used to assess a condition of a small intestinal disease.
[0134] An example embodiment may include the method described above, further including, via the at least one processor, assigning a second score to the image frame, the second score corresponding to the severity of the given feature.
[0135] An example embodiment may include the method described above, further including monitoring, via the at least one processor, disease worsening or disease improvement using the second score.
[0136] In one embodiment, the method may include the method described above, further including automatically arranging, via at least one processor, the set of image frames according to at least one of the first score or the second score to quantify the effect of the disease on the portion of the digestive tract.
[0137] An example embodiment may include a system including a memory and at least one processor, the processor communicatively coupled to the memory and via the at least one processor The system is configured to perform operations including receiving a set of image frames, at least a subset of the set of image frames depicting the inner surface of a given patient's digestive tract, and automatically assigning, via at least one processor, a first score to each image frame of the set of image frames, the first score corresponding to the image frame's informativeness regarding the likelihood of the image frame depicting the presence or absence of a given feature on the inner surface of the digestive tract, and the given feature being used to assess a condition of a small intestinal disease.
[0138] An example embodiment may include the system described above, wherein operating further includes assigning, via the at least one processor, a second score to the image frame, the second score corresponding to the severity of the given feature.
[0139] An example embodiment may include the system described above, wherein operating further includes monitoring, via the at least one processor, disease worsening or disease improvement using the second score.
[0140] An example embodiment may include the system described above, wherein operating further includes automatically arranging, via the at least one processor, the set of image frames according to at least one of the first score or the second score to quantify an effect of the disease on the portion of the digestive tract.
[0141] An example embodiment may include a non-transitory computer-readable storage medium having stored thereon instructions that, when executed by at least one processor, cause the at least one processor to perform operations including: receiving, via the at least one processor, a set of image frames, where at least a subset of the set of image frames depict an inner surface of a digestive tract of a given patient; and automatically assigning, via the at least one processor, a first score to each image frame of the set of image frames, where the first score corresponds to an informativeness of the image frame regarding a likelihood of the image frame depicting the presence or absence of a given feature on the inner surface of the digestive tract, and where the given feature is used to assess a condition of a small intestinal disease.
[0142] An example embodiment may include the non-transitory computer-readable storage medium described above, wherein operating further includes assigning, via the at least one processor, a second score to the image frame, the second score corresponding to the severity of the given feature.
[0143] An example embodiment may include the non-transitory computer-readable storage medium described above, wherein operating further includes monitoring, via the at least one processor, disease worsening or disease improvement using the second score.
[0144] An example embodiment may include the non-transitory computer-readable storage medium described above, wherein operating further includes automatically arranging, via the at least one processor, the set of image frames according to at least one of the first score or the second score to quantify an effect of disease on the portion of the digestive tract.
[0145] It is understood that the Detailed Description section below, and not the Summary or Abstract section, is intended to be used to interpret the claims. The Summary or Abstract section may describe one or more example embodiments of the enhanced battery management techniques described herein for battery life extension, although not all, and therefore is not intended to limit the scope of the appended claims in any way.
[0146] Other systems, methods, features, and advantages of the enhanced technology disclosed herein will be or become apparent to one with skill in the art upon examination of the following detailed description and the corresponding drawings filed herewith. It is intended that all such additional systems, methods, features, and advantages be included within this Detailed Description and protected by the following claims.
[0147] The accompanying drawings are incorporated into and form a part of this specification. [Brief explanation of the drawings]
[0148] [Figure 1] 1 is a flowchart illustrating a method for implementing some aspects of the enhanced image processing techniques disclosed herein, according to some embodiments. [Figure 2] 1 is a flowchart illustrating a method for capturing expected changes in disease characteristics of a given organ in response to a treatment, challenge, or another stimulus, according to some embodiments. [Figure 3] FIG. 1 is a block diagram illustrating a system implementing enhanced image processing techniques, according to some embodiments. [Figure 4] 1 is a flowchart illustrating a method for identifying and / or selecting at least one snippet in an image frame described herein, according to some embodiments. [Figure 5]1 is a flowchart illustrating a method for identifying or selecting at least one snippet in an image frame described herein, according to some embodiments. [Figure 6] 1 is a flowchart illustrating a method for training a machine learning (ML) algorithm using another ML algorithm, according to some embodiments. [Figure 7] FIG. 1 is a diagram of a single snippet within an image frame, according to some embodiments. [Figure 8] 1 is a diagram of multiple snippets in an image frame, according to some embodiments. [Figure 9] 1 is an exemplary computer system useful for implementing various embodiments. [Figure 10] 10 is an example of reassigning scores for already scored image frames with respect to a function that correlates with relative position, according to some embodiments. [Figure 11] 1 is an exemplary method for weighted selection of image frames, according to some embodiments. [Figure 12] 1 is an exemplary method for image transformation and reorientation with region of interest selection, according to some embodiments. [Figure 13] 1 is an exemplary method for image scoring and frame presentation to reduce bias in subsequent scoring, according to some embodiments.
[0149] In the drawings, identical reference numbers indicate identical or similar elements. Additionally, the leftmost digit(s) of a reference number generally identifies the figure in which the reference number first appears. DETAILED DESCRIPTION OF THE INVENTION
[0150] Provided herein are embodiments of systems, apparatus, devices, methods, and / or computer program products, and / or combinations and subcombinations thereof, for image scoring of bowel pathologies (bowel diseases). By scoring images according to the enhanced techniques disclosed herein, it is possible to reduce or overcome the difficulties described in the Background section above. These enhanced techniques for image scoring not only improve the reliability of diagnosis and monitoring by human reviewers (e.g., gastroenterologists, other medical professionals), but also improve the reliability of review, monitoring, diagnosis, or any of these. It also serves as the basis for automated techniques for the combination of
[0151] Some non-limiting examples mentioned in this Summary of the Invention include specific types of intestinal diseases, the work of medical professionals (e.g., gastroenterologists), and diagnosis or monitoring as non-limiting use cases for purposes of illustrating some of the enhanced technologies described herein. The enhanced technologies described herein may thereby be useful in establishing new standards of care in medical fields such as gastroenterology. Whereas current physician diagnosis and monitoring can be analogous to environmental scientists venturing into a forest to measure the impact of widespread plant pathology on individual trees, the enhanced technologies described herein can be considered analogous to using aerial imagery to map the entire forest, enabling comprehensive numerical metrics to assess the overall extent of pathology within the forest as a whole.
[0152] However, in addition to medical diagnosis, observation, and monitoring, these enhanced techniques of image scoring may also have broader applications than the non-limiting examples described herein, and elements of these enhanced techniques also improve aspects of other fields such as artificial intelligence, including computer vision and machine learning, to name a few.
[0153] General approach to scoring and sampling In some embodiments, a score may be calculated for the image frame on a frame-by-frame basis. Such a score may indicate, for example, how informative the image frame as a whole is with respect to features that may be evident with respect to determining other features, such as the intensity of a given disease feature in a given organ depicted in the image frame. For example, by applying an image filter, at least one computer processor (such as processor 904) may be configured to perform automatic segmentation of the image frame to determine regions of the image frame that may contribute to or impair the ability of a reviewer (human or artificial intelligence) to evaluate the image frame and score other aspects of the image frame, such as potential disease severity or features.
[0154] A one-dimensional numerical scale can be used to determine the informativeness of a given image frame. In a specific example use, when analyzing image frames depicting the inner wall of the small intestine to determine the severity of celiac disease characteristics, different information can be determined depending on whether the image frame shows the mucosal surface of the intestinal wall or whether the image frame shows, for example, the edge of the mucosa (the crest of the folds on the mucosal surface). In general, images that show mainly edges may be more informative than images that show mainly surfaces, because villi may be more visible in images of edges than surfaces. The visibility of villi can facilitate the determination of villous atrophy, which can be used, for example, to monitor the severity of celiac disease.
[0155] Image frames showing image regions of both the mucosal surface and the mucosal edge may be more informative than image frames showing only one. At the other extreme, image frames that do not show surface or edge characteristics (e.g., due to lack of light, lack of focus, obstructing artifacts in the image, etc.) may not be useful for determining significant disease characteristics, at least per computer vision, for automated scoring of severity.
[0156] The numerical scores can be analyzed and adjusted to compare the image frames to a threshold and / or make any other determinations about the images for use in subsequent scoring, such as disease severity. For example, any image frames in a set of image frames that have an informative score above a certain threshold may be passed to another scoring module to score the severity of the disease characteristic. To perform such numerical calculations and comparisons, points on a one-dimensional numerical scale can be assigned to specific characteristics and combinations thereof.
[0157] The actual values on the scale may be discrete in some embodiments, or alternatively continuous in other embodiments, with the assigned characteristic values being merely guidance. In one embodiment, the information score may be defined on a 4-point scale. For image frames that lack clear information (regarding a given disease), the information score may be zero or near zero. For image frames that primarily show images of mucosal surfaces, the information score may be 1 or near 1. For image frames that primarily show images of mucosal edges, the information score may be 2 or near 2. And for image frames that significantly show images of both mucosal surfaces and mucosal edges, the information score may be 3 or near 3. This 0-3 scale is a non-limiting example for illustrative purposes only.
[0158] Based on such informativeness scores of individual image frames, additional calculations may be performed to automatically determine the likely informativeness of another image frame. For example, the informativeness threshold may be determined manually or automatically. In some embodiments, regression rules may be used in conjunction with the informativeness scores to predict the informativeness of image frames within a frame set and explain most of the variance between image frames. In some embodiments, a random forest (RF) algorithm may be applied, as well as a machine learning (ML) model that may be trained using a manually scored training set. Other scores of the same image frames may be used to train an ML model that processes informativeness. For example, in some embodiments, image frames may first be independently scored for severity, which may then influence subsequent informativeness scoring and informativeness thresholds.
[0159] The attribution of any score to an image frame (or a sufficiently large region thereof) may be performed by any algorithm, such as an artificial intelligence (AI) model, computer vision (CV), machine learning (ML), random forest (RF) algorithm, etc. More specifically, and as described further below, artificial neural network (ANN) techniques may also be utilized in the classification and scoring process for informativeness, severity, or other measures. Examples of ANNs may include feed-forward neural networks and further include convolutional neural networks (CNNs), although other configurations are possible, as described in more detail below.
[0160] The informativeness score, in some embodiments, may be referred to as a quality score, a usefulness score, a reliability score, a reproducibility score, etc., or any combination thereof. For example, when filtering a set of image frames based on informativeness, image frames that fall below a certain informativeness threshold may be treated in various ways, such as deleting or removing them from the set, hiding or ignoring them for the purpose of assigning other scores, and / or withholding image frames below a certain information threshold for subsequent analysis, which may be similar or different to the initial score for informativeness or other scores for other values.
[0161] In addition to scoring, such as information scoring and sorting images above or below a given threshold, sampling can also be performed to select specific image frames from a set of image frames. For some portions of a given organ that are generally known to be more informative than other portions, more images can be probabilistically selected from those more informative portions than from less informative portions. For example, a particular disease of the small intestine may be proportionally more or less reliably expressed in a given portion of the small intestine (e.g., more sampling from the distal tertile, less sampling from the proximal end, etc.). When looking for features of celiac disease, to take one particular example, such features may generally be more evident in the proximal tertile of the small intestine than in the distal tertile. Thus, a sampling algorithm for observing celiac disease may select samples from the distal tertile without considering samples from the proximal tertile. More weight may be given to the proximal tertile than to the distal tertile. The specific location will depend on the organ and disease characteristics being targeted, as different diseases may be expressed differently even within the same organ and / or may be expressed in different relative locations of the same organ.
[0162] The actual information score can also be adaptively fed into sampling, so that regions with corresponding image frames having higher information scores can then be sampled more frequently than other regions with corresponding image frames having lower scores. Other sampling methods, such as square root or cube root based methods, may also be used. For informativeness, adaptive sampling and selection can be adjusted depending on whether subsequent scoring of another attribute (e.g., disease severity) is performed automatically (e.g., by AI) or manually by a human reviewer to reduce fatigue and bias.
[0163] Scores may also be aggregated, such as for a set of image frames and / or a particular patient, such as in instances where a set of image frames may have multiple patients. In some exemplary embodiments, the severity scores may be weighted by the corresponding informativeness scores to arrive at an overall score for the set of image frames and / or patient.
[0164] Additionally, image frame scores can be curve-fitted to determine informativeness and / or other score metrics, such as to determine expected values, anomalies, outliers, etc. Scores can be weighted, and a given value of relative position can be assigned to a weighted average of a particular score, such as to determine informativeness or other score metrics for a given organ portion according to its relative position, in some embodiments.
[0165] Weights may be normalized, for example, to have a mean value of 1 (or other default normal value). Weights may be determined using a probability density function (PDF), e.g., normal, uniform, logarithmic, etc. This may be modified based on known distributions of probabilities correlating with relative position along a given organ. For example, for the proximal region of the small intestine (the pyloric region transitioning into the duodenum), a particular score may be expected to differ from corresponding scores in medial or distal regions (the jejunum, ileum, and ileum transitioning into the cecum). These variations may be accounted for to determine local abnormalities or to ignore expected variations that may differ from the norm in other regions but may not otherwise be regions of interest.
[0166] According to a general approach that includes any combination of the techniques described herein, images may be segmented, scored, and / or arranged in a manner that aids in accurate identification, review, classification, and / or diagnosis. The selected examples of specific techniques described herein are not intended to be limiting, but rather to provide exemplary guidelines to those skilled in the art implementing techniques in accordance with these enhanced techniques.
[0167] Scoring the informativeness of an image frame or its region 1 is a flowchart illustrating a method 100 for identifying and / or selecting at least one snippet within an image frame in accordance with the enhanced image processing techniques described herein, according to some embodiments. Method 100 may be performed by processing logic, which may include hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions executing on a processing device), or a combination thereof. Not all steps of method 100 may be required in all cases to perform the enhanced techniques disclosed herein. Furthermore, some steps of method 100 may be performed simultaneously or in a different order than shown in FIG. 1, as will be understood by those skilled in the art.
[0168] Method 100 should be described with reference to FIGS. 7-9. However, method 100 The steps of method 100 may be performed by at least one computer processor coupled to at least one memory device. Exemplary processor and memory device(s) are described below with respect to 904 of FIG. 9. In some embodiments, method 100 may be performed by system 300 of FIG. 3, which may further include at least one processor and memory such as those of FIG. 9.
[0169] At 102, at least one processor 904 can be configured to receive a set of image frames, such as the image frames shown on the left side of Figures 7 and 8. The image frames can originate from at least one sensor, camera, or other imaging device, as described herein. Some non-limiting examples include an oral camera, an endoscope, a therioscope, a laparoscope, an ultrasound scanner, an X-ray detector, a computed tomography scanner, a positron emission tomography scanner, a magnetic resonance imaging scanner, an optical coherence tomography scanner, and a confocal microscope.
[0170] Any given image frame in the set of image frames may be generated in response to automated or event-driven input, manual input from a user, or any sequence or set of imaging procedures prescribed by a given imaging device. The generation of image frames may be based on sampling, such as, but not limited to, random sampling, sampling by time, sampling by position, sampling by frequency, sampling by relative position, sampling by relative proximity, and sampling by change in value of at least one score. For example, sampling by relative proximity may enable biologically based samples that focus on parts of an organ most likely to provide an informative representation of a particular disease signature.
[0171] An image frame, for purposes of this example, may be defined as an image that represents (shows, depicts) a portion of an organ that may exhibit, for example, features potentially indicative of a particular disease. In some cases, the features may indicate symptoms or other manifestations of a particular disease. In some embodiments, an image frame may include images from multiple cameras or sensors taken substantially simultaneously, e.g., images from bug-eye camera devices. An image frame need not be limited to a single output of one sensor or camera.
[0172] Each image frame can be received directly from the imaging device when the image is captured, from another device internal or external to the imaging device, such as a storage device capable of storing image data after the image is captured, or from any combination of the above. Accordingly, image frames can be received from virtually any other source of image data as would be understood by one skilled in the art. Some non-limiting examples of storage devices include local storage media and network storage media (e.g., on a local network or on a remote resource on a network or cloud computing infrastructure) accessible via standard protocols, a user interface (UI), an application programming interface (API), or the like.
[0173] At 104, the processor 904 may assign at least one score to the image frame based at least in part on the output of an artificial neural network (ANN) or by other means of determining the first score. In some embodiments, the first score may be an informative score (quality score, reliability score, etc.). Assigning such image frame scores to image frames in the set of image frames may allow the image frames of the set of image frames to be sorted, ranked (ascending or descending), or otherwise arranged, as further described with respect to 108 below. In some embodiments, the first score may be used for normalization, as described further below with respect to 106. Alternatively, in other embodiments, the first score for each frame assigned at 104 may instead be used as the basis for filtering using a threshold, as described further below with respect to 108, which bypasses normalization.
[0174] At 106, the processor 904 may normalize the first score of each image frame in the set of image frames by at least one curve fit, as described above, to have the first score within a specified range. The at least one curve fit (or fit to at least one curve) may be based on inputs such as at least one PDF and / or other local weightings, e.g., an expected value correlated with relative position or other parameters. In some embodiments, the normalized scores may be weighted with respect to a predetermined default value. In some embodiments, the scores may be iteratively normalized for adaptive sampling, e.g., reassigning scores and repeating 104 to select new corresponding image samples if the reassigned scores change. FIG. 10 illustrates an example of adaptive sampling for a PDF correlated with relative position.
[0175] At 108, processor 904 may be configured to extract at least one region of interest from a given image frame based on the normalized first score(s) of 106. For purposes of method 100, a snippet may include at least one region of interest of an image frame (e.g., as shown in FIGS. 7 and 8). Specifically, a snippet may be defined in some embodiments as a region of interest that may have a particular area, shape, and / or size, for example, encompassing a given edge value (e.g., exceeding a predetermined edge detection threshold) within the image frame that is identified as including a mucosal surface and / or an edge of a mucosal surface. Such snippets may be identified using at least several filters (e.g., edge filters, horizon filters, color filters, etc.). These filters may include several edge detection algorithms and / or predetermined thresholds, for example, as assigned at 104 and using the normalized score(s) described above at 106, or any other score metric. For desired results, a given neural network may be designed, customized, and / or modified in various ways in some embodiments, as described further below.
[0176] In some embodiments, a snippet may be defined as having a small area relative to the overall image frame size, e.g., 0.7% of the total image frame area per snippet. Other embodiments may define a snippet as up to 5% or 25% of the image frame area, depending on other factors such as the actual area of the target organ depicted in each image frame. Where an image frame may be defined as relatively small (e.g., a smaller component of a larger image captured at a given time), in some embodiments, a snippet may be defined as potentially occupying a relatively large amount of image frame area.
[0177] In embodiments in which the first score relates to the information (or reliability) of the image frame, 108 may optionally be selectively applied only to image frames having at least a minimum first score of informativeness or reliability, ignoring image frames having scores below a given threshold, since such image frames may not be sufficiently informative or reliable to yield a meaningful second score. For example, to selectively apply 108 to selected frame(s) forming a subset of the set frames received at 102, scored at 104, and normalized at 106, some embodiments may also perform certain steps described further below with respect to method 1100 of FIG. 11. Additionally or alternatively, a range may be specified as having upper and / or lower bounds, rather than just a threshold. Thus, the first score between any values x and y may be selected based on the first score. Image frames having a first score value below x or above y may be excluded from assignment of a second score, while image frames having a first score value below x or above y may be processed for snippet extraction, for example, at 108. Alternatively, the specified range may be discontinuous. Thus, in other embodiments, image frames having a first score value below x or above y may be excluded from snippet extraction, while image frames between x and y may be processed for snippet extraction, for example, at 108.
[0178] Additionally, at 108, any extracted snippet(s) may be oriented (or reoriented) in a particular direction. This orientation may be performed, for example, in the same operation as the extraction, or in other embodiments, as a separate step. In some embodiments, the particular direction may be relative to an edge within the region of interest. In some cases, within the intestine, for example, this orientation may be determined by ensuring that a mucosal edge is at the bottom of the oriented region of interest, or in any other preferred orientation with respect to a particular edge, face, or other image feature.
[0179] One example embodiment may operate by measuring the magnitude of a vector at a point within the region of interest relative to the centroid of a horizontal line within a circular region around each measured point. If the vector is determined to exceed a certain threshold (e.g., 25% of the radius of the circular region), this determines the bottom of the region of interest, and a rectangular patch may be drawn (e.g., parallel to the vector, with a length of x percent of the width of the image frame and a height of y percent of the height of the image frame). Each of these (re)oriented snippets may be evaluated separately from the image frame, thereby providing a score for the entire image frame, e.g., with respect to informativeness and / or severity of disease features, as further described in 110 below. In some embodiments, if there is no strong vector (e.g., 25% of the radius of the circular region), orientation may be skipped.
[0180] At 110, the at least one region of interest may be input to an artificial neural network (ANN) to determine at least one score for the image frame. The ANN may be at least one of a feedforward neural network, a recurrent neural network, a modular neural network, or a memory network, to name a few non-limiting examples. For example, in the case of a feedforward neural network, in some embodiments, it may further correspond to at least one of a convolutional neural network, a probabilistic neural network, a time-delay neural network, a perceptron neural network, an autoencoder, or any combination thereof.
[0181] To give one non-limiting example, an ANN can be a pseudo-recurrent convolutional neural network (CNN), which can have deep convolutional layers and include multiple initial features, for example, several more features per convolutional layer.
[0182] Furthermore, for a given neural network and set of image frames, regions of interest, and / or snippets, in some embodiments, the number of connected neurons in each layer of the (deep) neural network may be adjusted, e.g., to accommodate smaller image sizes. In deep neural networks, the outputs of convolutional layers may feed transition layers for further processing. The outputs of these layers may be merged into a densely connected multilayer perceptron (see FIG. 1 above) using multiple replications. In one embodiment, multiple layers may be merged with three layers using five replications (e.g., to survive dropout) to obtain a final output of 10 classes, each class representing a tenth of a class, e.g., ranging from 0 to 3. According to another example of a pseudo-regressive CNN, during training, normal noise with a standard deviation of 0.2 may be added (to prevent overfitting). (because of this) the location may be unclear.
[0183] By way of further example, the score may be refined and generalized to the image frame as a whole based on one or more snippets or regions of interest within the image frame. One way to achieve this result is to use a pseudo-regression CNN, in which a relatively small percentage of image frames in the training set are withheld from the CNN. For example, if 10% of the image frames are withheld and the snippets each produce a value of 10 after CNN processing, half of the withheld (reserved) snippets can be used to train a 100-tree random forest (RF) to estimate, in one embodiment, the local curve value. This estimate, along with the 10-value output from the other half of the reserved training snippets, can be used to build another random forest that can, for example, estimate the error.
[0184] According to some embodiments, at 112, the processor 904 may assign at least one score to the image frame based at least in part on the output of the ANN responsive to the input of 110. According to at least the examples described with respect to 110 above, a score for the entire image frame may be determined based on one or more snippets or regions of interest within the image frame. By assigning such image frame scores to image frames within the set of image frames, the image frames of the set of image frames may be sorted, ranked (ascending or descending), or otherwise arranged.
[0185] At 114, based on at least the first score, the second score, or both scores assigned to each image frame across the set of image frames (or for a particular patient), processor 904, according to some embodiments, can determine at least one aggregate score for the set (or patient). The aggregate score can be further used, for example, to diagnose or monitor a given disease of interest, or can be used to further summarize or categorize the results of image scoring and / or sampling, at least according to the techniques described above and below. For example, in exemplary embodiments, the aggregate score can indicate overall villous atrophy in a potential celiac disease patient to facilitate diagnosis. In some embodiments, aggregate scores can be generated for specific organ regions (e.g., the inner third of the small intestine, the distal fourth of the ileum, etc.).
[0186] Furthermore, the AI, in some embodiments, can perform more consistent and reliable automated diagnosis of a subject's disease based on at least one aggregate score. In the example of intestinal imaging, application of the enhanced technology described herein can enable significantly improved reliability in diagnosing celiac disease, to name one non-limiting example. Furthermore, diagnosis can be performed manually or automatically in response to observation of a given patient's reaction to irritants or in response to treatment administered to a given patient for a disease.
[0187] Such diagnosis can be facilitated by the enhanced techniques described herein even when a set of image frames contains largely homogeneous images, with the exception of certain landmark features of the intestinal tract (e.g., the pylorus, duodenum, cecum, etc.). Much of the small intestine may appear the same as other parts of the small intestine in any given section. This homogeneity can confuse or bias human reviewers, leading them to assume no changes even when subtle yet measurable differences exist. Such homogeneity can also cause differences to be lost in many conventional edge detection algorithms, depending on thresholds, smoothing, gradients, etc. However, by using the enhanced techniques described herein, detection of even subtle features of certain diseases can be significantly improved, resulting in more accurate diagnoses of those diseases.
[0188] Scoring the severity of disease features within an image frame FIG. 2 is a flowchart illustrating a method 200 for capturing predicted changes in disease characteristics of a given organ in response to a treatment, challenge, or another stimulus, according to some embodiments. Methods such as method 200 can be used in clinical trials of specific medical treatments, to name one non-limiting use example. Method 200 can be performed by processing logic, which can include hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions executing on a processing device), or a combination thereof. Not all steps of method 200 may be required in all cases to perform the enhanced image processing techniques disclosed herein. Furthermore, some steps of method 200 can be performed simultaneously or in a different order than that shown in FIG. 2, as will be understood by those skilled in the art.
[0189] Method 200 shall be described with reference to FIGS. 4-8. However, method 200 is not limited to only those exemplary embodiments. The steps of method 200 may be performed by at least one computer processor coupled to at least one memory device. Exemplary processor and memory device(s) are described below with respect to FIG. 9. In some embodiments, method 200 may be performed by system 300 of FIG. 3, which may further include at least one processor and memory such as those of FIG. 9.
[0190] The method may be initialized, executed, or otherwise started if not already running. According to some embodiments, from the initialization stage, after any implementation-specific initialization procedures, the method may proceed to 210. In some embodiments, this execution flow may be defined by finite state machines, procedural programming, object-oriented programming, and / or other logic in hardware, software, or some combination of the above.
[0191] At 210, regions of interest (referred to as "snippets" in some embodiments) may be extracted from one or more image frames (see method 500 of FIG. 5 for further discussion of regions of interest or snippets). Also at 210, these regions of interest or snippets may be processed in other ways, such as orienting and / or reorienting the regions of interest or snippets, and applying other filters for color, brightness, and normalization, to name a few non-limiting examples. From this stage at 210, the extracted regions of interest (for purposes of method 200, regions of interest are synonymously referred to as snippets) may be input into a pseudo-recurrent convolutional neural network (CNN) for processing at 220.
[0192] At 220, the regions or snippets of interest may be processed by at least one neural network, such as a pseudo-regression CNN in this example. While the non-limiting exemplary embodiment shown in FIG. 2 depicts a CNN, other types of neural networks and / or algorithms may perform the processing of 220 in other embodiments. A more detailed description of how this pseudo-regression CNN processing works is provided below with respect to FIG. 5. Output from the pseudo-regression CNN at 220 may be relayed to a CNN prediction random forest (RF) at 230A and / or a CNN error RF at 230B.
[0193] At 230A, the CNN prediction RF can receive at least one result from the pseudo-regression CNN at 220. After its own processing, the CNN prediction RF can output at least one result, for example, as input to snippet selection at 240 and / or as input to frame statistics at 250. The resulting value of the CNN prediction RF output may, in some embodiments, correspond to, for example, the local severity of a disease feature in a given snippet. Such predictions may be referred to as estimates or estimates. More details regarding predictions using RF (e.g., for edge detection) are discussed below with respect to FIG. 5. The CNN predicted RF output can also serve as input to the CNN error RF at 230B, described below.
[0194] At 230B, the CNN error RF can receive at least one result from the pseudo-regression CNN at 220. After its own processing, the CNN prediction RF can output at least one result, for example, as input to snippet selection at 240 and / or as input to frame statistics at 250. More details regarding prediction using an RF (e.g., for error estimation) are described below with respect to FIG. 5.
[0195] At 240, snippet selection may be performed. In some embodiments, snippet selection may be performed by edge detection and reorienting the extracted image around the detected edges (e.g., as a result of CNN RF prediction at 230A). Snippets may be evaluated (scored) and potentially ranked by image characteristics such as intensity and / or based on content such as severity of potential disease features. Details of snippet selection and scoring are described below with respect to FIGS. 4 and 5. Selected snippets may be used as input for computing frame statistics, for example, at 250. Additionally or alternatively, some embodiments may integrate snippet selection with steps of image transformation, masking, edge detection, reorientation, or combinations thereof, to name a few non-limiting examples. Embodiments incorporating such steps are further described below with respect to method 1200 of FIG. 12.
[0196] In some embodiments, frame statistics may be calculated at 250. Frame statistics may include, but are not limited to, the mean or standard deviation of a particular metric for a snippet, frame, or frame set, the number of snippets per frame or frame set, frame metadata per frame or frame set, error values estimated by the CNN error RF at 230B, etc. Frame statistics may be useful, for example, to determine an overall (“big picture”) score for a given patient at a given time, such as diagnosing a particular disease. In another example, frames with a higher-than-expected region-of-interest standard deviation may, in some cases, be determined to be less informative than other frames.
[0197] At 260A, the CNN prediction RF can receive at least one result from the pseudo-regression CNN at 250. After its own processing, the CNN prediction RF can output at least one result, for example, as an input to curve fit 270. The resulting value of the CNN prediction RF output, in some embodiments, can correspond, for example, to the local severity of a disease feature in a given snippet. Such a prediction can be referred to as an estimate or assessment. More details regarding prediction using RF (e.g., for edge detection) are described below with respect to FIG. 5. The CNN prediction RF output can also serve as an input to the CNN error RF at 260B, described below.
[0198] At 260B, the CNN error RF can receive at least one result from the pseudo-regression CNN at 250. After its own processing, the CNN prediction RF can output at least one result, for example, as input to a curve fit 270. More details regarding prediction using an RF (e.g., for error estimation) are described below with respect to FIG. 5.
[0199] At 270, curve fitting may be performed to further filter and / or normalize the results collected from the CNN / RF prediction and / or error estimation calculation. Normalization may be performed with respect to a curve (e.g., Gaussian, Lorentzian, etc.). In some embodiments, the image intensity for a snippet may also be referred to as the disease severity prediction for each snippet. In some embodiments, image intensity may refer to the disease severity score for the entire image frame. In some embodiments, image intensity may be measured, for example, by brightness. It may also refer to a particular image characteristic such as brightness or saturation. Further details of normalization, including curve fitting, are discussed below with respect to FIG.
[0200] 2 includes captions briefly describing the operation of method 200 using example values for a pseudo-regression CNN configured to generate 10 outputs and the RF used in a validation snippet with an error of less than 0.5. However, other embodiments may be implemented using, for example, different values for the number of outputs and the error tolerance. Appendix A attached hereto provides additional details regarding embodiments of method 200.
[0201] Various embodiments may be implemented using one or more known computer systems, such as, for example, computer system 900 shown in Figure 9. Computer system 900 may comprise system 300 of Figure 3 and may be used, for example, to implement method 200 of Figure 2.
[0202] Method 200 is disclosed in the order shown above in this exemplary embodiment of Figure 2. However, in practice, the operations disclosed above may be performed sequentially in any sequence, along with other operations, or multiple operations performed simultaneously may be performed simultaneously, or any combination of the above.
[0203] Implementing the above operations of Figure 2 may provide at least the following advantages: Using these enhanced techniques, such as the image processing outlined in Figure 2, the reliability of scoring and other assessments of disease characteristics, symptoms, and diagnoses may be improved, further enabling reliable clinical testing and / or treatment of diseases that have typically been difficult to diagnose and treat. Method 200 of Figure 2, as well as methods 400, 500, and 600 of Figures 4-6 described below, provide an efficient and reliable method for achieving specific results useful in patient monitoring, diagnosis, and treatment.
[0204] Example of a CNN application for scoring informativeness or severity FIG. 3 is a block diagram illustrating a system that implements some of the enhanced image processing techniques described herein, according to some embodiments.
[0205] System 300, in one embodiment, may include, for example, a multilayer perceptron 320 configured to receive at least one convolutional input 310 and any number of other inputs 340 to generate at least a voxel output 330. Multilayer perceptron 320 may be or include, for example, a feedforward neural network and / or classifier, and may have multiple layers. In some embodiments, these layers may be densely connected, such as when activation functions may be reused from particular layers, or in some embodiments, when activations of one or more layers are skipped for certain computations, such as determining a weight matrix.
[0206] The multilayer perceptron 320 or another neural network (e.g., at least one convolutional neural network (CNN) associated with the at least one convolutional input 310), in some embodiments, can map a given voxel output to a measure of severity of the disease feature corresponding to the voxel. Such a measure of severity can be used, for example, to attribute a score to the image frame corresponding to the at least one given voxel. In some embodiments, a voxel can additionally or alternatively be represented, for example, as a region of interest and / or a snippet.
[0207] Some of the other inputs may be, for example, color values (e.g., in RGB or PBG color spaces, to name a few non-limiting examples), and / or timestamp metadata, for example. , position metadata, or relative position metadata. In one embodiment, where relative position may indicate, for example, a scale or percentage along the length of the small intestine (start to end), such metadata may play any role in how informativeness and / or disease severity may be measured or scored by a neural network such as multilayer perceptron 320, and may affect voxel output 330.
[0208] Each of these elements may be modular components implemented in hardware, software, or some combination thereof. Additionally, in some embodiments, any of these elements may be integrated with computer system 900 or multiple instances thereof. Additionally, system 300 may be useful for implementing, for example, methods 200, 400, 500, 600, or any of methods similar to those shown in FIGS. 2 and 4-6.
[0209] The convolution input 310 data may originate from virtually any source, such as user input, machine input, or other automated input, such as periodic or event-driven input. The image frames may originate from at least one sensor, camera, or other imaging device, as described herein. Some non-limiting examples include an oral camera, an endoscope, a therioscope, a laparoscope, an ultrasound scanner, an X-ray detector, a computed tomography scanner, a positron emission tomography scanner, a magnetic resonance imaging scanner, an optical coherence tomography scanner, and a confocal microscope.
[0210] The image frames may be generated directly from the imaging device when the image is captured, from another device such as a storage device capable of storing image data after the image is captured, or from any combination of the above. Thus, the image frames may be received from virtually any other source of image data as would be understood by one of ordinary skill in the art. Some non-limiting examples of storage devices include local storage media and network storage media (e.g., on a local network or on a remote resource on a network or cloud computing infrastructure) accessible via standard protocols, a user interface (UI), an application programming interface (API), etc.
[0211] Additional logic in the multi-layer perceptron 320 and / or other neural networks may be implemented to perform further processing, some examples of which are shown in Figures 2 and 4-6 herein.
[0212] Scoring and Aligning Image Frames 4 is a flowchart illustrating a method 400 for scoring and arranging image frames according to the enhanced image processing techniques described herein, according to some embodiments. Method 400 may be performed by processing logic, which may include hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions executing on a processing device), or a combination thereof. Not all steps of method 400 may be required in all cases to perform the enhanced techniques disclosed herein. Furthermore, some steps of method 400 may be performed simultaneously or in a different order than that shown in FIG. 4, as will be understood by those skilled in the art.
[0213] At 402, at least one processor 904 may be configured to receive a set of image frames, such as the image frames shown on the left side of Figures 7 and 8. The image frames may be captured by at least one sensor, camera, or other imaging device, as described herein. Some non-limiting examples include oral cameras, endoscopes, therioscopes, laparoscopes, ultrasound scanners, x-ray detectors, computed tomography scanners, positron emission tomography scanners, magnetic resonance imaging scanners, optical coherence tomography scanners, and confocal microscopes.
[0214] Any given image frame in the set of image frames may be generated in response to automated or event-driven input, manual input from a user, or any sequence or set of imaging procedures prescribed by a given imaging device. The generation of image frames may be based on sampling, such as, but not limited to, random sampling, sampling by time, sampling by position, sampling by frequency, sampling by relative position, sampling by relative proximity, and sampling by change in value of at least one score. For example, sampling by relative proximity may enable biologically based samples that focus on parts of an organ most likely to provide an informative representation of a particular disease signature.
[0215] An image frame, for purposes of this example, may be defined as an image that represents (shows, depicts) a portion of an organ that may exhibit, for example, features potentially indicative of a particular disease. In some cases, the features may indicate symptoms or other manifestations of a particular disease. In some embodiments, an image frame may include images from multiple cameras or sensors taken substantially simultaneously, e.g., images from bug-eye camera devices. An image frame need not be limited to a single output of one sensor or camera. Sensor 8 Each image frame can be received directly from the imaging device when the image is captured, from another device internal or external to the imaging device, such as a storage device capable of storing image data after the image is captured, or from any combination of the above. Accordingly, image frames can be received from virtually any other source of image data as would be understood by one skilled in the art. Some non-limiting examples of storage devices include local storage media and network storage media (e.g., on a local network or on a remote resource on a network or cloud computing infrastructure) accessible via standard protocols, a user interface (UI), an application programming interface (API), or the like.
[0216] According to some embodiments, at 404, the processor 904 may assign at least one score to the image frame based at least in part on the output of the ANN responsive to the input of 510. According to at least the example described with respect to 510 above, a score for the entire image frame may be determined based on one or more snippets or regions of interest within the image frame. Assigning such image frame scores to image frames within the set of image frames may allow the image frames of the set of image frames to be sorted, ranked (ascending or descending), or otherwise arranged, as further described with respect to 408 below.
[0217] At 406, processor 904 may assign a second score to each image frame having a first score within a specified range. For example, in embodiments in which the first score relates to the information (or reliability) of the image frame, 406 may, as the case may be, selectively apply only to image frames having at least a minimum first score of informativeness or reliability, and ignore image frames having scores below a given threshold, since such image frames may not be sufficiently informative or reliable to yield a meaningful second score.
[0218] In this example, the second score may refer to the intensity or severity of a disease feature in the portion of the target organ depicted in the image frame. However, in some embodiments, the second score may additionally or alternatively refer to any quantifiable metric or qualitative property, such as, for example, any image property or image frame metadata. Furthermore, the specified range may have upper and / or lower bounds rather than just a threshold. Thus, image frames having a first score value between any values x and y may be excluded from the assignment of a second score, while image frames having a first score value below x or above y may be assigned the second score, for example, at 406. Alternatively, the specified range may be discontinuous. Thus, in other embodiments, image frames having a first score value below x or above y may be excluded from the assignment of a second score, while image frames between x and y may be assigned the second score, for example, at 406.
[0219] At 408, the processor 904, according to some embodiments, can arrange the set of image frames according to at least one of the first score or the second score. Arranging the (set of) image frames in this manner further enables curating the set of image frames to be presented to a human reviewer for diagnosis, rather than simply by distance along its length, for example, in some embodiments, in a manner designed to reduce reviewer bias. If the human reviewer can reduce bias by techniques such as randomizing and classifying images in a random order by time, location, and / or predicted (estimated) score determined by an ANN, removing less reliable (less informative) image frames from the set and automatically presenting the classified images to the human reviewer can result in more accurate and consistent results. The sort order can also be determined by image frame metadata, such as timestamp, location, relative position within a given organ, and at least one image property that can be analyzed by the processor 904. The processor 904 can additionally or alternatively remove image frames from the set based on, for example, their corresponding metadata.
[0220] Furthermore, in some embodiments, artificial intelligence (AI) can be used to perform more consistent and reliable automated diagnosis of a disease of interest. In the example of intestinal imaging, application of the enhanced technology described herein can enable significantly improved reliability in diagnosing celiac disease, to name one non-limiting example. Furthermore, diagnosis can be performed manually or automatically in response to observation of a given patient's reaction to irritants or in response to treatment administered to a given patient for a disease.
[0221] Such diagnosis can be facilitated by the enhanced techniques described herein even when a set of image frames contains largely homogeneous images, with the exception of certain landmark features of the intestinal tract (e.g., the pylorus, duodenum, cecum, etc.). Much of the small intestine may appear the same as other parts of the small intestine in any given section. This homogeneity can confuse or bias human reviewers, leading them to assume no changes even when subtle yet measurable differences exist. Such homogeneity can also cause differences to be lost in many conventional edge detection algorithms, depending on thresholds, smoothing, gradients, etc. However, by using the enhanced techniques described herein, detection of even subtle features of certain diseases can be significantly improved, resulting in more accurate diagnoses of those diseases.
[0222] Selecting a snippet (region of interest) within an image frame 5 is a flow chart illustrating a method 500 for identifying and / or selecting image frames according to the enhanced image processing techniques described herein, in accordance with some embodiments. Method 500 may be implemented using hardware (e.g., circuitry, dedicated logic, programmable logic, etc.). 5. The method 500 may be performed by processing logic, which may include hardware (e.g., logic, microcode, etc.), software (e.g., instructions executing on a processing device), or a combination thereof. Not all steps of method 500 may be required in all cases to implement the enhanced techniques disclosed herein. Furthermore, some steps of method 500 may be performed simultaneously or in a different order than that shown in FIG. 5, as will be understood by one of ordinary skill in the art.
[0223] Method 500 shall be described with reference to FIGS. 2, 3, and 6-8. However, method 500 is not limited to only those exemplary embodiments. The steps of method 500 may be performed by at least one computer processor coupled to at least one memory device. Exemplary processor and memory device(s) are described below with respect to 904 of FIG. 9. In some embodiments, method 500 may be performed by system 300 of FIG. 3, which may further include at least one processor and memory such as those of FIG. 9.
[0224] In 502, at least one processor 904 may be configured to receive a set of image frames, such as those shown on the left side of Figures 7 and 8. The image frames may originate from at least one sensor, camera, or other imaging device, as described herein. Some non-limiting examples include an oral camera, an endoscope, a therioscope, a laparoscope, an ultrasound scanner, an X-ray detector, a computed tomography scanner, a positron emission tomography scanner, a magnetic resonance imaging scanner, an optical coherence tomography scanner, and a confocal microscope.
[0225] Any given image frame in the set of image frames may be generated in response to automated or event-driven input, manual input from a user, or any sequence or set of imaging procedures prescribed by a given imaging device. The generation of image frames may be based on sampling, including, but not limited to, random sampling, sampling by time, sampling by position, sampling by frequency, sampling by relative position, sampling by relative proximity, and sampling by, for example, change in value of at least one score. For example, sampling by relative proximity may enable biologically based samples that focus on parts of an organ most likely to provide an informative representation of a particular disease feature.
[0226] An image frame, for purposes of this example, may be defined as an image that represents (shows, depicts) a portion of an organ that may exhibit, for example, features potentially indicative of a particular disease. In some cases, the features may indicate symptoms or other manifestations of a particular disease. In some embodiments, an image frame may include images from multiple cameras or sensors taken substantially simultaneously, e.g., images from bug-eye camera devices. An image frame need not be limited to a single output of one sensor or camera.
[0227] Each image frame can be received directly from the imaging device when the image is captured, from another device, such as a storage device, capable of storing image data after the image is captured, or from any combination of the above. Accordingly, image frames can be received from virtually any other source of image data, as would be understood by one skilled in the art. Some non-limiting examples of storage devices include local storage media, network storage media (e.g., on a local network or on a remote resource on a network or cloud computing infrastructure) accessible via standard protocols, a user interface (UI), an application programming interface (API), or the like.
[0228] At 504, the processor 904 may be configured to determine that an edge value exceeds a predetermined threshold in an image frame of the set of image frames. This determination, in some embodiments, may ultimately be used to determine the informativeness of the image frame (how likely the image frame is to indicate the presence, absence, or severity of a disease feature) and / or to determine the severity of a given feature (sometimes referred to as disease intensity). In making this determination, it is assumed that the analyzed image frame is sufficiently informative. In some embodiments, this determination may be attempted at any given image frame to identify one or more edges within the image frame that may form the basis of one or more snippets of the image frame (see FIGS. 7 and 8 ), or if the image frame does not have any edges that are useful, for example, for determining a particular disease feature.
[0229] To determine that an edge value exceeds a predetermined threshold, any of a variety of edge detection techniques can be used, such as, for example, calculating a local maximum or zero crossing of a differential operator or Laplacian function, or other differential edge detection methods, to name a few non-limiting examples. Any threshold value for edge detection purposes described herein may be determined by various techniques to balance signal and noise in a given image set, such as, for example, using adaptive thresholding techniques and / or threshold hysteresis. Additional steps may also accompany edge detection, such as Gaussian smoothing, transforming specific color ranges, separating color channels, and calculating gradients and / or gradient estimates.
[0230] In some embodiments, further aiding edge detection can be the use of at least one machine learning (ML) algorithm, neural network (NN), and / or random forest (RF). For example, an ML algorithm can be trained using images from inside the intestinal tract to separate a mucosal foreground from a background, or to separate obstructions (floating material, bile flow, etc.) or other artifacts (air bubbles, bright spots, reflections, etc.) in, for example, intestinal image frames. While such anomalies can reduce the effectiveness of relatively simple edge detection algorithms, neural networks, machine learning, random forests, or any combination thereof, can be configured in some embodiments to detect such images or regions as spurious (uninformative or unreliable) and remove them from consideration in the set of image frames.
[0231] For example, bright spots can be detected by at least two of the three channels of an RGB (red-green-blue) image, each of which has a relatively high intensity. Beyond the bright spots, bright lines can indicate air bubbles. In some embodiments, bile can be determined by casting the RGB color space to PBG (pink-brown-green) and detecting high-intensity green channels. These are just a few examples of how some artifacts can be determined when analyzing intestinal images. Image frames with such artifacts can be classified as low or uninformative.
[0232] Mucosal surfaces may be determined to have at least one strong edge (detected with a high-threshold edge filter). For example, image frames containing mucosal surfaces may be determined to be more informative than other image frames or relatively more informative with respect to determining celiac disease and its severity. Among informative image frames free of obstructive artifacts, the severity of disease features can be reliably scored, for example, based on edge and surface features corresponding to the intestinal surface. Such scoring may be performed manually, for example, by a human expert (e.g., a gastroenterologist) reviewing image frames within a set of image frames. In some embodiments, such scoring may be performed automatically, for example, by applying similar or different algorithms for edge and / or surface features.
[0233] At 506, processor 904 may be configured to extract at least one region of interest from a given image frame based on the determination of 504. For purposes of method 500, a region of interest may include at least one snippet of an image frame (e.g., as shown in FIGS. 7 and 8). Specifically, a snippet may be defined in some embodiments as a region of interest that may have a particular area, shape, and / or size, for example, encompassing a given edge value (e.g., exceeding a predetermined edge detection threshold) within the image frame that is identified as containing a mucosal surface and / or an edge of a mucosal surface. A snippet may be defined in some embodiments as having a small area relative to the overall image frame size, for example, 0.7% of the total image frame area per snippet. Other embodiments may define a snippet as up to 5% or 25% of the image frame area, depending on other factors, such as the actual area of the target organ depicted in each image frame. Where an image frame may be defined as relatively small (e.g., a smaller component of a larger image captured at a given time), in some embodiments, a snippet may be defined as potentially occupying a relatively large amount of image frame area.
[0234] At 508, at least one region of interest, such as any extracted from a given image frame based on the determination of 504, may be oriented (or reoriented) in a particular direction. In some embodiments, the particular direction may be relative to an edge within the region of interest. In some cases, within the intestine, for example, this orientation may be determined by ensuring that an edge of the mucosa is at the bottom of the oriented region of interest.
[0235] One example embodiment may operate by measuring the magnitude of a vector at a point within the region of interest relative to the centroid of a horizontal line within a circular region around each measured point. If the vector is determined to exceed a certain threshold (e.g., 25% of the radius of the circular region), this determines the bottom of the region of interest, and a rectangular patch may be drawn (e.g., parallel to the vector, with a length of x percent of the width of the image frame and a height of y percent of the height of the image frame). Each of these (re)oriented snippets may be evaluated separately from the image frame, thereby providing a score for the entire image frame, e.g., with respect to informativeness and / or severity of disease features, as further described below at 510. In some embodiments, if there is no strong vector (e.g., 25% of the radius of the circular region), orientation may be skipped.
[0236] At 510, the at least one region of interest may be input to an artificial neural network (ANN) to determine at least one score for the image frame. The ANN may be at least one of a feedforward neural network, a recurrent neural network, a modular neural network, or a memory network, to name a few non-limiting examples. For example, in the case of a feedforward neural network, in some embodiments, it may further correspond to at least one of a convolutional neural network, a probabilistic neural network, a time-delay neural network, a perceptron neural network, an autoencoder, or any combination thereof.
[0237] In further embodiments based on convolutional neural network (CNN) processing, for example, such CNNs can further integrate at least some filters (e.g., edge filters, horizon filters, color filters, etc.). These filters may include, for example, some edge detection algorithms and / or predetermined thresholds, as described with respect to 504 above. For a desired result, a given ANN, such as a CNN, may in some embodiments be designed, customized, and / or modified in various ways, for example, according to some criteria. In some embodiments, the filters may be of different sizes. Convolutional filters may be implemented, for example, by a CNN, and may be generated using training data. The CNN-based filter may be trained using a random selection of data or noise generated therefrom. In some embodiments, the CNN-based filter may be configured to output additional information regarding the reliability of the regression, e.g., for predicting features identified in a region of interest. In general, the CNN-based filter may improve the reliability of the results, at the expense of increased complexity in the design and implementation of the filter.
[0238] To give one non-limiting example, the CNN may be a pseudo-regression CNN based on a DenseNet architecture (a dense convolutional neural network architecture or a residual neural network with multi-layer skip). More specifically, in a further embodiment according to this example, the CNN may be based on DenseNet121 and configured, for example, to have one dense block (block configuration=(4,0)) that is four convolutional layers deep, 16 initial features, and more than 16 initial features per convolutional layer.
[0239] Furthermore, for a given neural network and set of image frames, regions of interest, and / or snippets, in some embodiments, the number of connected neurons in each layer of the (deep) neural network can be adjusted, for example, to accommodate smaller image sizes. In deep neural networks, the outputs of convolutional layers can feed transition layers for further processing. The outputs of these layers can be merged into a densely connected multilayer perceptron using multiple replications (see Figure 1 above). In one embodiment, multiple layers can be merged with three layers using five replications (e.g., to survive dropout) to obtain a final output of 10 classes, each class representing a decade ranging, for example, from 0 to 3. According to another example of a pseudo-regressive CNN, during training, locations can be obscured (to prevent overfitting) by adding normally distributed noise with a standard deviation of 0.2.
[0240] By way of further example, the score may be refined and generalized to the image frame as a whole based on one or more snippets or regions of interest within the image frame. One way to achieve this result is to use a pseudo-regression CNN, in which a relatively small percentage of image frames in the training set are withheld from the CNN. For example, if 10% of the image frames are withheld and the snippets each produce a value of 10 after CNN processing, half of the withheld (reserved) snippets can be used to train a 100-tree random forest (RF) to estimate, in one embodiment, the local curve value. This estimate, along with the 10-value output from the other half of the reserved training snippets, can be used to build another random forest that can, for example, estimate the error.
[0241] According to some embodiments, at 512, the processor 904 may assign at least one score to the image frame based at least in part on the output of the ANN responsive to the input of 510. According to at least the examples described with respect to 510 above, a score for the entire image frame may be determined based on one or more snippets or regions of interest within the image frame. By assigning such image frame scores to image frames within the set of image frames, the image frames of the set of image frames may be sorted, ranked (ascending or descending), or otherwise arranged.
[0242] Arranging the (set of) image frames in this manner further enables, for example, in some embodiments, curating the set of image frames to be presented to a human reviewer for diagnosis in a manner designed to reduce reviewer bias. The human reviewer may be able to reduce bias by techniques such as randomizing and categorizing images in a random order by time, location, and / or predicted (estimated) score determined by the ANN, thereby eliminating unreliable (less informative) images from the set. ) Removing image frames and automatically presenting the classified images to a human reviewer may result in more accurate and consistent results. The sort order may also be determined by image frame metadata, such as timestamp, location, relative position within a given organ, and at least one image property that may be analyzed by processor 904. Processor 904 may additionally or alternatively remove image frames from the set based on, for example, their corresponding metadata.
[0243] Furthermore, in some embodiments, artificial intelligence (AI) can be used to perform more consistent and reliable automated diagnosis of the disease of interest. Thus, in the example of intestinal imaging, application of the enhanced technology described herein can enable significantly improved reliability in diagnosing celiac disease, to name one non-limiting example. Furthermore, diagnosis can be performed manually or automatically in response to observation of a given patient's reaction to irritants or in response to treatment administered to a given patient for the disease.
[0244] Such diagnosis can be facilitated by the enhanced techniques described herein even when a set of image frames contains largely homogeneous images. Except for certain landmark features of the intestinal tract (e.g., the pylorus, duodenum, cecum, etc.), much of the small intestine may appear the same as other parts of the small intestine in any given segment. This homogeneity can confuse or bias human reviewers, leading them to assume no changes even when subtle yet measurable differences exist. Such homogeneity can also cause differences to be lost in many conventional edge detection algorithms, depending on thresholds, smoothing, gradients, etc. However, by using the enhanced techniques described herein, detection of even subtle features of certain diseases can be significantly improved, resulting in more accurate diagnosis of those diseases.
[0245] Training machine learning algorithms 6 is a flowchart illustrating a method 600 for training a machine learning (ML) algorithm using another ML algorithm, according to some embodiments. Method 600 may be performed by processing logic, which may include hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions executing on a processing device), or a combination thereof. Not all steps of method 600 may be required in all cases to perform the enhanced techniques disclosed herein. Furthermore, some steps of method 600 may be performed simultaneously or in a different order than shown in FIG. 6, as will be understood by those skilled in the art.
[0246] Method 600 shall be described with reference to FIGS. 2, 3, and 7-9. However, method 600 is not limited to only those exemplary embodiments. The steps of method 600 may be performed by at least one computer processor coupled to at least one memory device. Exemplary processor and memory device(s) are described below with respect to 904 of FIG. 9. In some embodiments, method 600 may be performed by system 300 of FIG. 3, which may further include at least one processor and memory such as those of FIG. 9.
[0247] In 602, at least one processor 904 may be configured to receive a first set of image frames, such as the image frames shown on the left side of FIGS. 7 and 8, and a corresponding initial set of metadata. The image frames may originate from at least one sensor, camera, or other imaging device, as described herein. Some non-limiting examples include an oral camera, an endoscope, a therioscope, a laparoscope, an ultrasound scanner, an X-ray detector, a computed tomography scanner, a positron emission tomography scanner, a magnetic resonance imaging scanner, an optical coherence tomography scanner, and a confocal microscope. Examples include microscopes.
[0248] Any given image frame in the set of image frames may be generated in response to automated or event-driven input, manual input from a user, or any sequence or set of imaging procedures defined by a given imaging device. Image frame generation can be based on sampling, including, but not limited to, random sampling, time-based sampling, location-based sampling, frequency-based sampling, relative position-based sampling, relative proximity-based sampling, and sampling based on, for example, a change in the value of at least one score. For example, relative proximity-based sampling can enable biology-based sampling, focusing on parts of an organ most likely to provide an informative representation of a particular disease signature. Additional information regarding imaging is described above with respect to FIG. 5.
[0249] The first metadata set may include, for example, but not limited to, a timestamp, a location, a relative location within a given organ, and at least one image property that can be analyzed by processor 904. Processor 904 may additionally or alternatively add or remove image frames to or from the set based on, for example, the image's corresponding metadata.
[0250] At 604, the processor 904 may input the first set of image frames and the corresponding first set of metadata to a first ML algorithm. The first ML algorithm may include any suitable algorithm, such as algorithms directed to regression, deep learning, reinforcement learning, active learning, supervised learning, semi-supervised learning, unsupervised learning, and other equivalent aspects within the scope of ML. The ML algorithm, in some embodiments, may rely on an artificial neural network, such as, for example, a deep convolutional neural network.
[0251] At 606, the processor 904 may generate a second set of image frames and corresponding metadata sets based on the first set of image frames. This second set of image frames and / or corresponding metadata sets may, in some embodiments, be based on at least one output value from the first ML algorithm. Such output may serve as a basis for evaluating the performance and accuracy of the first ML algorithm, for example. In some embodiments, such output may also serve as input to one or more other ML algorithm(s), such as for training the other ML algorithm(s), as shown at 608 below. In this example, the one or more other ML algorithm(s) may be of the same type as the first ML algorithm, a different type from the first ML algorithm, or a combination thereof.
[0252] At 608, at least a portion of the output generated at 606 may be input, such as by processor 904, in the form of a second set of image frames and their corresponding second set of metadata being provided to a second ML algorithm. Depending on the output of the first ML algorithm (e.g., at 606), the second set of image frames may be a subset of the first set of image frames and / or may be otherwise refined. As a further example, the corresponding second metadata may, in some embodiments, include more detail than the first set of metadata depending on the output of the first ML algorithm.
[0253] Depending on what the first and second ML algorithms are and how they may operate, such a configuration may allow ML algorithms to cooperatively train other ML algorithms to improve the performance and effectiveness of one or more ML algorithms with at least a given processing task(s) in the immediate vicinity, which may be further refined by 610 below.
[0254] At 610, the processor 904 may generate a refined set of image frames, refined metadata corresponding to the refined set of image frames, or a combination of the above. The refined set of image frames, in some embodiments, may be different from either the first set of image frames or the second set of image frames. The refined metadata may similarly be a subset or superset of the first metadata set and / or the second metadata set, or, in some embodiments, may not overlap with either the first metadata set or the second metadata set. In some embodiments, the refined image frames and / or refined metadata may be suitable for presentation as a diagnosis, as informational images for review, or some combination thereof, to name a few examples.
[0255] Method 600 is disclosed in the order shown above in this exemplary embodiment of Figure 6. However, in practice, the operations disclosed above may be performed sequentially in any sequence, along with other operations, or multiple operations performed simultaneously may be performed simultaneously, or any combination of the above.
[0256] 4-6 may provide at least the following advantages: By using a computer-implemented method employing the above-described enhanced technology, images can be scored and evaluated in an objective manner that, to date, has not otherwise been achieved by human reviewers (gastroenterologists) for at least certain bowel disorders. Thus, this application of the enhanced technology described herein achieves an objective methodology and, therefore, a sound basis for efficiently observing, diagnosing, and treating patients with such bowel disorders, thereby avoiding the costly and unreliable trial-and-error of conventional methods.
[0257] Orienting and extracting regions of interest (snippets) from image frames FIG. 7 illustrates an example 700 of identifying, orienting, and extracting at least one region of interest from an image frame, for example, in the form of a snippet, according to some embodiments.
[0258] An image frame is shown at 710, and certain edge values are identified at 712 as exceeding a predetermined threshold, having strong edge values. Regions of interest around edges with strong edge values may be extracted and / or oriented in the form of snippets at 720. It will be appreciated that extraction and (re)orientation may occur in any sequence. Further processing may be performed at 730 and 740 to further refine the results with respect to scoring, such as to determine predictions at the level of individual snippets and for purposes of considering multiple snippets within an image frame (see FIG. 8 below).
[0259] At 730, the color image (snippet) may be converted to a different color map, such as grayscale. This conversion may be performed in a variety of ways, from a linear flattening approximation to desaturation per color channel. One implementation may be in the form of, for example, an "rgb2gray" function as part of various standard libraries for mathematical operations. In some embodiments, a geometric mean may be used to convert the color map to grayscale, for example, which may result in improved contrast for some use cases.
[0260] At 740, the image intensities of the snippets may be normalized. The normalization may be performed with respect to a curve (e.g., Gaussian, Lorentzian, etc.). In some embodiments, the image intensities for the snippets may also be referred to as the disease severity prediction for each snippet. In some embodiments, the image intensity may refer to the disease severity score for the entire image frame. In some embodiments, the image intensity may refer to a particular image characteristic, such as brightness or saturation.
[0261] Orienting and Extracting Multiple Regions of Interest (Snippets) from an Image Frame FIG. 8 illustrates an example 800 of identifying, orienting, and extracting multiple regions of interest from an image frame in the form of snippets, according to some embodiments.
[0262] An image frame is shown at 810, and certain edge values are identified as exceeding a predetermined threshold at 822, 832, and 842, each having a strong edge value. For each of 822, 832, and 842, a respective region of interest in the form of snippets 820, 830, and 840 can be extracted, (re)oriented, grayscaled, and / or normalized. Snippets 820, 830, and 840, among other possible snippets, can be evaluated individually, such as based on a first algorithm for generating a prediction for each snippet. Based on a second algorithm, the predicted scores can be weighted and combined to form a composite image score in the form of a single representative score that is assigned to the corresponding image frame as a whole, applying any of the enhanced techniques described herein.
[0263] Adaptive sampling of image frames FIG. 10 shows an example 1000 of reassigning scores for already scored image frames with respect to a PDF that correlates with relative position, according to some embodiments.
[0264] In this example, the relative position may be in terms of distance along the small intestine. In some embodiments, the previously scored frames may correspond to positions along the small intestine, which in turn may correspond to probability values. Executing a given algorithm, at least one processor 904 may calculate the area under the curve (AUC) of the PDF, such as the AUC for any given point, which may correspond to a particular sample image frame and / or one of the minimum or maximum limits of the PDF or relative position (e.g., 0 or 1), according to some embodiments.
[0265] After running the above given algorithm, for any two points along the PDF with the largest AUC between them in a given set of image frames, another algorithm can evaluate the cost function of any intervening or adjacent sample image frames with respect to the informationality scores already calculated based on previously scored frames. Further details regarding determining informationality scores are described elsewhere herein, for example, with respect to 504 in Figure 5 above.
[0266] If the cost function determines that the AUC can be more evenly balanced between sample image frames with different (or different scores), the different (or differently scored) image frame may be selected as the sample image frame for subsequent processing, which may include snippet selection, which, according to some embodiments, may be performed in accordance with Figures 1-3, 7, 8, and 12, as further described herein.
[0267] Weighted selection of image frames 11 illustrates an exemplary method 1100 for weighted selection of image frames, according to some embodiments. Method 1100 may be performed by processing logic, which may include hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions executing on a processing device), or a combination thereof. Not all steps of method 1100 may be required in all cases to perform the enhanced techniques disclosed herein. Furthermore, some steps of method 1100 may be performed simultaneously or in a different order than that shown in FIG. 11, as will be understood by those skilled in the art.
[0268] Method 1100 shall be described with reference to FIGS. 9-11. However, method 1100 is not limited to only those exemplary embodiments. The steps of method 1100 may be performed by at least one computer processor coupled to at least one memory device. Exemplary processor and memory device(s) are described below with respect to 904 of FIG. 9. In some embodiments, method 1100 may be performed by system 300 of FIG. 3, which may further include at least one processor and memory such as those of FIG. 9.
[0269] At 1102, at least one processor 904 may be configured to determine, for a given point along a PDF correlated with a relative position, a closest point corresponding to a given sample image frame of the set of image frames, the given sample image frame being the sample image frame closest to the given point along the PDF correlated with the relative position. An example of a PDF correlated with a relative position is shown in FIG. 10 , along with points corresponding to the relative positions of the sample image frames and correlated with probability values of the PDF. In some embodiments, not all points of a relative position correspond to a sample image frame, but any given point may have a closest point corresponding to a sample image frame in the set of image frames, unless the set of image frames is empty.
[0270] At 1104, the processor 904 may be configured to calculate an integral along the PDF between a given point and the closest point corresponding to the given sample image frame (the closest sample image frame). Various integration techniques may be implemented for performing 1104, such as measuring the AUC between specified points along the PDF.
[0271] At 1106, the processor 904 may be configured to divide the integral obtained from 1104 by the mean value of the PDF. In some embodiments, the mean value may be the mean value of the PDF for a given organ relative location. In some embodiments, the mean value may be the median value of the PDF. In some embodiments, the mean value of the PDF may be calculated for a given subset or range(s) of relative locations of a given organ (e.g., all but the mucus proximal to the small intestine), for example, depending on the target symptom being measured or monitored. The result of this division 1106 may indicate a measure of expected benefit for a given sample image frame based on the PDF for the relative location. This division, in some embodiments, may result in a normalized estimated benefit measure (B) with a mean value of 1, which may depend on the PDF.
[0272] At 1108, processor 904 may perform a smoothing operation (e.g., a linear transform, a matrix transform, a moving average, a spline, a filter, a similar algorithm, or any combination thereof) on the information scores of the set of image frames, thereby reassigning the information value of at least one image frame in the set of image frames. In the same or a separate operation, processor 904 may divide the information score of a given sample image frame by the mean value of the scale to which the information score is assigned. For example, as described elsewhere herein, on an informationality scale of 0 to 3, the mean value is 1.5. Dividing by the mean value as in 1108 may result in a normalized information measure (I), which in some embodiments has a mean value of 1, depending on the overall information content of the available sample image frames.
[0273] According to some embodiments, a nonlinear smoothing operation may reduce bias toward the selection of individual image frames, while a linear smoothing operation may improve the overall correlation of a given set of image frames. In some embodiments, removing outlier frames with convergence or divergence values above a predetermined threshold may improve the informative correlation of the remaining image frames. Additionally or alternatively, a polynomial transformation, e.g., a second-order polynomial trained to relate mean values to known curve values (e.g., PDF), may be used to smooth the remaining image frames. In other embodiments, for example, polynomials of other orders may be used.
[0274] At 1110, the processor 904 may calculate a penalty value based on the number of skipped image frames for a given set of image frames or a subset of image frames, e.g., corresponding to a given region or range of relative values. In some embodiments, the penalty may be calculated based on a Gaussian window having a width defined as a predetermined percentage of the measurement range of the relative position (e.g., 1 / 18 of the small intestine) and, in some embodiments, an AUC defined as a calculated fraction (e.g., 1 / 50) of the product of the normalized measurements of estimated benefit and informativeness (e.g., B*I / 50). The calculated fraction may be a factor of the total frames per selected sample image frame (e.g., 1 frame selected out of 50 frames). Other calculated or predetermined values may be used, for example, for other use cases, organs, and / or target symptoms, according to various embodiments. The sum of the calculated penalty values may be represented as P.
[0275] At 1112, the processor 904 may calculate a net desirability score (N) for a given sample image frame. In some embodiments, the calculation at 1112 may evaluate N where N=B*IP, where the terms B, I, and P may be defined as described above with respect to method 1100. In some embodiments, N may be calculated, for example, for each image frame in a set of image frames or a subset thereof. In some cases, to facilitate spacing of representative samples, rounding of the N value may be performed to improve separation of the sampled image frames and reduce the likelihood of over-fitting with a predictive model based on the sampled image frames.
[0276] At 1114, the processor 904 may determine whether to continue sampling image frames for subsequent evaluation (e.g., snippet selection). The determination at 1114 may involve one or more conditions, such as evaluating whether a desired number of frames have been obtained and / or whether the net desirability score of the set of image frames or a subset thereof is within a predetermined range of an acceptable sample.
[0277] Image transformation and reorientation by region of interest selection 12 illustrates an exemplary method 1200 for image transformation and reorientation with region-of-interest selection, according to some embodiments. Method 1200 may be performed by processing logic, which may include hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions executing on a processing device), or a combination thereof. Not all steps of method 1200 may be required in all cases to perform the enhanced techniques disclosed herein. Furthermore, some steps of method 1200 may be performed simultaneously or in a different order than that shown in FIG. 12, as will be understood by those skilled in the art.
[0278] Method 1200 shall be described with reference to FIGS. 2, 5, 7-9, and 12. However, method 1200 is not limited to only those exemplary embodiments. The steps of method 1200 may be performed by at least one computer processor coupled to at least one memory device. Exemplary processor and memory device(s) are described below with respect to 904 of FIG. 9. In some embodiments, method 1200 may be performed by system 300 of FIG. 3, which may further include at least one processor and memory such as those of FIG. 9.
[0279] At 1202, the at least one processor 904 can perform edge detection on the sample image frames of the set of sample image frames. Techniques are described elsewhere herein. The edge detection performed in 1202 may include, for example, the use of random forest (RF) prediction, such as with a neural network. In some embodiments, the edge detection performed in 1202 may be part of an iterative process that is performed entirely in 1202 or that is divided into multiple stages, some of which may be performed in 1202 and other parts of method 1200.
[0280] At 1204, the processor 904 may determine, from the set of sample image frames, sample image frames that include edge regions that form at least a certain percentage of the image frame area, e.g., exceeding a predetermined threshold. In some cases, the processor 904 may select sample image frames if they are determined at 1204 to have at least 3% edge region per frame area, and ignore or discard sample image frames that have less than 3% edge region per frame area. However, according to one embodiment, in other cases where more samples may be required, the threshold may be lowered to allow more frames to be processed, even if the frames have less edge region area. The determination of edge regions may result from the processor 904 executing at least one specified stochastic or randomized algorithm, such as data masking, simulated annealing, stochastic perturbation, approximation, optimization, or any combination thereof, to name a few non-limiting examples.
[0281] At 1206, the processor 904 may convert the color map of at least one selected image frame, for example, from color (pink-brown-green, red-green-blue, etc.) to grayscale, if the image frame contains color. If the image frame is not in color, 1206 may be skipped. To convert from color to grayscale, the processor 904 may execute an algorithm that can remove color components from the original image (e.g., setting saturation and hue values to zero). In some embodiments, grayscale images may be encoded with color data corresponding to the grayscale, for example, for compatibility with a particular format, while other embodiments may work directly with gray image values, e.g., monochrome or relative luminance values.
[0282] In some embodiments, the red-green-blue (RGB) values can be averaged, e.g., the RGB values can be summed and divided by 3. In some embodiments, a weighted average of the RGB channels can be calculated, using the relative luminance or intensity of each channel as a weight for each component to determine the weighted average of the RGB components as a whole. In some embodiments, the conversion to grayscale can be performed by calculating the geometric mean of the RGB values, e.g., [ka] For the purpose of edge detection and evaluation of specific conditions in specific organs, contrast enhancement may be achieved by the geometric mean method of grayscaling (conversion from color to grayscale) to improve the reliability (and reproducibility) of the results. For other types of evaluation, e.g., different conditions and / or different organs, other grayscaling methods may be performed as appropriate for a given purpose.
[0283] At 1208, the processor 904 can convolve the grayscale image frames. The convolution at 1208 can identify or otherwise detect curved edges, folds, and / or other horizontal lines within the selected grayscale image frames. In some embodiments, for example, depending on the training of the CNN, edge curves can be successfully predicted with relatively stable convergence. Such horizontal lines (mucosal folds / edges) can be detected within the image. These may be detected via algorithms that can convolve monochrome (bi-level) filtered images at a 50% threshold (e.g., half black, half white) at various angles, and / or by convolving light or dark lines at various angles within the image. Depending on the data present in a given set of image frames, 1208 convolutions may be skipped.
[0284] At 1210, the processor 904 can identify candidate points within a given image frame, the candidate points being within edge regions that are spaced from one another by at least a predetermined minimum amount. For example, the predetermined minimum amount of space can be in absolute terms (e.g., 7 pixels); according to some embodiments, it can be in relative terms with respect to the included image frame, resolution, number or density of pixels (e.g., 4% of the image frame), etc., and / or a statistical analysis of each edge region within the given image frame (e.g., standard deviation of edge characteristics, where the candidate points have edge values at least one standard deviation higher than the mean edge value of each edge region).
[0285] Thus, any minimum spacing criteria described above may limit the number of possible regions of interest (e.g., snippets) within any given image frame. Similar to the minimum edge region threshold with respect to image frame area determined at 1204, the minimum distance threshold between candidate points or regions of interest may be reduced at 1210 if additional regions of interest are needed or desired for a given use case. In some embodiments, each image frame, e.g., candidate point, within the set of image frames or a subset thereof may be evaluated to identify candidate points, or at least until enough candidate points have been identified across enough samples to produce what is likely to be a reproducible result.
[0286] At 1212, the processor 904 may orient (or reorient) at least one region of interest within a given image frame. Further examples of orienting or reorienting regions of interest are further shown and described above with respect to the snippets (regions of interest) of Figures 2, 7, and / or 8. In some embodiments, which may include a use case of evaluating images taken from inside the small intestine to evaluate symptoms of celiac disease, the regions of interest (snippets) may be oriented or reoriented in a common direction relative to a common feature throughout the regions of interest, such that, for example, sections of the regions of interest depicting the mucosal wall are located at the bottom of each corresponding region of interest when the at least one region of interest is oriented or reoriented according to 1212.
[0287] According to some embodiments, examples of (re)oriented regions of interest (snippets) are shown in Figures 7 and 8. The orientation or reorientation can be repeated for each selected image frame of the set of image frames. If the regions of interest do not contain common features or common strong vectors, such as features (e.g., mucosa) that occupy at least 25% of the radius of the region of interest in any particular frame, then the orientation or reorientation of 1212 may be skipped. Additionally or alternatively, according to some embodiments, 1212 may be skipped if a rotation-invariant filter can be used for convolution with multiple existing orientations, e.g., an 8-way convolution.
[0288] At 1214, processor 904 may reject, ignore, or otherwise discard regions of interest that may result in spurious results. For example, regions of interest that may contain artifacts such as air bubbles, bile flow, or unclear images, to name a few non-limiting examples, may result in spurious results. Rejecting such regions of interest may improve the reproducibility of diagnoses resulting from any monitoring or evaluation or decision made based on the remaining regions of interest in a given image frame or set of image frames. Examples of how processor 904 may identify regions of interest for rejection are described further above, such as with respect to 504 in FIG. 5.
[0289] Some non-limiting examples include processing using neural networks, RF algorithms, ML algorithms, etc., filtering regions with high-intensity channels above a given threshold, or filtering for a particular color balance or color gradient in the corresponding original image frame, according to some embodiments. Additionally or alternatively, the gradients may be averaged with the standard deviation of their Laplacian function for another metric to reject frames, e.g., to assess blurriness. Statistical analysis (mean, standard deviation, etc.) of PGB gradients, multi-channel intensity values, detected dark lines / edges, mucosa scores, and / or intensity PDFs, or similar factors, may also be used to determine whether to reject, ignore, or otherwise discard a snippet or region of interest, according to some embodiments.
[0290] Any or all of the above steps may, in some embodiments, be performed as part of snippet selection 240, as shown and described further above with respect to Figure 2. Additionally or alternatively, any or all of the above steps may be performed as part of the processes shown in Figures 5 and 7-9, for example.
[0291] Image scoring and frame display to reduce bias in subsequent scoring 13 illustrates an exemplary method 1300 for image scoring and frame presentation to reduce bias in subsequent scoring, according to some embodiments. Method 1300 may be performed by processing logic, which may include hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions executing on a processing device), or a combination thereof. Not all steps of method 1300 may be required in all cases to perform the enhanced techniques disclosed herein. Furthermore, some steps of method 1300 may be performed simultaneously or in a different order than that shown in FIG. 13, as will be understood by those skilled in the art.
[0292] Method 1300 shall be described with reference to FIGS. 9-13. However, method 1300 is not limited to only those exemplary embodiments. The steps of method 1300 may be performed by at least one computer processor coupled to at least one memory device. Exemplary processor and memory device(s) are described below with respect to processor 904 of FIG. 9. In some embodiments, method 1300 may be performed by system 300 of FIG. 3, which may further include at least one processor and memory such as those of FIG. 9.
[0293] At 1302, at least one processor 904 may receive output of an imaging device. The output of the imaging device may include image frames forming at least a subset of a set of image frames that may depict the interior surface of a given patient's digestive tract, such as the small intestine. In some embodiments, not all image frames of the set may reliably depict the interior surface of the digestive tract due to poor quality of the given image frame (e.g., insufficient light, reflections, etc.) or obstructions between the image sensor and the interior surface (e.g., air bubbles, bile flow, etc.).
[0294] Membership of the subset is determined by whether the score exceeds a predetermined threshold, such as the first score assigned in 1306 below, or another previously assigned score, which may indicate, for example, the reliability, informativeness, or other measure of a given image frame. Additionally or alternatively, membership of the subset may be determined by other factors, such as relative location within the digestive tract, elapsed time, or other external metrics, which in some embodiments may include the likelihood of an image being symptomatic of a given disease in a given patient. It may be correlated with the reliability or informativeness of the corresponding image frame.
[0295] At 1304, the processor 904 may automatically decompose at least one image frame of the plurality of image frames into multiple regions of interest via at least one machine learning (ML) algorithm. The at least one region of interest may be defined by the processor 904 by determining that an edge value (e.g., within each image frame of the at least one image frame) exceeds a predetermined threshold. The edge value may be determined via at least one type of edge detection algorithm and may be compared by the processor 904 to the predetermined threshold. For a given edge having an edge value that exceeds the predetermined threshold, the region of interest may, in some embodiments, be defined by, for example, a minimum area, a radius or other perimeter of a neighborhood of the edge having a shape and size that may match a maximum percentage of the area of the corresponding image frame, and / or proximity to another region of interest in the corresponding image frame.
[0296] At 1306, the processor 904 may automatically assign a first score based at least in part on the edge value of each region of interest when decomposed from the at least one corresponding image frame determined at 1304. The automatic assignment may, in some embodiments, be performed by at least one neural network, which may implement at least one additional ML algorithm, such as random forest regression, as described further above. The automatic assignment of the first score may be responsive to an event, such as edge detection, sequential processing of the regions of interest, random processing of the regions of interest, etc.
[0297] In some embodiments, the first score may be attributed to the corresponding image frame as a whole, and the first score may be based on at least one region of interest within the corresponding image frame, e.g., to represent the informativeness of the entire corresponding image frame. In some embodiments, an instance of the first score may be automatically assigned to each region of interest, and a composite value of the first score may be calculated for the corresponding image frame.
[0298] At 1308, the processor 904 may automatically shuffle the set of image frames. In some embodiments, the shuffling may involve randomization, such as selecting and / or ordering by numbers generated by a pseudo-random number generator. Some embodiments may implement the shuffling systematically according to at least one algorithm calculated to reduce bias in subsequent review and / or scoring of the subset of image frames.
[0299] For example, image frames that are members of a set of image frames but not a subset of image frames, or image frames with a first score below a predetermined threshold, can be systematically interspersed in a fully automated manner among image frames with a first score that meets or exceeds the predetermined threshold. Thus, in such embodiments, for example, the shuffling algorithm can prevent or otherwise interrupt relatively long sequences of image frames that may have scores above or below the threshold. By avoiding or breaking up such sequences, the shuffling algorithm, according to some embodiments, can reduce bias or fatigue that may result from subsequent review and / or scoring by an automated ML algorithm and / or a manual human reviewer as a result of a relatively long sequence of image frames that exceed or fall below the predetermined threshold.
[0300] At 1310, the processor 904 can output a presentation of the set of image frames after shuffling. The presentation can be a visual display, such as a slide show, with a presentation of the shuffled image frames. An automatic presentation can automatically proceed in the shuffled order. The automatic presentation may be further arranged or presented according to additional organizational schemes, such as, for example, time intervals and / or position intervals. Additionally or alternatively, as a complement or alternative to any shuffling algorithm, the presentation may be organized in a manner that may further reduce bias and / or fatigue in subsequent review and / or scoring of subsets or sets of image frames.
[0301] For example, the organized presentation of image frames may include a temporal interval in which each of the image frames (or particular selected image frames) may, in some embodiments, be presented in a visual slideshow, e.g., by predetermined intervals, randomized intervals, or systematically separated intervals, such as with respect to the display time or visual interval (e.g., the distance between simultaneously displayed image frames) of any given image frame in succession. As used herein, "order" may, in some embodiments, refer to the order from one image frame to the next in a shuffled deck of image frames shuffled by 1308 above, although this need not be any predetermined order by, for example, any output of an imaging device or a gastrointestinal tract. Other forms of organizing the visual display or other presentation of members of a set of image frames to reduce bias and / or fatigue may be realized within the scope of this disclosure.
[0302] Some embodiments of method 1300 may further include assigning a second score to the representation of the set of image frames. The second score may include, for example, an expression of the severity of the symptom. In some embodiments, the second score may be assigned automatically, for example, via processor 904, or manually, for example, via a reviewer evaluating the representation of 1310.
[0303] After method 1300, the information can be shuffled or otherwise rearranged and presented in an organized manner by assigning associated values (e.g., edge values, scores, etc.) to information such as regions of interest, corresponding image frames, and / or (sub)sets of image frames. Thus, the information can be reorganized in new and unique ways with respect to the display of the same information.
[0304] Further iterations of method 1300 may be used to monitor or train at least one ML algorithm based at least in part on the output of 1310. The monitoring or training of the at least one ML algorithm may involve predefined start and end points for the output of the imaging device. After fully monitoring or training the at least one ML algorithm, the at least one ML algorithm may perform method 1300 in a fully automated manner, such as via at least one processor 904.
[0305] Example of a computer system Various embodiments may be implemented using one or more computer systems, such as, for example, computer system 900 shown in Figure 9. One or more computer systems 900 may be used to implement, for example, any of the embodiments described herein, as well as combinations and subcombinations thereof.
[0306] Computer system 900 may include one or more processors (also referred to as central processing units or CPUs), such as processor 904. Processor 904 may be connected to a bus or communication infrastructure 906.
[0307] The computer system 900 may also include user input / output device(s) 803 such as a monitor, keyboard, pointing device, etc., which can communicate with the communication infrastructure 906 via the user input / output interface(s) 902.
[0308] One or more of the processors 904 may be a graphics processing unit (GPU). In one embodiment, a GPU may be a processor, which is a specialized electronic circuit designed to process mathematically intensive applications. GPUs may have a parallel structure that is efficient for parallel processing of large blocks of data, such as mathematically intensive data common in computer graphics applications, images, video, vector processing, array processing, and cryptography (e.g., brute-force cracking), generating cryptographic hashes or hash sequences, solving partial hash reversal problems, and / or producing results for other proof-of-work calculations for some blockchain-based applications. GPUs may be particularly useful in at least the image analysis and ML aspects described herein.
[0309] Additionally, one or more of the processors 904 may include a coprocessor or other implementation of logic for accelerating cryptographic calculations or other specialized mathematical functions, such as a hardware-accelerated cryptographic coprocessor. Such accelerated processors may further include an instruction set(s) for acceleration using the coprocessor and / or other logic to facilitate such acceleration.
[0310] The computer system 900 may also include a main or primary memory 908, such as random access memory (RAM). The main memory 908 may include one or more levels of cache. The main memory 908 may store control logic (i.e., computer software) and / or data therein.
[0311] Computer system 900 may also include one or more secondary storage devices or secondary memory 910. Secondary memory 910 may include, for example, a main storage drive 912 and a removable storage drive or drive 914. Main storage drive 912 may be, for example, a hard disk drive or a solid state drive. Removable storage drive 914 may be, for example, a floppy disk drive, a magnetic tape drive, a compact disk drive, an optical storage device, a tape backup device, and / or any other storage device / drive.
[0312] The removable storage drive 914 may interface with a removable storage unit 918. The removable storage unit 918 may comprise a computer-usable or readable storage device having computer software (control logic) and / or data stored thereon. The removable storage unit 918 may be a floppy disk, magnetic tape, compact disk, DVD, optical storage disk, and / or any other computer data storage device. The removable storage drive 914 may read from and / or write to the removable storage unit 918.
[0313] Secondary memory 910 may include other means, devices, components, appliances, or other approaches for allowing computer programs and / or other instructions and / or data to be accessed by computer system 900. Such means, devices, components, appliances, or other approaches may include, for example, removable storage unit 922 and interface 920. Examples of removable storage unit 922 and interface 920 include a program cartridge and cartridge interface (such as found in a video game device), a removable memory chip (such as an EPROM or PROM) and associated socket, a memory stick, or the like. Examples of such removable storage devices include hard drives and memory sticks, USB ports, memory cards and associated memory card slots, and / or other removable storage units and associated interfaces.
[0314] Computer system 900 may further include a communications or network interface 924. Communications interface 924 may enable computer system 900 to communicate and interact with any combination of external devices, external networks, external entities, etc. (individually and collectively referred to by reference numeral 928). For example, communications interface 924 may enable computer system 900 to communicate with external or remote devices 928 via communications path 926, which may be wired and / or wireless (or a combination thereof), and which may include any combination of a LAN, a WAN, the Internet, etc. Control logic and / or data may be transmitted to and from computer system 900 via communications path 926.
[0315] The computer system 900 may also be any of a personal digital assistant (PDA), a desktop workstation, a laptop or notebook computer, a netbook, a tablet, a smartphone, a smartwatch or other wearable, an appliance, part of the Internet of Things (IoT), and / or an embedded system, or any combination thereof, to name a few non-limiting examples.
[0316] It should be understood that the framework described herein may be implemented as a method, process, apparatus, system, or article of manufacture, such as a non-transitory computer-readable medium or device. For illustrative purposes, the framework of the present invention may be described in the context of a distributed ledger that is publicly available, or at least available to untrusted third parties. One example of a contemporary use case is a blockchain-based system. However, it should be understood that the framework of the present invention may also be applied to other settings where sensitive or confidential information may need to pass by or through untrusted third parties, and that the technology is not limited to the use of distributed ledgers or blockchains.
[0317] The computer system 900 may be a client or server accessing or hosting any application and / or data via any delivery paradigm, including, but not limited to, remote or distributed cloud computing solutions; local or on-premise software (e.g., "on-premise" cloud-based solutions); "as a service" models (e.g., Content as a Service (CaaS), Digital Content as a Service (DCaaS), Software as a Service (SaaS), Managed Software as a Service (MSaaS), Platform as a Service (PaaS), Desktop as a Service (DaaS), Framework as a Service (FaaS), Backend as a Service (BaaS), Mobile Backend as a Service (MBaaS), Infrastructure as a Service (IaaS), Database as a Service (DBaaS), etc.); and / or hybrid models including any combination of the foregoing examples or other service or delivery paradigms.
[0318] Any applicable data structures, file formats, and schemas may be used, including but not limited to JavaScript Object Notation (JSON), Extensible Markup Language (XML), Yet Another Markup Language (YAML), Extensible Hypertext Markup Language (XHTML), Wireless Markup Language (WML), and the like. Language (WML), MessagePack, XML user interface The data structure may be a representation of the underlying data structure, format, or schema, either alone or in combination with the XHTML, XHTML-based XML, XML, XML, or XHTML-based XML, or any other functionally similar representation. Alternatively, the data structure, format, or schema may be proprietary, either exclusively or in combination with known or open standards.
[0319] The associated data, files, and / or databases may be stored, retrieved, accessed, and / or transmitted in a human-readable format, such as a numerical, textual, graphical, or multimedia format, as well as various types of markup languages, among other possible formats. Alternatively, or in combination with the above formats, the data, files, and / or databases may be stored, retrieved, accessed, and / or transmitted in binary, encoded, compressed, and / or encrypted format, or any other machine-readable format.
[0320] The interfaces or interconnections between the various systems and layers may use any number of protocols, program frameworks, floorplans, or application programming interfaces (APIs), such as, but not limited to, the Document Object Model (DOM), Discovery Service (DS), NSUserDefaults, Web Services Description Language (WSDL), Message Exchange Patterns (MEP), Web Distributed Data Exchange (WDDX), Web Hypertext Application Technology Working Group (WHATWG), HTML5 Web Messaging, Representational State Transfer (REST or RESTful Web Services), Extensible User Interface Protocol (XUP), Simple Object Access Protocol (SOAP), XML Schema Definition (XSD), XML Remote Procedure Call (XML-RPC), or any other mechanism, open or proprietary, that may achieve similar functionality and results.
[0321] Such interfaces or interconnections may also utilize Uniform Resource Identifiers (URIs), which may further include Uniform Resource Locators (URLs) or Uniform Resource Names (URNs). Other forms of uniform and / or unique identifiers, locators, or names may be used exclusively or in combination with such forms.
[0322] Any of the above protocols or APIs may interface with or be implemented in any programming language, procedural, functional, or object-oriented, and may be compiled or interpreted, including, by way of non-limiting example, C, C++, C#, Objective-C, Java, Swift, Go, Ruby, Perl, Python, JavaScript, WebAssembly, or virtually any other language, with any other libraries or schemas, in any kind of framework, runtime environment, virtual machine, interpreter, stack, engine, or similar mechanism, such as, but not limited to, Node.js, V8, Knockout, jQuery, Dojo, Dijit, OpenUI5, AngularJS, Express.js, Backbone.js, Ember.js, DHTMLX, Vue, React, Electron, among many other non-limiting examples.
[0323] In some embodiments, a tangible, non-transitory apparatus or article of manufacture that includes a tangible, non-transitory computer usable or readable medium having control logic (software) stored thereon is Also referred to herein as a computer program product or program storage device, including but not limited to computer system 900, main memory 908, secondary memory 910, and removable storage units 918 and 922, as well as tangible articles of manufacture incorporating any combination of the above. Such control logic, when executed by one or more data processing devices (such as computer system 900), may cause such data processing devices to operate as described herein.
[0324] Based on the teachings contained herein, it will be apparent to one skilled in the relevant art(s) how to make and use embodiments of the present disclosure using data processing devices, computer systems, and / or computer architectures other than those shown in Figure 9. In particular, embodiments may operate with software, hardware, and / or operating system implementations other than those described herein.
[0325] conclusion It is understood that the Detailed Description section is intended to be used to interpret the claims, and not any other sections, which may set forth one or more, but not all, example embodiments as contemplated by the inventor(s), and therefore do not limit the scope of the disclosure or the appended claims in any way.
[0326] While this disclosure describes exemplary embodiments for exemplary fields and applications, it should be understood that the disclosure is not limited thereto. Other embodiments and modifications of the invention are possible and are within the scope and spirit of the present disclosure. For example, without limiting the generality of this paragraph, embodiments are not limited to the software, hardware, firmware, and / or entities shown in the figures and / or described herein. Moreover, embodiments (whether or not explicitly described herein) have significant utility for fields and applications beyond the examples described herein.
[0327] The embodiments have been described herein using functional blocks illustrating implementation of specified functions and relationships thereof. The boundaries of these functional building blocks have been arbitrarily defined herein for convenience of description. Alternative boundaries may be defined as long as the implementation and relationships of the specified functions are properly performed. Also, alternative embodiments may execute the functional blocks, steps, operations, methods, etc. using an order different from that described herein.
[0328] References herein to “one embodiment,” “embodiment,” “example embodiment,” “some embodiments,” or similar phrases indicate that the described embodiment may include a particular feature, structure, or characteristic, but not all embodiments necessarily include the particular feature, structure, or characteristic. Moreover, such phrases may not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in connection with an embodiment, it would be within the knowledge of one of ordinary skill in the relevant art(s) to incorporate such feature, structure, or characteristic into other embodiments, whether or not explicitly mentioned or described herein. Furthermore, some embodiments may be described using the terms “coupled” and “connected,” along with their derivatives. These terms are not necessarily intended as synonyms for each other. For example, some embodiments may be described using the terms “connected” and / or “coupled” to indicate that two or more elements are in direct physical or electrical contact with each other. However, the term “coupled” can also mean that two or more elements cooperate or interact with each other, even though they are not in direct contact with each other.
[0329] The breadth and scope of the present disclosure should not be limited by any of the above-described exemplary embodiments, but should be defined only in accordance with the following claims and their equivalents.
Claims
1. 1. A computer-implemented method comprising: Operations performed by at least one processor, said at least one processor including: receiving a set of image frames; at least a subset of the set of image frames depicting an interior surface of a digestive tract of a given patient; determining at least one region of interest within an image frame of the set of image frames; determining the at least one region of interest based on an edge value and a predetermined threshold; extracting the at least one region of interest from the image frames of the set of image frames; orienting the at least one region of interest based on the edge value; determining at least one score corresponding to the image frame using the oriented at least one region of interest and the artificial network, the at least one score representing at least one of information in the image frame indicative of a given feature affecting the digestive tract or a severity of the given feature; assigning said at least one score to said image frame; The method of claim 1,
2. said manipulation removing at least some color data from said at least one region of interest; or The method of claim 1 , further comprising at least one of: converting the at least one region of interest to grayscale.
3. The method of claim 1 , wherein the manipulation further comprises normalizing at least one image attribute of the at least one region of interest.
4. The method of claim 3 , wherein the at least one image attribute includes intensity.
5. The method of claim 1 , wherein the region of interest within the image frame does not exceed 25 percent of the total area of the image frame.
6. The method of claim 1 , wherein each region of interest in the image frame has the same aspect ratio as any other region of interest in the image frame.
7. The operation is identifying at least one additional region of interest within an additional image frame of the set of image frames, the additional image frame and the image frame each representing a homogenous image from the set of image frames; determining at least one additional score using the artificial network and the at least one additional region of interest; and The method of claim 1 , further comprising: assigning the at least one additional score to the additional image frame of the set of image frames.
8. The method of claim 1 , wherein the at least one score is determined by at least one decision tree.
9. The method of claim 8 , wherein the operations further comprise generating at least one regression based at least in part on the at least one decision tree.
10. The at least one regression Linear models, Cutoff, weighting curve, Gaussian function, Gaussian field, or Bayesian prediction.
11. at least one processor; and memory hardware in communication with the at least one processor, The memory hardware stores the following instructions: When performed by the at least one processor, the operations include causing the at least one processor to: receiving a set of image frames; at least a subset of the set of image frames depicting an interior surface of a digestive tract of a given patient; determining at least one region of interest within an image frame of the set of image frames; determining the at least one region of interest based on an edge value and a predetermined threshold; extracting the at least one region of interest from the image frames of the set of image frames; orienting the at least one region of interest based on the edge value; determining at least one score corresponding to the image frame using the oriented at least one region of interest and the artificial network, the at least one score representing at least one of information in the image frame indicative of a given feature affecting the digestive tract or a severity of the given feature; assigning said at least one score to said image frame; The system executes the above.
12. said manipulation removing at least some color data from said at least one region of interest; or The system of claim 11 , further comprising at least one of: converting the at least one region of interest to grayscale.
13. The system of claim 11 , wherein the manipulation further comprises normalizing at least one image attribute of the at least one region of interest.
14. The system of claim 13 , wherein the at least one image attribute includes intensity.
15. The system of claim 11 , wherein the region of interest within the image frame does not exceed 25 percent of the total area of the image frame.
16. The system of claim 11 , wherein each region of interest in the image frame has the same aspect ratio as any other region of interest in the image frame.
17. The operation is identifying at least one additional region of interest within an additional image frame of the set of image frames, the additional image frame and the image frame each representing a homogenous image from the set of image frames; determining at least one additional score using the artificial network and the at least one additional region of interest; and The system of claim 11 , further comprising: assigning the at least one additional score to the additional image frame of the set of image frames.
18. The system of claim 11 , wherein the at least one score is determined by at least one decision tree.
19. 20. The system of claim 18, wherein the operations further comprise generating at least one regression based at least in part on the at least one decision tree.
20. The at least one regression Linear models, Cutoff, weighting curve, Gaussian function, Gaussian field, or 20. The system of claim 19, comprising at least one of: a Bayesian prediction.
Citation Information
Patent Citations
Image processing apparatus, image processing method and image processing program
JP2006304995A
Image processing device and image processing method
WO2018105062A1