Systems and methods for automated detection of colon polyps using depth-in-color encoding and machine learning
The method of depth-in-color encoding and machine learning improves colon polyp detection in colonoscopy by analyzing OCT images, addressing subjective screening issues and reducing colorectal cancer incidence and mortality.
Patent Information
- Application Number
- PCT/US2025/030032
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-17
- Filing Date
- 2025-05-19
- Publication Date
- 2025-11-20
AI Technical Summary
Current colonoscopy screening methods for colorectal cancer are subjective and lead to disparate care, with adenoma-detection rates varying among physicians, resulting in missed polyps and increased incidence and mortality rates of colorectal cancer.
A method using depth-in-color encoding and machine learning to analyze optical coherence tomography (OCT) images, combining convolutional neural networks (CNNs) with support vector machines (SVMs) to characterize biological structures by encoding depth projections, enabling enhanced morphological detection of colon polyps.
Improves the accuracy of colon polyp detection, reducing missed polyps and enhancing objective screening, thereby decreasing colorectal cancer incidence and mortality.
Smart Images

Figure US2025030032_20112025_PF_FP_ABST
Abstract
Description
SYSTEMS AND METHODS FOR AUTOMATED DETECTION OF COLON POLYPS USING DEPTH-IN-COLOR ENCODING AND MACHINE LEARNINGCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Application No. 63 / 649,110, filed on May 17, 2024, which is incorporated herein by reference in its entirety for all purposes.BACKGROUND
[0002] Colorectal cancer (CRC) is the third most common cancer with prevalence rising among people under 50 years old. Colonoscopy is the most prevalent screening technique, demonstrated to clearly improve patient outcomes. Even though colonoscopy is highly effective, missed polyps can lead to interval cancers (5-10% screenings) that develop between surveillance intervals. Subjective screening leads to disparate care as, for instance, the adenoma-detection rate (ADR) can vary around 6% between physicians and every 1% increase in ADR correlates with a 3% reduced incidence rate of CRC and 5% reduction in mortality. Objective screening may improve outcomes in the colon as well as other tissues throughout the body.SUMMARY
[0003] Accordingly, improved methods for screening of tissues such as the colon are needed. Disclosed herein are systems, methods, and computer readable medium for characterizing a biological structure.
[0004] In one embodiment, a method of characterizing a biological structure including: obtaining image data for the biological structure using an imaging system; constructing a volumetric dataset of the biological structure based on the image data; determining at least one of a plurality of sections along a plane or curve within the volumetric dataset; and characterizing, using a trained algorithm and based on at least one of the plurality of sections, the biological structure for potential tissue type or disease state.
[0005] In another embodiment, a system for characterizing a biological structure including: a processor operably coupled to an imaging system, the processor being configured to: obtain image data for the biological structure using the imaging system; construct a volumetricdataset of the biological structure based on the image data; determine at least one of a plurality of sections along a plane or curve within the volumetric dataset; and characterize, using a trained algorithm and based on at least one of the plurality of sections, the biological structure for potential tissue type or disease state.BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Various objects, features, and advantages of the disclosed subject matter can be more fully appreciated with reference to the following detailed description of the disclosed subject matter when considered in connection with the following drawings, in which like reference numerals identify like elements.
[0007] FIG. 1 shows an illustration of colon polyp crypt cross sectional representations. Surface (0-150 pm), Mid (151-300 pm), Deep (301-451 pm). Typical hyperplastic serration is in the top third of crypt, typical sessile serrated adenoma serration depth is entire crypt.
[0008] FIGS. 2A-2B shows a schematic of an optical coherence tomography system in accordance with some embodiments of the disclosure. FIG. 2A: Center wavelength: 1310 nm, field-of-view: 1 cm x 1 cm, resolution: 27 / 14 pm (lateral / axial); PC: personal computer; SS: Swept source laser; DAQ: data acquisition; BD1-2: balanced detectors; BS1: 90 / 10 beam splitter; BS2: 50 / 50 beam splitter; PCT1-4: polarization controller; CIR102: circulator; VA: variable attenuator; VDL: variable delay line; SH: scanning head; SM: single mode fiber; PM: polarization maintaining fiber; MM: multimode fiber. FIG. 2B: schematic of an OCT system in communication with a computer / di splay system and a probe.
[0009] FIG. 3 shows an example ex vivo study process. Tissue is excised as part of standard-of-care, imaged, diagnosed, and annotated.
[0010] FIG. 4 shows a polyp annotation flow chart. Only samples which could be identified with high confidence were annotated by the pathologist.
[0011] FIG. 5 shows an example naive inception module showing how multiple convolutions give sensitivity to scale. The GoogLeNet architectures was chosen because it is based on these modules and polyps have spatial frequency differences that appear at multiple scales.
[0012] FIG. 6 shows an exemplary machine learning classification pipeline for ensemble decision level fusion using a GoogLeNet CNN and SVM. NMP: non-malignant potential.
[0013] FIGS. 7A-7D show the method of depth-in-color slice and sum encodings. FIG. 7A shows slice sections: red, surface (0-150 pm); green, mid (151-300 pm); blue, deep (301-450 pm). FIG. 7B shows depth-in-color shown the slice enface projection. FIG. 7C shows sum sections; red, surface (0-150 pm); green, mid (0-300 pm); blue, deep (0-450 pm). FIG. 7D shows depth-in-color shown the slice enface projection.
[0014] FIG. 8 shows patches from slice reconstructions of normal (uninvolved mucosa adjacent to polyp area), normal (part of a sample polyp with no diagnostic abnormality), hyperplastic, sessile serrated adenomas / polyps (SSA / P), and adenoma. Roman numerals in parenthesis after class are Kudos pit pattern criteria applied to the images. Conventional enface to a 0-450 pm (full section) sum. Layers color encoded as red (surface, slice 0-150 pm, sum 0- 150 pm), green (mid, slice 151-300 pm, sum 0-300 pm), and deep (blue, slice 301-450 pm, sum 0-450 pm). Contrast differences such as additional redness in Normal (I) and Adenoma (Ills) are an artifact of autoscaling.
[0015] FIGS. 9A-9B show performance metrics from classification of polyps of all sizes, excluding uninvolved mucosa. (FIG. 9A) receiver operating characteristic (ROC), (FIG. 9B) shows a confusion matrix (CM) for a test set of N=94 samples. Classification model trained on N=123 samples (all sizes). Two classes used: non-malignant potential (NMP) including Normal, and Hyperplastic and malignant potential (MP) including Adenoma, and SSAP. ROC shown for malignant potential. TPR - true positive rate, FPR - false positive rate, TNR - true negative rate, PPV - positive predictive value, NPV - negative predictive value, AUC - area under the curve. ROC - receiver operating characteristic, CM - confusion matrix.
[0016] FIGS. 10A-10B show performance metrics from classification of diminutive polyps (<5 mm) excluding uninvolved mucosa. (FIG. 10A) Receiver operating characteristic (ROC), and (FIG. 10B) confusion matrix (CM) using a test set N=78 of diminutive polyps. Classification model trained on N=123 samples (all sizes). Two classes used: non-malignant potential (NMP) including Normal, and Hyperplastic and malignant potential (MP) including Adenoma, and SSAP. TPR - true positive rate, FPR - false positive rate, TNR - true negativerate, PPV - positive predictive value, NPV - negative predictive value, AUC - area under the curve.
[0017] FIGS. 11A-11B show performance metrics (FIG. 11 A) Receiver operating characteristic (ROC), and (FIG. 1 IB) confusion matrix (CM) using a test set (N=106) of all samples including tissue that is endoscopically deemed suspicious but upon histophathology analysis is deemed to be normal, and tissue that is next to a polyp and is inherently sampled during extraction of the polyp. “Uninvolved normal mucosa” includes all tissue inadvertently samples near a polyp as well as tissue which was endoscopically deemed suspicious and extracted. Classification model trained on N=205 samples (all sizes). Two classes: non-malignant potential (NMP) including Normal, and Hyperplastic and malignant potential (MP) including Adenoma, and SSAP. TPR - true positive rate, FPR - false positive rate, TNR - true negative rate, PPV - positive predictive value, NPV - negative predictive value, AUC - area under the curve.
[0018] FIG. 12 shows an example process in accordance with some embodiments of the methods and systems described herein.
[0019] FIG. 13 shows an example computer system in accordance with some embodiments of the methods and systems described herein.DETAILED DESCRIPTION
[0020] In accordance with some embodiments of the disclosed subject matter, mechanisms (which can include, for example, systems, computer readable media, and methods) for characterizing a biological structure.
[0021] Colon polyps have structural differences that are depth dependent. This can be seen as a difference in crypt depths between hyperplastic and sessile serrated polyps (FIG. 1). This is also not unique to colon polyps. Cardiovascular tissue has a similar depth dependent nature where plaque morphology varies as a function of depth. This is also true in the esophagus, duodenum, lungs, skin, cornea, retina, nails, and this method would equally apply to those organs.
[0022] The emergence of machine learning approaches has introduced novel techniques for optical coherence tomography image analysis. Often depth dependent information isextracted on a B-Scan basis (2d: depth and 1 lateral coordinate). Enface (top down) views of tissue have relevant contextual information that often goes ignored when analyzing B-Scans. Hence, there is typically a tradeoff between analyzing enface projections, and their associated spatial frequency components, and maintaining adequate depth sensitivity.
[0023] Disclosed herein is a method to color encode depth projections to enable enhanced morphological detection capabilities. We combine this encoding with a neural network and a support vector machine. A general method is described below:
[0024] 1. OCT volumes are acquired by a system (colon polyp acquisition system: center wavelength 1310 bandwidth 100 nm, 100 kHz A-line rate, Axsun, 27 pm lateral resolution)
[0025] 2. OCT volumes are processed as linear, 2*log2, 10*logl0, or attenuation corrected using either attenuation measurement and intensity flattening, or ultrasound like attenuation correction. Each yields different levels of surface sensitivity.
[0026] 3. The surface and / or lumen is identified (the top interface of the tissue and air or tissue and water / saline / contrast medium) this is stored for later use.
[0027] 4. On a per B-Scan basis depth slices can be extracted. The slices are extracted by taking 3 uniform (or non-uniform) thickness slices (around 1 pm-2 mm). In the case of colon polyps, a thickness of 150 pm was used implying the following slices: Superficial (Red: 0-150 pm), Middle (Green: 151-300 pm), Deep (Blue: 301-450 pm).
[0028] 5. The slices are then summed for each B-Scan in the volume to make an enface projection. The most superficial slice is assigned a color channel (e.g. Red), the middle slice another channel (e.g. Green), and the deepest slice another channel (e.g. Blue).
[0029] 6. On a per B-Scan basis, depth sums can be extracted. The slices are extracted by taking 3 uniform (or non-uniform) thickness slices, going to the same depths but starting from the surface. In the case of colon polyps, a thickness difference of 150 p m was used implying the following sums: Superficial (Red: 0-150 pm), Middle (Green: 0- 300 pm), Deep (Blue: 0-450 pm).
[0030] 7. The enface slice and sums are used to train separate convolutional neural networks (CNNs) (GoogLeNet in the case of colon polyps) and finally, with the same trainingdata, used to train a support vector machine (if only slices or sums are used, a simple threshold is also sufficient).
[0031] 7a. When used to classify the CNN can output a probability score that acts as the input of the SVM.
[0032] 8. The SVM inputs can include other features such as those derived from the gray level co-occurrence I gray level tone difference matrices or Haralick texture parameters.
[0033] 9. Features can be included by ranking them using ChiA2 analysis / MRMR / ANOVA or any other statistical tools then excluding feature with the least predictive power. In the case of colon polyps ChiA2 was used and the convolutional neural networks were deemed to have adequate predictive power (>25%) to both be included in the support vector machine. This is not a strict criteria and can be set based on the threshold that maximizes the accuracy of the training and / or validation set.
[0034] Note: Enface slices are primarily included to extract depth sensitivity, while sums are primarily included to enhance contrast. Low SNR OCT can be caused by device variability, environmental conditions, or tissue factors such as highly scattering tissue. In a low SNR environment enface slice projections may have insufficient contrast for a convolutional neural network to extract relevant features. In addition, the slices may be insufficiently robust to artifacts such as fecal contents. Spectroscopic OCT can be used to reject those artifacts.
[0035] Depth-in-color encoding can be done independently as part of a visualization mechanism. Slices can be done either beam normal or surface normal.
[0036] Thus, in some embodiments the method includes one or more of: obtaining image data for the biological structure using an imaging system; constructing a volumetric representation of the biological structure based on the three-dimensional image data; determining at least one of a plurality of slice sections or a plurality of sum sections based on the volumetric representation of the biological structure; and / or characterizing, using a trained algorithm and based on at least one of the plurality of slice sections or the plurality of sum sections, the biological structure for potential tissue type or disease state.
[0037] Image data may be obtained using any suitable method that is able to image a sample at a specific depth and resolution to detect biological structures within the sample. Insome embodiments, the imaging system includes at least one of an optical coherence tomography (OCT) system, LC-OCT, FFOCT, FFOCM, pOCT, ultrasound system, computerized tomography (CT) system, ultrasound system, magnetic imaging technique (MRI) system, PET, SPECT, confocal microscopy system, SECM, or multiphoton microscopy. In some embodiments, an OCT system is used. FIG. 2A shows a schematic of an OCT system in accordance with some embodiments of the present disclosure.
[0038] In some embodiments, a sample is visualized ex vivo (e.g., a sample is extracted from a subject and then imaged). In some embodiments, a sample is visualized in vivo. When the sample is visualized in vivo, the imaging system (e.g., OCT or other imaging system) may be disposed in a probe or housing system that is capable of navigating an area of the subject (see FIG. 2B). For instance, an OCT system may be disposed in a probe such as a colonoscope and used to image and endothelial surface such as the surface of the colon tissue. In various embodiments, a subject (e.g., a human or other animal) may have the probe inserted into a particular body location or cavity, where the probe (e.g., a colonoscope) may include a light, camera, optics, and various instruments for excising tissue portions, along with an imaging probe (e.g., OCT imaging probe) for obtaining high-resolution image information. The probe may be used to identify structures or regions of a tissue that require further analysis and diagnosis. In some embodiments, the information from the probe and imaging system may be used to diagnose the subject based on characterizing the biological structure and / or recommend a treatment option based on the diagnosis. In certain embodiments, the treatment option may include excising the polyp or other tissue from the colon of a subject, where the tissue may be subject to further analysis upon removal.
[0039] It is widely recognized that tissue pathology is frequently depth-dependent; thus, evaluating morphology and related parameters as a function of depth is essential for accurate pathological diagnosis. Anotable example is the maturation pattern observed in epithelial tissues, where cells originate at the basal layer and progressively mature as they migrate toward the surface, eventually dying and being shed. This characteristic epithelial maturation pattern, consistently visible to varying degrees across epithelia, underscores the importance of depth- resolved automated analysis. Such an approach is particularly valuable in diagnosing epithelial diseases, including monitoring the progression from normal tissue through dysplasia and carcinoma in situ to invasive cancer. Additional pertinent examples include:
[0040] 1. The identification and classification of psoriasis, where characteristic depthdependent changes include hyperkeratosis and parakeratosis in superficial layers, and elongation of rete ridges in deeper layers.
[0041] 2. Diagnosis of lichen planus, characterized by basal cell degeneration and inflammatory infiltrate immediately beneath the epidermis.
[0042] 3. Assessment of dermatitis, where epidermal spongiosis (fluid accumulation within the epidermis) and depth-specific inflammatory infiltrates help differentiate between subtypes.
[0043] 4. Evaluation of oral leukoplakia, where depth-dependent architectural and cytological changes assist in distinguishing benign hyperplasia from dysplasia.
[0044] 5. Diagnosis of Barrett’s esophagus, characterized by depth-dependent glandular metaplasia replacing the normal squamous epithelium.
[0045] 6. Identification of duodenal pathology, including celiac disease, where villous atrophy and crypt hyperplasia are depth-dependent markers crucial for diagnosis.
[0046] 7. Diagnosis of colorectal polyps, where the depth of glandular dysplasia helps differentiate between benign and pre-cancerous lesions.
[0047] 8. Evaluation of anal intraepithelial neoplasia, which demonstrates characteristic depth-dependent cellular atypia and differentiation patterns.
[0048] 9. Detection of pre-cancerous conditions in the lung, such as bronchial dysplasia, characterized by depth-dependent alterations in the respiratory epithelium.
[0049] 10. Identification of laryngeal epithelial dysplasia, where the depth and extent of cellular atypia correlate with diagnostic severity.
[0050] 11. Gynecological diagnosis, such as cervical intraepithelial neoplasia (CIN), where depth-dependent grading of epithelial dysplasia is critical.
[0051] 12. Evaluation of urologic conditions such as urothelial carcinoma in situ of the bladder, characterized by depth-dependent cellular atypia within the urothelium.
[0052] Once a primary diagnosis is established, depth-dependent pathology analysis also becomes valuable in assessing the stage and extent of diseases, guiding treatment decisions and prognosis.
[0053] In various embodiments the disclosed systems, computer readable media, and methods are also applicable to other tissues and conditions (particularly those having depthdependent differences in normal and diseased tissues) including tissues from the heart, coronary arteries, peripheral vasculature, or any other catheter accessible by blood vessel, esophagus, duodenum, lungs, skin, cornea, retina, and nails, among others, as well as diseases and conditions of these tissues including but not limited to cancer, early disease states such as Barrett’s esophagus (both dysplastic and non-dysplastic), adenocarcinoma, wounds such as burns, skin cancers, and pre-cancers. In some embodiments, the systems and methods may also be applicable to brain tissue including tumor types in which nuclear clusters form well- delineated boundaries, or stomach tissue including gastric cardia.
[0054] In some embodiments, the image data includes a plurality of two-dimensional (2D) images. In some embodiments, each 2D image is collected at a specific vertical position (z- position). In other words, the imaging system may acquire a series of images along a vertical axis or z-axis. By doing so, a volumetric image may be constructed from the series of 2D images. The vertical spacing between images and the total number of images may be determined based on the tissue being imaged.
[0055] The intensity data from the volumetric image may be processed in various ways to reveal structural features that can then be analyzed, e.g., by a neural network. One type of processing produces slice sections, which may be determined based on a series of 2D images collected along the vertical axis. Slice sections may be generated by counting pixels from the top layer of a tissue to a target depth in tissue (see FIG. 7A). As shown in FIG. 7A, a first slice section layer can be identified within a first distance band from the surface (e.g., the surface (red) layer in FIG. 7A is made of the image intensities in the first 150 pm of tissue in the volumetric image); a second layer can be identified within a second distance band from the surface (e.g., the mid (green) layer in FIG. 7A is made of the image intensities of the tissue in the 151-300 pm range of the volumetric image); and a third layer can be identified within a third distance bandfrom the surface (e.g., the deep (blue) layer in FIG. 7A is made of the image intensities of the tissue in the 301-450 pm range of the volumetric image).
[0056] Sum sections may be determined based on a series of 2D images collected along the vertical axis. Sum sections may be generated by aggregating pixels from the top layer of a tissue to a variety of different target depths within the tissue. FIG. 7C shows example sum sections. Sum sections may be generated by counting pixels from the top layer of a tissue to a target depth within the tissue (see FIG. 7C) to produce bands of different thicknesses, all of which are measured from the surface of the sample. As shown in FIG. 7C, a first sum section layer can be identified within a first distance band from the surface (e.g., the surface (red) layer in FIG. 7A is made of the image intensities in the 0-150 pm portion of tissue in the volumetric image starting from the surface); a second layer can be identified within a second distance band from the surface (e.g., the mid (green) layer in FIG. 7C is made of the image intensities of the tissue in the 0-300 pm range of the volumetric image); and a third layer can be identified within a third distance band from the surface (e.g., the deep (blue) layer in FIG. 7C is made of the image intensities of the tissue in the 0-450 pm range of the volumetric image).
[0057] In some embodiments, the method includes dynamically assigning boundaries of the slice sections based on features of the biological structures. This may include determining layers which can be resolved using methods such as adaptive thresholding. This may further include determining the number of layers (e.g., number of images taken along the vertical axis) as well as the thicknesses of each layer (which do not have to be the same as one another).
[0058] In some embodiments, the slice sections or sum sections may be stored as separate channels of a single image frame. As noted above, this is comparable to an RGB image in which the image has three channels and a distinct value associated with each channel. Therefore, multi-channel images constructed from slice sections or sum sections are referred to as depth-in-color images. For clarity, these images are frequently shown in red, green, and blue. In some embodiments, depth-in-color images comprise more than three channels.
[0059] For instance, a slice section with a depth of 0-150 pm may be stored as one channel, a second slice section may have a depth of 151-300 pm, and a third slice section with a depth of 301-450 pm may be stored as a third channel. FIG. 7A shows an example slice section extraction, and FIG. 7B shows a multi-channel image constructed from slice sections. Regardingsum sections, a first sum section may have a depth of 0-150 pm, a second sum section may have a depth of 0-300 pm, and a third sum section may have a depth of 0-450 pm. FIG. 7C shows an example sum section extraction, and FIG. 7D shows an example multi-channel image constructed from the sum sections.
[0060] The multi-channel images are not affected by traditional limitations for multicolor images (e.g., RGB images). There may be two channels rather than three. Alternatively, there may be more than three channels.
[0061] In the particular embodiment shown in FIG. 7, the number of bands (slice or sum) was chosen to be 3 and each band was arbitrarily assigned the colors red, green, and blue in order to place the image data into an RGB format that could be used with pre-existing neural networks that have been trained on RGB image data. Nevertheless, the particular ranges of values for the bands can be varied and do not need to be the same thickness for a given sample. In addition, the procedures are not limited to using 3 bands and in various embodiments 2, 4, 5, or any other number of bands can be identified and used for further analysis. Furthermore, other methods of segmenting or dividing the volumetric information can also be applied to the data besides the slice and sum procedures discussed above.
[0062] The dimensions or thicknesses of the slice or sum sections (e.g., the upper and lower limits of the ranges of each slice or sum section) are not limited to those described above. The limits may be narrower or broader, for example based on the biological structure and / or image quality. In some embodiments, the limits of slice or sum ranges are manually defined. In some embodiments, ranges may be manually adapted based on pathology. For instance, if certain depths of a tissue are known to be important, ranges can be manually set to best show those depths.
[0063] The method may include dynamically assigning boundaries of the slice sections based on features of the biological structure. For instance, the different slice sections do not need to be equal to one another; one slice could have a depth of 50 pm while another could have a depth of 200 pm. In some embodiments, the number of sections may also be dynamically assigned.
[0064] In some embodiments, a trained algorithm may be used to characterize the biological structure using the multi-channel slice section images and multi-color sum sectionimages. A trained algorithm may be any supervised or unsupervised machine learning algorithm. In some embodiments, A trained algorithm may include, but is not limited to, regression algorithms, classification algorithms such as decision tress, discriminant analysis, support vector machines (SVMs), nearest-neighbors, naive Bayes, random forest, neural networks, convolutional neural networks, or autoencoders.
[0065] In some embodiments, the systems and methods described herein may be used to build hybrid trained networks which use patient data (demographics, previous patient visit information (e.g., patient chart), or a diagnosis obtained from another modality or a continuous variable which include disease state). The data can be plugged in raw and the network can learn how the aggregate of the inputs predict disease progression over time. Additionally or alternatively, the networks can be fused at the decision level where separate networks are trained to give a class, probability score and then those networks are fused with our network in an SVM.
[0066] In some embodiments, the trained algorithm is trained on multiple types of images. For instance, a trained algorithm may be trained based on slice sections and sum sections. Training the algorithm using two or more types of images provides improved accuracy compared to using a single type of data. Additionally or alternatively, the trained algorithm may be trained on single sections at specific depths (e.g., a section at 0 pm, a section at 100 pm, a section at 200 pm, etc.). In some embodiments, a probability density function moment such as mean, standard deviation, skew, kurtosis, etc. may be used instead of sum. In some embodiments, it is not necessary to take exactly a single depth. For example, if your surface of the tissue was tilted at 20 degrees, each section could be at 20 degrees to offset the surface tilt. Another important point is that often the tissue is not oriented perfectly so taking any of these sections at an angle makes sense to overcome this orientation issue. In a more sophisticated algorithm, any curved section reconstruction may be used. For example, epithelia are not flat and the basal layer (or basement membrane) may follow a curve in a 3D anatomical space. It is possible to follow the same 3D curve and create similar curves around which you would sum or average that are depth dependent.
[0067] In some embodiments, the trained algorithm is trained on images collected with different modalities or systems. The trained algorithm may be trained on images of polyps taken with white light endoscopy or narrow band imaging. Light has a fixed penetration depth whichmeans that the detected image is similar to a sum reconstruction. As such, transfer learning can be with white light endoscopy (low resolution and high resolution) to pre-train the networks. The specific modalities which concern depth that can be used do include both standard depth techniques such as point and line scan reflectance confocal, structured illumination. White light or hyperspectral images and spectral cubes can be used as differences in light penetration and knowledge about reflectance spectra can be used to construct similar slice and sum images that parameterize depth.
[0068] In some embodiments, the trained algorithm includes multiple trained algorithms. For instance, a trained algorithm may include a neural network as well as a classifier.
[0069] In some embodiments, the trained algorithm is used to classify the biological structures. For instance, a trained algorithm may be used to classify a biological structure as healthy or disease state, or used to indicate which tissue is likely to progress and which is not. Without being limited as to theory, analyzing the segmented volumetric image data may permit the neural network to identify structural features in the biological structure that are characteristic of particular conditions and then this information can then be used to characterize the biological structure. The systems and methods described herein may use spatial frequency differences (e.g., learns differences in scale of honeycomb structure) so while fractal patterns may be emergent and learned, it would still be related to spatial frequency differences. In addition to characterizing this particular structure, it could be used to diagnose the particular disease of the particular structure.
[0070] Examples
[0071] Example 1:
[0072] Colorectal cancer (CRC) is the third most common cancer with prevalence rising among people <50 years old. Colonoscopy is the most prevalent screening technique, demonstrated to clearly improve patient outcomes. Even though it is highly effective, missed polyps can lead to interval cancers (5-10% screenings), that develop between surveillance intervals. Subjective screening leads to disparate care, for instance, the adenoma-detection rate (ADR) can vary around 6% between physicians, and every 1% increase in ADR correlates with a3% reduced incidence rate of CRC and 5% reduction in mortality. Objective screening may improve outcomes.
[0073] New optical modalities are being developed which can screen more objectively, but effective translation requires benchmarking against field specific standards. The Preservation and Incorporation of Valuable Endoscopic Innovations (PIVI) standards apply to new in-vivo diagnostic modalities. Two thresholds were established for colon polyps: (1) if suspected rectosigmoid diminutive (<=5 mm) can be identified with a Negative Predictive Value (NPV)>=90% a see-and-leave protocol can be adopted, and (2) there is >=90% agreement with post-polypectomy surveillance intervals for all polyps (<=5 mm) then a resect-and-discard protocol can be adopted.
[0074] In CRC, tissue undergoes dysplastic changes then follows one of the two main architectural sequences. The most common, responsible for -70% of sporadic cancers, is the adenoma-carcinoma sequence, characterized by chromosomal instability (CIN). The remaining (-30% of sporadic cancers) are associated with the sessile serrated lesions that arise from hyperplastic lesions; genetically this is characterized by the CpG island methylator pathway (CIMP). Some tissue architecture is superficial and can be described using Kudo’s pit pattern classification, designed for White Light Endoscopy (WLE) and Narrow Band Imaging (NBI). Differences can be more subtle; hyperplastic polyps, are benign, but have surface features that are architecturally similar to SSA / P. Features that vary with depth fall within the imaging depth of an optical coherence tomography (OCT) system (1-2 mm).
[0075] OCT imaging probes exist which can image colonic mucosa. This includes a forward-viewing enface probe (lat. res. 5 pm), a side viewing balloon-based (lat. res. 20 pm), and capsule variants (lat. res. 27 pm / lat. res. 40 pm). Our group has also disclosed a capsule device which can image an entire swine colon.
[0076] Algorithms have been developed to classify polyps with traditional computer vision as well as deep learning. Zeng et al. used OCT to derive scattering coefficients (ps) in an ex-vivo study (lat. res. 10 pm) of 33 patients.
[0077] Different texture-derived features have been input into a support vector machine (SVM). They showed 95% sensitivity and 94% specificity for detecting cancer and adenomatous polyps. The same group later used an Angular Scattering Index (ASI), derived from the psmapsand showed that the ASI can discriminate polyps. Deep learning has also been used, with one group showing that lesions in a rat model (lat. res. 10 pm) can be classified using the Xception model using B-scan patches, reporting 98% specificity and 78% specificity. Another group has reported an ex -vivo OCT polyp study (lat. res. 5 pm) using multi-modal fusion deep learning. Scalar features input into a feed-forward model were fused with OCT derived spectral data input into a bi-directional Recurrent Neural Network (RNN). This was used to classify A-lines with a sensitivity of 98% and specificity of 95%, notably randomization was on a per A-line basis, not per-patient. It is not clear that animal models will generalize to humans with different imaging conditions. Sample numbers in human tissue studies are generally low (<15 polyps) and it is not clear that the models can adequately account for biological variability. The lateral resolution of all studies was <=10 pm, and it is unclear whether this will generalize to probes that have been shown to screen large portions of the organ. Though normal uninvolved mucosa is often included in data analysis, this is pathologically different than polypoid aggregates, and its inclusion may over represent the model’s performance.
[0078] Here, we present an ex-vivo OCT study comprising polyps and polyp fragments excised from 300 patients (lat. res. 27 pm). An analysis was performed on polyps (all sizes), and a sub-analysis of diminutive polyps (<5 mm). An ensemble network was used; two convolutional neural networks (CNNs) were fused at the decision-level into an SVM. Performance metrics were benchmarked against the PIVI criteria. To the best of our knowledge, this is the largest human polyp / polyp fragment study using OCT to date including SSA / P.
[0079] Methods
[0080] Imaging System Description
[0081] The imaging system is based on a 1310 nm swept source Axsun engine (FIG. 2). Light is guided from the source through a beam splitter (90 / 10, Thorlabs Inc. TW1300R2A1), on one path to the sample via circulators (Thorlabs CIR1310-APC). The sample reflectance is recombined with the reference arm in BS2 (50 / 50 Thorlabs Inc. TW1300R5A2). A variable reference arm attenuator (Thorlabs VOA50-APC) is used to adjust reference arm power. The VDL has a 7.5 mm travel range and can cover an optical path of 15 mm. Fiber-based polarization controllers (Thorlabs FPC020) are used to adjust power in each balanced detector. Three fiber types are used including single mode (SMF28), polarization maintaining (Panda 1310), and shortsegment multimode. The scanning head has a collimator (Thorlabs F220APC-1310), 2 galvanometer scanners (Thorlabs GVS002), and a 2.5 cm diameter scan lens (Thorlabs LSM03).
[0082] Study Using Excited Hyman Polyps / Polyp Fragments Using Excised Polyps of All Sizes
[0083] An ex-vivo study was conducted using polyps and polyp fragments excised from 300 patients (Massachusetts General Brigham Protocol #2015P000328). Following standard colonoscopic resection, the polyps were intercepted immediately post-excision, placed in a petri dish, and imaged using an Axsun-based OCT system. Samples were subsequently sent for standard-of-care histological processing. Independent pathological diagnoses were rendered by the pathology core. The study protocol was optimized part way through imaging to reduce sampling errors; polyp features are not macroscopically visible making it difficult to orient the polyps on the sample tray, consequently, a number of polyps were imaged upside down. To reduce this, polyps from patient numbers 91-300 were imaged on both sides by flipping them over after the initial imaging pass.
[0084] Performance metrics, including sensitivity, specificity, positive predictive value (PPV), and NPV, were calculated separately for the diminutive polyps. ROC analyses were conducted to benchmark the system’s performance, with the diminutive subset analyzed to ensure the system could distinguish between polyps with no malignant potential (NMP) including suspicious normal samples, hyperplastic polyps, and polyps with malignant potential (MP) including conventional adenomas, and sessile serrated adenomas and polyps.
[0085] Diminutive Polyp Sub-Analysis
[0086] The first PIVI criteria was designed to determine if diminutive (< 5 mm) polyps have malignant potential. To ensure that architectural changes associated with polyp size did not affect our results a specific sub-analysis was performed. Classification was a two-step process, involving training a classifier and subsequent classification of test set using two convolutional neural networks (CNNs) which output class probabilities, which were then used as inputs into an SVM. In the sub analysis, the existing CNN was used (trained on all polyp sizes). The SVM, however, was re-trained on only diminutive polyps. The test set only included diminutive polyps.
[0087] Annotation Protocol
[0088] Enface projections were annotated by an expert pathologist. Four classes were annotated: Normal, Hyperplastic, Adenoma, Sessile Serrate Adenomas I Polyps (SSAP). The pathologist received four maps for annotation, comprising three intensity-based en face projections: surface, mid, and deep and a single attenuation coefficient ( / it) map. High-confidence regions were annotated (QuPath) using the pathology core diagnosis, raw histology data, and Kudo's pit pattern criteria. Polyps that exhibited suboptimal orientation, tissue folding, or poor imaging quality (e.g., artifacts, out-of-focus) were excluded from the analysis if the pathologist could not annotate the images with high confidence (FIG. 4).
[0089] Pre-Processing
[0090] Annotation: A scan with tissue classified as normal by the pathologist was included in the dataset only if it represented the sole class present. This strategy was implemented to mitigate the influence of uninvolved normal mucosa on NPV and to more accurately assess the model’s performance concerning polypoid aggregates. When both the original and flipped polyp scans were labeled, the results were analyzed separately, followed by calculating an area-weighted average for each metric. This methodology effectively treats the combined scans as a single, large area sample. In instances where samples contained discontinuous regions of a single class, they were similarly aggregated into a single sample. Scans exhibiting multiple classes, with the exception of normal-labeled tissue, were analyzed as separate samples.
[0091] OCT Pre-Processing: OCT volumes were generated using standard swept-source processing methods, a Hanning window was used before the Fourier transform. To ensure that the OCT images were properly weighted for intensity and depth dependent differences were attributable to attenuation in a sample, we performed a confocal gate correction using a mirror translated away from focus using a z stage and an attenuator. Volumes used in processing were displayed in the square magnitude representation (201og10) to better capture depth dependent features. The polyp surface was automatically segmented using adaptive thresholding. Enface slice maps were generated by counting pixels from the top layer of the tissue to a target depth in tissue (assumed constant n=1.4). Three slices thickness were used: surface (0-150 pm), mid (151-300 pm), and deep (301-450 pm). Enface projections were individually autoscaled to fill the dynamic range (1% underflow, 1% saturation) and combined into each channel of an RGB reconstruction.The implementation of autoscaling enables the full utilization of the system's dynamic range; however, it may introduce image distortions. This issue is most pronounced in tissues with high attenuation, particularly in superficial (top) slices, where deeper layers become increasingly distorted. Practically, this distortion disrupts the visual appearance and shading of crypt structures, impairing their accurate representation. Such artifacts would not be present if attenuation was not a factor. The attenuating structures as well as the autoscaling process also remove the context clues necessary to resolve depth-aggregated features. It was critical to recover this and to do that sum images were also generated which contain the depth-aggregated features and then perform identical autoscaling. The sum and slice images, once generated, were then saved into memory for the classification step.
[0092] Classifier Design
[0093] Colon polyps, viewed from the top, exhibit differences in crypt size and shape depending on their class, viewed from the top down this appears as frequency differences. These differences have been shown to exist in multi-scale including having a fractal dimension. One option to extract the spatial frequency differences was using a series of fixed Gabor filters. This is similar to the inception module which instead learns convolutional kernels. The module is sensitive to scale as it combines 1x1, 3x3, and 5x5 convolutions within the same layer (FIG. 5). The GoogLeNet architecture is a network built on these modules. In brief, the GoogLeNet architecture uses global average pooling for parameter reduction instead of fully connected layers to preserve spatial information and mitigate over fitting, and has a depth of 22 layers making it ideal to learn rich, hierarchical feature representations. The network architecture was not modified, except the last fully connected layer was replaced to match the dimension of the classifier (Output Size = 2) and a pretrained implementation was used.
[0094] Classification Pipeline
[0095] The polyps were randomized on a per-patient basis 50% into a training and validation set (N = 123: Nor = 47, Hyp = 20, Adn = 44, SSA / P = 12) and 50% into a test set (N = 94: Nor = 16, Hyp = 21, Adn = 38, SSA / P = 19) the diminutive polyps were a subset of this same test set (N = 78, Nor = 16, Hyp = 19, Adn = 28, SSA / P = 15). Randomization was done on a per- patient basis to ensure that the model was adequately robust to biological variability. Data was augmented with random x- and y-axis reflection, rotation of 0-90 degrees, and translation. Scalewas not used for augmentation as pit size is indicative of class. Machine learning hyperparameter optimization was initially done using a validation set extracted from the training data (25% random patches). Once optimal hyperparameters were fixed and the entire training and validation dataset was used to train the model. The slice and sum images were converted into patches using a scanning square (1 x 1 mm, 100 x 100 pix, overlap: 80%). A patch was only counted as a certain class if 90% of pixels in the patch area had a single-class annotation. The training set was then used to train a convolutional neural network (GoogLeNet). The model was training for 30 epochs (Learning rate = 0.01). The resulting CNN predicted class probability was input as a score into a linear support vector machine (SVM). The SVM was trained using the training set (5-fold cross validation) with a linear kernel, automatic kernel scaling, box constraint level 1, and standardized data.
[0096] Data preprocessing, feature extraction, and machine learning model development were conducted using MATLAB R2023a (MathWorks, Natick, MA, USA). The implementation utilized the Neural Network Toolbox and the Statistics and Machine Learning Toolbox to facilitate algorithm development and evaluation.
[0097] Results
[0098] The increased biology-informed contrast provided depth-in-color is shown for different polyp types. The performance of the classifier is then shown on all polyps, and a subanalysis of diminutive polyps.
[0099] Increase Pit and Crypt Contrast Using Depth-In-Color OCT
[0100] The enface slice and sum reconstructions were generated using depth-in-color encoding and the resulting full-polyp images were generated. The projections show clear differences between the crypt patterns of different classes: Normal, Adenoma, Hyperplastic, and SSA / P. Normal tissue, as well as uninvolved normal mucosa, maintains the typically honeycomb crypt distribution and -150 pm size, while adenoma maintains long, spaghetti like predominantly blue (deep) crypts, hyperplastic shows differences in crypt patterns serration project as a blurringof the crypts when contrast to normal tissue, finally sessile serrated lesions can be visualized as enlarged crypts with blue rings around the crypt entrance.
[0101] Network Performance for Polyps of All Sizes
[0102] The receiver operating characteristic (ROC) analysis and confusion matrix (FIGS. 10A, 10B) highlight the effectiveness of the proposed classification method for colon polyps of all sizes. The performance metrics for the CNN based approach was an area under the curve (AUC) of 0.90, an accuracy of 86%, Sensitivity was 95% [85-100%], indicating a strong ability to correctly identify neoplastic polyps, while specificity was 73% [59-87%]. The positive predictive value (PPV) was 0.84 [0.75-0.94], and the negative predictive value (NPV) was 0.90 [0.82-0.98], suggesting the system reliably differentiates between malignant and benign polyps. The ROC operating point selected was the one that met PIVI.
[0103] Sub-Analysis of Diminutive Polyps and Benchmarking Against the PIVI Criteria
[0104] A sub-analysis of diminutive polyps (<5 mm) was conducted with the following performance metrics: an AUC was 0.88, with an accuracy of 86%, sensitivity of 93% [82-100%], and specificity of 77% [63-91%]. The PPV was 0.83 [0.72-0.95], and the NPV was 0.90 [0.81- 0.99], If these results are replicated in vivo, our system would meet the first PIVI threshold.
[0105] Sub-Analysis Including Normal Uninvolved Mucosa in Test Set
[0106] Previous studies report accuracy exceeding 90%. However, these studies are not directly comparable to ours. One key difference was the presence of normal uninvolved mucosa in the test set. This factor significantly influences classification accuracy. When normal tissue is included in the test set for our model, performance metrics align more closely with those that have been previously reported: AUC: 0.95, Accuracy: 90%, Sensitivity: 94% [83-100%], Specificity: 86% [76-97%], PPV: 0.85 [0.74-0.96], and NPV: 0.94 [0.87-1.00], In this sub-analysis there were more total samples, so the entire model (CNN and SVM) was retrained with a larger training set (66% training, 33% test). Although these metrics suggest improved performance, this introduces bias into the NPV calculation. Normal tissue exhibits distinct structural characteristics that CNNs can effectively classify, primarily by detecting spatial frequency differences. Patches of uninvolved mucosa often get accurately classified as “negative.” Suspicious tissue, which is later identified as normal pathologically, may have undergone structural changes. Thus, it is crucial toremove uninvolved mucosa from the test set, to avoid overstating performance. There are limits to this approach, as determining what the clinician considered suspicious during endoscopy is inherently complex. The exclusion of normal mucosa unless it is the sole class, likely improves performance, but it may not have eliminated bias.
[0107] Translation of Classification Approach to in vivo
[0108] The first PIVI threshold, which requires a negative predictive value (NPV) of 90% or greater, would be met if this approach holds in-vivo. The second PIVI (surveillance interval agreement of 90%) requires a larger study. The performance metrics of our system are comparable to WLE I NBI, each have a surveillance interval agreement of 95%. It is likely the approach will perform better in-vivo. In-vivo and ex-vivo imaging each have distinct advantages and challenges. In the case of in-vivo imaging, the organ is intact but there are also fecal contents, although these can be reduced with bowel preparation. Other artifacts may occur depending on probe type. In the case of capsule-like probes these are motion artifacts, limited focus and OCT range control. Motion artifacts have been reported in a capsule with high resolution driveshaft scanning to be visible in -20% of frames when using extremely high magnification. Tissue folding is a problem was observed in capsule probes and will likely render parts of the image un-analyzable. Defocus will also need to be compensated and can be improved with either extended depth of focus or numerical refocusing. The imaging range exceeding the bowel wall can be adjusted in real time using a tracker with a dynamically adjustable variable delay line. Forward viewing probes which can capture enface OCT projections can be used with localization performed via a conventional endoscope. These probes will not be subject to tissue folding, motion artifacts, or OCT range control. They would, however, suffer from sampling error.
[0109] Conclusion
[0110] This study demonstrates that an OCT system, with in-vivo resolution, combined with machine learning algorithms can effectively classify ex-vivo colorectal polyps, achieving performance metrics that meet the AGSE’s PIVI criteria. Specifically, the system achieved high accuracy, sensitivity, and specificity in distinguishing polyps with malignant potential, including diminutive hyperplastic polyps, which could potentially be left in place without resection. The ability to meet these clinical benchmarks suggests that OCT, when integrated with automatedanalysis, holds a significant promise as a diagnostic tool for real-time classification of colorectal lesions.
[0111] Further Details
[0112] The polyps were primarily excised from the rectosigmoid. Error! Reference source not found, shows the distribution of polyps by location in the training and test sets.
[0113] Table 1 : Location from where polyps were extracted from and how they were distributed in training and test sets.
[0114] Surface Map Generation Using a Custom Adaptive Thresholding Routine
[0115] To generate the surface map of the tissue, we utilized a custom adaptive thresholding algorithm designed to segment the enface projections from 3D OCT volumes. The process involved several key steps to ensure accurate surface delineation:
[0116] 1. Data Preprocessing: The raw OCT volume data was loaded from the respective directory and preprocessed. The enface projection for each B-scan was generated by summing the pixel intensities across the depth dimension. This provided a 2D representation of the tissue surface in each frame. A pre-existing tissue mask was applied to isolate tissue regions and exclude noise.
[0117] 2 Low Pass Filtering and Identifying a High-Confidence Region: For each B-scan, the pixel intensities were processed to remove low-signal noise through rectification and intensity normalization. A Gaussian filter (o = 5) was applied to smooth the image and eliminate high- frequency noise. A pixel-wise threshold was then calculated based on the mean intensity of non-zero values, retaining only the pixels with intensities above 25% of the mean to highlight tissue structures. The binary mask was then used as a high confidence region.
[0118] 3. Surface Detection: For each frame, the location of the tissue surface was then identified by detecting maximum intensity in each column of the image (inside the high confidence region). A moving average (window = 5) was applied to smooth the detected surface. To account for irregularities, an adaptive reset check was introduced. A “for-loop” compared the detected surface with the smoothed mean. If the deviation exceeded 25%, the detected surface was reset at the values where that was exceeded, this led to interpolation through multiple iterations of smoothing.
[0119] 4 Fine-tuning Surface Segmentation: The surface was further refined by iteratively correcting outliers using the moving mean approach. This adaptive adjustment ensured that the surface map accurately followed the contour of the tissue, even in the presence of noise or artifacts.
[0120] 5. Interpolation and Fin-Tuning: To handle dropped frames and missing data, interpolation was applied between frames where the surface could not be detected. A low-pass Gaussian filter (o = 1 and 3) was used to smooth the final surface map. Regions that deviated significantly from the smoothed surface were adaptively corrected based on the difference between the raw surface and the smoothed data.
[0121] References for Example 1
[0122] 1. Szegedy, C., et al. Going deeper with convolutions, in Proceedings of theIEEE conference on computer vision and pattern recognition. 2015.
[0123] 2. Crockett, S.D., et al., Sessile Serrated Adenomas: An Evidence-BasedGuide to Management. Clinical Gastroenterology and Hepatology, 2015. 13(1): p. 11-26. el.
[0124] 3. Kudo, S.-e., et al., Diagnosis of colorectal tumorous lesions by magnifying endoscopy. Gastrointestinal Endoscopy, 1996. 44(1): p. 8-14.
[0125] 4. Weinberg, B.A. and J.L. Marshall, Colon Cancer in Young Adults: Trends and Their Implications. Curr Oncol Rep, 2019. 21(1): p. 3.
[0126] 5. Winawer, S. J , Colorectal cancer screening. Best practice & researchClinical gastroenterology, 2007. 21(6): p. 1031-1048.
[0127] 6. Sanduleanu, S., A.M. Masclee, and G.A. Meijer, Interval cancers after colonoscopy — insights and recommendations. Nature reviews Gastroenterology & hepatology, 2012. 9(9): p. 550-554.
[0128] 7. le Clercq, C., et al., Interval Colorectal Cancers Frequently Have SubtleMacroscopic Appearance: A 10 Year-Experience in an Academic Center. Gastroenterology, 2011. 140(5): p. S-112-S-113.
[0129] 8. Hassan, C., et al., Variability in adenoma detection rate in control groups of randomized colonoscopy trials: a systematic review and meta-analysis. Gastrointestinal endoscopy, 2023. 97(2): p. 212-225. e7.
[0130] 9. Pohl, H. and D. J. Robertson, Colorectal Cancers Detected AfterColonoscopy Frequently Result From Missed Lesions. Clinical Gastroenterology and Hepatology, 2010. 8(10): p. 858-864.
[0131] 10. Rex, D.K., et al., Quality indicators for colonoscopy. GastrointestinalEndoscopy, 2015. 81(1): p. 31-53.
[0132] 11. Fearon, E.R., Molecular Genetics of Colorectal Cancer. Annual Review ofPathology: Mechanisms of Disease, 2011. 6(1): p. 479-507.
[0133] 12. Pino, M.S. and D.C. Chung, The Chromosomal Instability Pathway inColon Cancer. Gastroenterology, 2010. 138(6): p. 2059-2072.
[0134] 13. Bettington, M., et al., The serrated pathway to colorectal carcinoma: current concepts and challenges. Histopathology, 2013. 62(3): p. 367-386.
[0135] 14. Amaro, A., S. Chiara, and U. Pfeffer, Molecular evolution of colorectal cancer: from multistep carcinogenesis to the big bang. Cancer and Metastasis Reviews, 2016. 35(1): p. 63-74.
[0136] 15. Liang, K., et al., Endoscopic forward-viewing optical coherence tomography and angiography with MHz swept source. Optics Letters, 2017. 42(16): p. 3193- 3196.
[0137] 16. Adler, D.C., et al., Three-dimensional endomicroscopy of the human colon using optical coherence tomography. Opt Express, 2009. 17(2): p. 784-96.
[0138] 17. Liang, K., et al., Ultrahigh speed en face OCT capsule for endoscopic imaging. Biomedical optics express, 2015. 6(4): p. 1146-1163.
[0139] 18. Song, D.-R., et al. Safety study of tethered capsule endomicroscopy (TCE) pull back through long segments of the small intestine, in Endoscopic Microscopy XVII. 2022. SPIE.
[0140] 19. Song, D.-R,, et al. Self-propelled retrograde tethered capsule endomicroscopy (R-TCE) in the colon, in Optical Coherence Tomography and Coherence Domain Optical Methods in Biomedicine XXVII. 2023. SPIE.
[0141] 20. Zeng, Y, et al., Diagnosing colorectal abnormalities using scattering coefficient maps acquired from optical coherence tomography. J Biophotonics, 2021. 14(1): p. e202000276.
[0142] 21. Zeng, Y, et al., The Angular Spectrum of the Scattering Coefficient MapReveals Subsurface Colorectal Cancer. Sci Rep, 2019. 9(1): p. 2998.
[0143] 22. Saratxaga, C.L., et al., Characterization of Optical Coherence TomographyImages for Colon Lesion Differentiation under Deep Learning. Applied Sciences, 2021. 11(7): p. 3119-.
[0144] 23. Kendall, W., et al., Deep learning classification of ex vivo human colon tissues using spectroscopic OCT. bioRxiv, 2023: p. 2023.09. 04.555974.
[0145] 24. Drexler, W., Optical Coherence Tomography. 2015.
[0146] 25. Faber, D.J., et al., Quantitative measurement of attenuation coefficients of weakly scattering media using optical coherence tomography. Optics Express, 2004. 12(19): p. 4353-4365.
[0147] 26. Zeng, Y, et al., The angular spectrum of the scattering coefficient map reveals subsurface colorectal Cancer. Scientific reports, 2019. 9(1): p. 2998.
[0148] 27. Costas, P., T. Andrew, and J.T.M.D. Guillermo. Morphological segmentation and fractal analysis for the classification of colon polyps from en face optical coherence tomography (OCT) images, in Proc. SPIE. 2023.
[0149] 28. Serre, T, Robust Object Recognition with Cortex-Like Mechanisms. IEEETrans. Pattern Anal. Mach. Intell, 2007. 29(3): p. 411-426.
[0150] 29. Luo, H., et al., Human colorectal cancer tissue assessment using optical coherence tomography catheter and deep learning. J Biophotonics, 2022. 15(6): p. e202100349.
[0151] 30. Haolin, N., et al. In vivo colorectal polyp evaluation using an optical coherence tomography catheter and deep learning: results of a feasibility study, in Proc.SPIE. 2024.
[0152] 31. Wallace, M.B.M.D., et al., Accuracy of in vivo colorectal polyp discrimination by using dual-focus high-definition narrow-band imaging colonoscopy. Gastrointestinal Endoscopy, 2014. 80(6): p. 1072-1087.
[0153] 32. Yin, B., et al., Extended depth of focus for coherence-based cellular imaging. Optica, 2017. 4(8): p. 959-965.
[0154] 33. Ralston, T.S., et al., Inverse scattering for optical coherence tomography.Journal of the Optical Society of America A, 2006, 23(5): p. 1027-1037.
[0155] 34. Tearney, G.J., B.E. Bouma, and LG. Fujimoto, High-speed phase- and group-delay scanning with a grating-based phase control delay line. Optics Letters, 1997. 22(23): p. 1811-1813.
[0156] Example 2: Illustrative Embodiments of Methods and Systems Described Herein
[0157] FIG. 12 shows a flow chart of an exemplary method 1200 of characterizing a biological structure. The method includes obtaining image data for a biological structure at step 1202. At step 1204, a volumetries representation of the biological structure may be constructed. The volumetric representation is based on the obtained image data. At step 1206, the method may include determining at least one of a plurality of slice sections or a plurality of sum sections. The plurality of slice sections and sum sections may be based on the volumetric representation. At step 1208, the method includes characterizing the biological structure using a trained algorithm based on at least one of the plurality of slice sections or the plurality of sum sections.
[0158] In FIG. 13, an example 1300 of a system (e.g., a data processing system) for 1300 in accordance with some embodiments of the disclosed subject matter is shown.
[0159] In some embodiments, computing device 1304 and / or server 1316 can be any suitable computing device or combination of devices, such as a desktop computer, a laptop computer, a smartphone, a tablet computer, a wearable computer, a server computer, a virtual machine being executed by a physical computing device, etc. As described herein, system 1300 can present information about a characterized biological structure to a user.
[0160] In some embodiments, communication network 1302 can be any suitable communication network or combination of communication networks. In some embodiments, communication network 1302 can be any suitable communication network or combination of communication networks. For example, communication network 1302 can include a Wi-Fi network (which can include one or more wireless routers, one or more switches, etc.), a peer-to- peer network (e.g., a Bluetooth network), a cellular network (e.g., a 4G network, a 5G network, etc., complying with any suitable standard, such as CDMA, GSM, LTE, LTE Advanced, WiMAX, etc.), a wired network, etc. In some embodiments, communication network 1002 can be a local area network, a wide area network, a public network (e.g., the Internet), a private or semi-private network (e.g., a corporate or university intranet), any other suitable type of network, or any suitable combination of networks. Communications links shown in FIG. 13 can each be any suitable communications link or combination of communications links, such as wired links, fiber optic links, Wi-Fi links, Bluetooth links, cellular links, etc.
[0161] FIG. 13 additionally shows an example of hardware that can be used to implement computing device 1304 and server 1316 in accordance with some embodiments of the disclosed subject matter. In some embodiments, computing device 1304 can be used to execute one or more set of instructions to identify a behavioral catalog. In other embodiments, computing device 1304 can be used to identify therapeutic interventions. In still other embodiments, computing device 1304 can be used to identify a configuration of parameter of a gene regulatory network to perform a desired function.
[0162] As shown in FIG. 13, computing device 1304 can include one or more hardware processor 1306, one or more displays 1308, one or more inputs 1310, one or more communications 1312, and / or memory 1314. In some embodiments, processor 1306 can be any suitable hardware processor or combination of processors, such as central processing unit, a graphics processing unit, etc. In some embodiments, display 1308 can include any suitabledisplay devices, such as a computer monitor, a touchscreen, a television, etc. In some embodiments, inputs 1310 can include any suitable input device and / or sensors that can be used to receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, etc.
[0163] In some embodiments, communication systems 1312 can include any suitable hardware, firmware, and / or software for communicating information over communication network 1302 and / or any other suitable communication networks. For example, communications systems 1312 can include one or more transceivers, one or more communication chips and / or chip sets, etc. In a more particular example, communications systems 1312 can include hardware, firmware and / or software that can be used to establish a Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, etc.
[0164] In some embodiments, memory 1314 can include any suitable storage device or devices that can be used to store instructions, values, etc., that can be used, for example, by processor 1306 to present content using display 1308, to communicate with server 1316 via communications system(s) 1312, etc.
[0165] Memory 1314 can include any suitable volatile memory, non-volatile memory, storage, or any suitable combination thereof. For example, memory 1314 can include RAM, ROM, EEPROM, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, etc. In some embodiments, memory 1314 can have encoded thereon a computer program for controlling operation of computing device 1304. In such embodiments, processor 1306 can execute at least a portion of the computer program to present content (e.g., images, user interfaces, graphics, tables, etc ), receive content from server 1316, transmit information to server 1316, etc.
[0166] In some embodiments, server 1316 can include a processor 1318, a display 1320, one or more inputs 1322, one or more communications systems 1324, and / or memory 1326. In some embodiments, processor 1318 can be any suitable hardware processor or combination of processors, such as a central processing unit, a graphics processing unit, etc. In some embodiments, display 1320 can include any suitable display devices, such as a computer monitor, a touchscreen, a television, etc. In some embodiments, inputs 1322 can include any suitable input devices and / or sensors that can be used to receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, etc.
[0167] In some embodiments, communications systems 1324 can include any suitable hardware, firmware, and / or software for communicating information over communication network 1302 and / or any other suitable communication networks. For example, communications systems 1324 can include one or more transceivers, one or more communication chips and / or chip sets, etc. In a more particular example, communications systems 1324 can include hardware, firmware and / or software that can be used to establish a Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, etc.
[0168] In some embodiments, memory 1326 can include any suitable storage device or devices that can be used to store instructions, values, etc., that can be used, for example, by processor 1318 to present content using display 1320, to communicate with one or more computing devices 1304, etc. Memory 1326 can include any suitable volatile memory, nonvolatile memory, storage, or any suitable combination thereof. For example, memory 1326 can include RAM, ROM, EEPROM, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, etc. In some embodiments, memory 1326 can have encoded thereon a server program for controlling operation of server 1316. In such embodiments, processor 1318 can execute at least a portion of the server program to transmit information and / or content (e.g., results of a tissue identification and / or classification, a user interface, etc.) to one or more computing devices 1304, receive information and / or content from one or more computing devices 1304, receive instructions from one or more devices (e.g., a personal computer, a laptop computer, a tablet computer, a smartphone, etc.), etc.
[0169] In some embodiments, any suitable computer readable media can be used for storing instructions for performing the functions and / or processes described herein. For example, in some embodiments, computer readable media can be transitory or non-transitory. For example, non-transitory computer readable media can include media such as magnetic media (such as hard disks, floppy disks, etc.), optical media (such as compact discs, digital video discs, Blu-ray discs, etc ), semiconductor media (such as RAM, Flash memory, electrically programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM), etc.), any suitable media that is not fleeting or devoid of any semblance of permanence during transmission, and / or any suitable tangible media. As another example, transitory computer readable media can include signals on networks, in wires, conductors,optical fibers, circuits, or any suitable media that is fleeting and devoid of any semblance of permanence during transmission, and / or any suitable intangible media.
[0170] A number of references to patent and non-patent documents are made throughout the publication, each of which is herein incorporated by reference in its entirety.
[0171] While the invention has been described above in connection with particular embodiments and examples, the invention is not necessarily so limited, and that numerous other embodiments, examples, uses, modifications and departures from the embodiments, examples and uses are intended to be encompassed by the claims attached hereto.
Claims
CLAIMSWhat is claimed is:
1. A method of characterizing a biological structure comprising: obtaining image data for the biological structure using an imaging system; constructing a volumetric dataset representing the biological structure based on the image data; determining a plurality of sections along a plane or curve within the volumetric dataset; and characterizing, using a trained algorithm and based on at least one of the plurality of sections, the biological structure for potential tissue type or disease state.
2. The method of claim 1, wherein at least one of the plurality of sections is a slice section or a sum section.
3. The method of claim 1 or 2, wherein at least one of the data along the at least one plane or curve and adjacent to the at least one plane or curve is further analyzed to create a section based on a moment of the probability distribution function of the data.
4. The method of any one of the preceding claims, wherein the image data comprises a plurality of two-dimensional images.
5. The method of any one of the preceding claims, wherein each image in the plurality of 2D images comprises a 2D image acquired at a different z-position.
6. The method of any one of claims 2-5, wherein determining the plurality of slice sections further comprises:generating the plurality of slice sections from the volumetric representation of the biological structure wherein each of the plurality of slice sections comprises a different layer of the volumetric representation of the biological structure.
7. The method of claim 6, wherein each of the plurality of slice sections is stored as a separate channel of a single image frame.
8. The method of claim 6 or 7, wherein the method further comprises dynamically assigning boundaries of the slice sections based on features of the biological structure.
9. The method of any one of claims 2-8, wherein determining the plurality of sum sections further comprises: generating the plurality of sum sections from the volumetric representation of the biological structure wherein each of the plurality of sum sections comprises a different thickness of the volumetric representation of the biological structure.
10. The method of claim 9, wherein each of the plurality of sum sections is stored as a separate channel of a single image frame.
11. The method of claim 9 or 10, wherein the method further includes dynamically assigning boundaries of the sum sections based on features of the biological structure.
12. The method of any one of the preceding claims, wherein the trained algorithm comprises at least one of a supervised or unsupervised machine learning algorithm, and wherein characterizing the biological structure further comprises:characterizing the biological structure based on analyzing at least one of the plurality of slice sections or the plurality of sum sections to characterize the biological structure.
13. The method of any one of the preceding claims, wherein the trained algorithm comprises a convolutional neural network (CNN).
14. The method of any one of the preceding claims, wherein the trained algorithm further comprises a trained classifier, wherein the trained classifier comprises at least one of a decision tree, discriminant analysis, Bayes classifier, ensemble classifier, or nearest neighbor classifier.
15. The method of claim 14, wherein the trained classifier is a linear support vector machine.
16. The method of any one of the preceding claims, wherein the method further comprises: diagnosing the subject based on characterizing the biological structure, and recommending a treatment option based on the diagnosis.
17. The method of claim 16, wherein the treatment option comprises excising the polyp from the colon of a subject.
18. The method of any one of the preceding claims, wherein the disease state is at least one of dysplasia or malignancy.
19. The method of any one of the preceding claims, wherein the tissue types is at least one of normal tissue, lesions with unknown malignant potential but other diagnostic abnormality, polyp, inflamed tissue, lesion with malignant potential, or lesion without malignant potential.
20. The method of any one of the preceding claims, wherein characterizing the biological structure for potential malignancy comprises classifying the biological structure as at least one of malignant, potentially malignant, or benign.
21. The method of any one of the preceding claims, wherein the imaging system comprises at least one of an optical coherence tomography system, ultrasound system, computerized tomography system, MRI system, or confocal microscopy system.
22. The method of claim 21, wherein the imaging system comprises an optical coherence tomography system.
23. A system for characterizing a biological structure comprising: a processor operably coupled to an imaging system, the processor being configured to: obtain image data for the biological structure using the imaging system; construct a volumetric data set of the biological structure based on the image data; determine at least one of a plurality of sections along a plane or curve within the volumetric dataset; and characterize, using a trained algorithm and based on at least one of the plurality of sections, the biological structure for potential tissue type or disease state.
24. The system of claim 23, wherein at least one of the plurality of sections is a slice section or a sum section.
25. The system of claim 23 or 24, wherein at least one of the data along the at least one plane or curve and adjacent to the at least one plane or curve is further analyzed to create a section based on a moment of the probability distribution function of the data.
26. The system of claim 25, wherein the image data comprises a plurality of two-dimensional images.
27. The system of claim 25 or 26, wherein each image in the plurality of 2D images comprises a 2D image acquired at a different z-position.
28. The system of any one of claims 24-27, wherein the processor, when determining the plurality of slice sections, is further configured to: generate the plurality of slice sections from the volumetric representation of the biological structure wherein each of the plurality of slice sections comprises a different layer of the volumetric representation of the biological structure.
29. The system of claim 28, wherein each of the plurality of slice sections is stored as a separate channel of a single image frame.
30. The system of claim 28 or 29, wherein the processor is further configured to dynamically assign boundaries of the slice sections based on features of the biological structure.
31. The system of any one of claims 24-30, wherein the processor, when determining the plurality of sum sections, is further configured to: generate the plurality of sum sections from the volumetric representation of the biological structure wherein each of the plurality of sum sections comprises a different thickness of the volumetric representation of the biological structure.
32. The system of claim 31, wherein each of the plurality of sum sections is stored as a separate channel of a single image frame.
33. The system of claim 31 or 32, wherein the processor is further configured to dynamically assign boundaries of the sum sections based on features of the biological structure.
34. The system of any one of claims 23-33, wherein the trained algorithm comprises at least one of a supervised or unsupervised machine learning algorithm, andwherein the processor, when characterizing the biological structure, is further configured to: characterize the biological structure based on analyzing at least one of the plurality of slice sections or the plurality of sum sections to characterize the biological structure.
35. The system of any one of claims 23-34, wherein the trained neural network comprises a convolutional neural network (CNN).
36. The system of any one of claims 23-35, wherein the trained algorithm further comprises a trained classifier, wherein the trained classifier comprises at least one of a decision tree, discriminant analysis, Bayes classifier, ensemble classifier, or nearest neighbor classifier.
37. The system of claim 36, wherein the trained classifier is a linear support vector machine.
38. The system of any one of claims 23-37, wherein the processor is further configured to: diagnose the subject based on characterizing the biological structure, and recommend a treatment option based on the diagnosis.
39. The system of claim 38, wherein the treatment option comprises excising the polyp from the colon of a subject.
40. The system of any one of claims 23-39, wherein the disease state is at least one of dysplasia or malignancy.
41. The system of any one of claims 23-40, wherein the tissue types is at least one of normal tissue, polyp, inflamed tissue, lesion with malignant potential, or lesion without malignant potential.
42. The system of any one of claims 23-41, wherein the processor, when characterizing the biological structure for potential malignancy, is further configured to classify the biological structure as at least one of malignant, potentially malignant, or benign.
43. The system of any one of claims 23-42, wherein the imaging system comprises at least one of an optical coherence tomography system, ultrasound system, computerized tomography system, MRI system, or confocal microscopy system.
44. The system of claim 43, wherein the imaging system comprises an optical coherence tomography system.
Citation Information
Patent Citations
Method for detecting polyps in a three dimensional image volume
US20060079761A1
Apparatus, device and method for capsule microscopy
US20190261840A1
Processing image data for assessing a clinical question
WO2023083710A1