Machine learning systems for wound assessment, healing prediction, and treatment
A non-invasive optical imaging system with machine learning algorithms predicts wound healing potential, addressing the imprecision and delay of current methods by enabling immediate, accurate therapy selection for diabetic foot ulcers and other wounds.
Patent Information
- Application Number
- JP2024180995
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-03-14
- Filing Date
- 2024-10-16
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2039-12-11
AI Technical Summary
Current methods for assessing and predicting wound healing, particularly diabetic foot ulcers, are invasive, imprecise, and delayed, leading to prolonged treatment times and increased risk of amputation.
A non-invasive, non-contact optical imaging system using machine learning algorithms to analyze tissue regions at or near a wound, providing quantitative healing predictions based on reflectance intensity values and patient health measures, enabling immediate determination of appropriate wound care therapy.
Enables accurate, real-time prediction of wound healing potential, allowing for timely administration of standard or advanced therapies, reducing healing time and amputation risk.
Smart Images

Figure 0007793018000018 
Figure 0007793018000019 
Figure 0007793018000020
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Application No. 62 / 780,854, filed December 17, 2018, entitled "Predicting Healing of Diabetic Foot Ulcers at the First Visit Using Artificial Intelligence," U.S. Provisional Application No. 62 / 780,121, filed December 14, 2018, entitled "System and Method for High Precision Multi-Aperture Spectral Imaging," and U.S. Provisional Application No. 62 / 818,375, filed March 14, 2019, entitled "System and Method for High Precision Multi-Aperture Spectral Imaging," each of which is expressly incorporated herein by reference in its entirety for all purposes.
[0002] STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT Portions of the inventions described in this disclosure were made with U.S. government support awarded by the Biomedical Advanced Research and Development Authority (BARDA) within the Office of the Assistant Secretary for Preparedness and Response, U.S. Department of Health and Human Services, under Contract No. HHSO100201300022C. Portions of the inventions described in this disclosure were made with U.S. government support awarded by the Defense Health Agency (DHA) under Contract No. W81XWH-17-C-0170 and / or Contract No. W81XWH-18-C-0114. The U.S. government may have certain rights in this invention.
[0003] The systems and methods disclosed herein relate to medical imaging, and more particularly to wound assessment, healing prediction, and treatment using machine learning techniques. [Background technology]
[0004] Optical imaging is an emerging technology with the potential to improve disease prevention, diagnosis, and treatment in the clinic, hospital bed, or operating room during emergency situations. Optical imaging techniques can noninvasively distinguish between tissues, or between tissues labeled with endogenous or exogenous contrast agents, by measuring the photon absorption or scattering properties of each tissue at various wavelengths. These differences in photon absorption and scattering properties provide specific tissue contrasts, enabling studies of functional and molecular activity that may be indicative of health and disease.
[0005] The electromagnetic spectrum is a band of wavelengths or frequencies of electromagnetic radiation (e.g., light). From longest to shortest wavelength, the electromagnetic spectrum includes radio waves, microwaves, infrared light (IR), visible light (i.e., light detectable by structures in the human eye), ultraviolet light (UV), X-rays, and gamma rays. Spectral imaging also refers to a branch of spectroscopy and photography that collects partial or complete spectra of spectral information at various locations on an image plane. Some types of spectral imaging systems can capture one or more spectral bands. Multispectral imaging systems can capture multiple spectral bands (typically a dozen or fewer spectral bands in different spectral regions) and collect measurements in each spectral band at each pixel, referencing a bandwidth of tens of nanometers per spectral channel. Hyperspectral imaging systems, on the other hand, can measure many more spectral bands, e.g., 200 or more spectral bands, and some types of hyperspectral imaging systems can continuously sample narrow bands (e.g., spectral bandwidths of a few nanometers or less) along a portion of the electromagnetic spectrum. Summary of the Invention [Means for solving the problem]
[0006] Aspects of the technology described herein relate to devices and methods that can be used to evaluate and / or classify tissue regions at or near a wound using non-contact, non-invasive, non-radioactive optical imaging. Such devices and methods may, for example, identify tissue regions corresponding to various tissue health classifications associated with the wound and / or measure healing-predictive parameters for the wound or portions thereof, and output a visual display of the identified regions and / or parameters that can be used by a clinician to determine a prognosis for wound healing, select an appropriate wound care therapy, or both. In some embodiments, devices and methods consistent with the technology of the present invention can provide such classification and / or prediction based on imaging at a single wavelength or multiple wavelengths. There has long been a need for such non-invasive imaging technology that can provide physicians with quantitative information that can predict the healing of a wound or portions thereof.
[0007] In one embodiment, a system for assessing or predicting wound healing comprises: at least one light-detecting element configured to collect light of at least a first wavelength reflected from a tissue region including a wound; and One or more processors in communication with the at least one light detecting element. the one or more processors: receiving a signal from the at least one light-detecting element indicative of light at the first wavelength reflected from the tissue region; generating an image having a plurality of pixels representing the tissue region based on the signals; measuring a reflected intensity value of each pixel at a first wavelength for at least a subset of the plurality of pixels based on the signal; determining one or more quantitative characteristics of the subset of the plurality of pixels based on the reflectance intensity value of each pixel in the subset; The method is configured to use one or more machine learning algorithms to generate at least one scalar value corresponding to a healing prediction parameter or a healing assessment parameter after a predetermined period of time based on the one or more quantitative features of the subset of the plurality of pixels.
[0008] In some embodiments, the wound is a diabetic foot ulcer. In some embodiments, the healing prediction parameter is a predicted amount of healing of the wound. In some embodiments, the healing prediction parameter is a predicted % area reduction of the wound. In some embodiments, the at least one scalar value comprises a plurality of scalar values, each scalar value corresponding to the healing potential of an individual pixel of the subset or a subgroup of pixels of the subset. In some embodiments, the one or more processors are further configured to output a visual representation of the plurality of scalar values for display to a user. In some embodiments, the visual representation comprises the image in which each pixel of the subset is displayed with a particular visual representation selected based on the healing potential of each pixel, with pixels with different healing potentials being displayed with different visual representations. In some embodiments, the one or more machine learning algorithms comprise SegNet pre-trained using an image database of wounds, burns, or ulcers. In some embodiments, the wound image database comprises a diabetic foot ulcer image database. In some embodiments, the wound image database comprises a burn image database. In some embodiments, the predetermined period is 30 days. In some embodiments, the one or more processors are further configured to identify at least one patient health measure corresponding to a patient having the tissue region, wherein the at least one scalar value is generated based on the one or more quantitative features of the subset of the plurality of pixels and the at least one patient health measure. In some embodiments, the at least one patient health measure includes at least one variable selected from the group consisting of demographic variables, diabetic foot ulcer history variables, compliance variables, endocrine system variables, cardiovascular system variables, musculoskeletal system variables, nutritional status variables, infection variables, renal variables, obstetric-gynecological variables, medication use variables, other disease variables, and clinical laboratory values.In some embodiments, the at least one patient health measure comprises one or more clinical features. In some embodiments, the one or more clinical features comprise at least one feature selected from the group consisting of the patient's age, the patient's degree of chronic kidney disease, the length of the wound on the day the image was created, and the width of the wound on the day the image was created. In some embodiments, the first wavelength is within the range of 420 nm ± 20 nm, 525 nm ± 35 nm, 581 nm ± 20 nm, 620 nm ± 20 nm, 660 nm ± 20 nm, 726 nm ± 41 nm, 820 nm ± 20 nm, or 855 nm ± 30 nm. In some embodiments, the first wavelength is within the range of 620 nm ± 20 nm, 660 nm ± 20 nm, or 420 nm ± 20 nm. In some embodiments, the one or more machine learning algorithms comprise a random forest ensemble. In some embodiments, the first wavelength is within the range of 726 nm±41 nm, 855 nm±30 nm, 525 nm±35 nm, 581 nm±20 nm, or 820 nm±20 nm. In some embodiments, the one or more machine learning algorithms include an ensemble of classifiers. In some embodiments, the system further includes an optical bandpass filter configured to pass light of at least the first wavelength. In some embodiments, the one or more processors are further configured to automatically segment pixels in the image into wound pixels and non-wound pixels; and select a subset of the plurality of pixels that includes the wound pixels. In some embodiments, the one or more processors are further configured to automatically segment the non-wound pixels into callus pixels and background pixels. In some embodiments, the one or more processors are further configured to automatically segment the non-wound pixels into callus pixels, normal skin pixels, and background pixels. In some embodiments, the one or more processors automatically segment the plurality of pixels using a segmentation algorithm that includes a convolutional neural network.In some embodiments, the segmentation algorithm is at least one of a U-Net including multiple convolutional layers and a SegNet including multiple convolutional layers. In some embodiments, the one or more quantitative features of the subset of the plurality of pixels include one or more aggregated quantitative features of the plurality of pixels. In some embodiments, the aggregated one or more quantitative features of the subset of the plurality of pixels are selected from the group consisting of a mean of the reflectance intensity values of pixels of the subset, a standard deviation of the reflectance intensity values of pixels of the subset, and a median of the reflectance intensity values of pixels of the subset. In some embodiments, the one or more processors are further configured to: create a plurality of transformed images by individually applying a plurality of filter kernels to the image by convolution; construct a three-dimensional matrix from the plurality of transformed images; and determine one or more quantitative features of the three-dimensional matrix, wherein the at least one scalar value is generated based on the one or more quantitative features of the subset of the plurality of pixels and based on the one or more quantitative features of the three-dimensional matrix. In some embodiments, the one or more quantitative features of the three-dimensional matrix are selected from the group consisting of a mean value of the numerical values of the three-dimensional matrix, a standard deviation of the numerical values of the three-dimensional matrix, a median of the three-dimensional matrix, and a product of the mean and median of the three-dimensional matrix. In some embodiments, the at least one scalar value is generated based on the mean value of the reflectance intensity values of the pixels of the subset, the standard deviation of the reflectance intensity values of the pixels of the subset, the median of the reflectance intensity values of the pixels of the subset, the mean value of the numerical values of the three-dimensional matrix, the standard deviation of the numerical values of the three-dimensional matrix, and the median of the three-dimensional matrix.In some embodiments, the at least one light detecting element is further configured to collect light of at least a second wavelength reflected from the tissue region, and the one or more processors are further configured to: receive from the at least one light detecting element a second signal indicative of the light of the second wavelength reflected from the tissue region; measure a reflection intensity value of each pixel at the second wavelength for at least a subset of the plurality of pixels based on the second signal; determine one or more additional quantitative features of the subset of the plurality of pixels based on the reflection intensity values of each pixel at the second wavelength; and generate the at least one scalar value based at least in part on the one or more additional quantitative features of the subset of the plurality of pixels.
[0009] In a second aspect, a system for wound assessment comprises: at least one light-detecting element configured to collect light of at least a first wavelength reflected from a tissue region including a wound; and One or more processors in communication with the at least one light detecting element. the one or more processors: receiving a signal from the at least one light-detecting element indicative of light at the first wavelength reflected from the tissue region; generating an image having a plurality of pixels representing the tissue region based on the signals; measuring a reflected intensity value of each of the plurality of pixels at a first wavelength based on the signal; The method is configured to use a machine learning algorithm to automatically segment each of the plurality of pixels into at least a first subset of the plurality of pixels including wound pixels and a second subset of the plurality of pixels including non-wound pixels based on the reflectance intensity value of each of the plurality of pixels.
[0010] In some embodiments, the one or more processors are further configured to automatically segment the second subset of pixels into at least two types of non-wound pixels, the at least two types of pixels being selected from the group consisting of callus pixels, normal skin pixels, and background pixels. In some embodiments, the machine learning algorithm comprises a convolutional neural network. In some embodiments, the machine learning algorithm is at least one of a U-Net including multiple convolutional layers and a SegNet including multiple convolutional layers. In some embodiments, the machine learning algorithm is trained based on a dataset including a plurality of segmented images of wounds, ulcers, or burns. In some embodiments, the wound is a diabetic foot ulcer. In some embodiments, the one or more processors are further configured to output a visual representation of the segmented plurality of pixels for display to a user. In some embodiments, the visual representation includes the image in which each pixel is displayed with a particular visual representation selected based on the segmented pixels, and wound pixels and non-wound pixels are displayed with different visual representations.
[0011] In another aspect, a method for predicting wound healing using a system for assessing or predicting wound healing comprises: illuminating a tissue region with light at least at a first wavelength and reflecting at least a portion of the light from the tissue region onto at least one light-detecting element; generating at least one scalar value using the wound healing assessment or prediction system; and Determining healing prediction parameters after a predetermined period of time Includes.
[0012] In some embodiments, illuminating the tissue region comprises activating one or more light emitting elements configured to emit light at least at the first wavelength. In some embodiments, illuminating the tissue region comprises exposing the tissue region to ambient light. In some embodiments, determining the healing prediction parameter comprises determining an expected % reduction in wound area after a predetermined period of time. In some embodiments, the method further comprises: measuring one or more dimensions of the wound a predetermined period of time after determining the predicted amount of healing of the wound; measuring the actual amount of healing of the wound after the predetermined period of time; and updating at least one of one or more machine learning algorithms by providing at least the images and the actual amount of healing of the wound as training data. Further includes: In some embodiments, the method further comprises selecting a standard wound care therapy or an advanced wound care therapy based at least in part on the healing prediction parameter. In some embodiments, the step of selecting a standard or advanced wound care therapy comprises prescribing or applying one or more standard therapies selected from the group consisting of optimizing nutritional status, debridement to remove necrotic tissue by any means, maintaining a clean and moist granulation tissue bed with an appropriate moist dressing, therapy required to resolve any infection that may exist, addressing lack of intravascular perfusion in the limb with the diabetic foot ulcer, reducing pressure load caused by the diabetic foot ulcer, and appropriate glycemic control if the healing prediction parameter indicates that the wound (preferably the diabetic foot ulcer) will not heal or close by more than 50% after 30 days; and prescribing or applying one or more advanced therapies selected from the group consisting of hyperbaric oxygen therapy, negative pressure wound therapy, bioengineered skin substitutes, synthetic growth factors, extracellular matrix proteins, matrix metalloproteinase modulators, and electrical stimulation therapy if the healing prediction parameter indicates that the wound (preferably the diabetic foot ulcer) will not heal or close by more than 50% after 30 days. [Brief explanation of the drawings]
[0013] [Figure 1A] 1 shows an example of light incident on a filter at various chief ray angles of incidence.
[0014] [Figure 1B] 1B is a graph showing an example of transmission efficiency obtained from the filter of FIG. 1A through which light is transmitted at various chief ray angles of incidence.
[0015] [Figure 2A] 1 shows an example of a multispectral image data cube.
[0016] [Figure 2B]An example of how a particular multispectral imaging technique generates the data cube of FIG. 2A is shown.
[0017] [Figure 2C] 2B shows an example of a snapshot imaging system capable of generating the data cube of FIG. 2A.
[0018] [Figure 3A] 1 shows a cross-sectional schematic diagram illustrating the optical design of an example multi-aperture imaging system with curved multi-bandpass filters according to the present disclosure.
[0019] [Figure 3B-3D] 3B shows an example of an optical design of optical components that constitute one optical path of the multi-aperture imaging system shown in FIG. 3A.
[0020] [Figures 4A-4E] 3A and 3B illustrate an embodiment of a multispectral, multi-aperture imaging system with the optical design described in connection with FIGS. 3A and 3B.
[0021] [Figure 5] 3C illustrates another embodiment of a multispectral, multi-aperture imaging system with the optical design described in connection with FIGS. 3A and 3B.
[0022] [Figures 6A-6C] 3C illustrates another embodiment of a multispectral, multi-aperture imaging system with the optical design described in connection with FIGS. 3A and 3B.
[0023] [Figures 7A-7B] 3C illustrates another embodiment of a multispectral, multi-aperture imaging system with the optical design described in connection with FIGS. 3A and 3B.
[0024] [Figure 8A-8B]3C illustrates another embodiment of a multispectral, multi-aperture imaging system with the optical design described in connection with FIGS. 3A and 3B.
[0025] [Figures 9A-9C] 3C illustrates another embodiment of a multispectral, multi-aperture imaging system with the optical design described in connection with FIGS. 3A and 3B.
[0026] [Figures 10A-10B] 3C illustrates another embodiment of a multispectral, multi-aperture imaging system with the optical design described in connection with FIGS. 3A and 3B.
[0027] [Figures 11A-11B] 1 shows an example of a set of frequency bands that can pass through the filters of the multispectral, multi-aperture imaging system shown in FIGS.
[0028] [Figure 12] A simplified block diagram of an imaging system that can be used for the multispectral and multi-aperture imaging system shown in Figures 3A-10B is shown.
[0029] [Figure 13] 10B. FIG. 10 is a flowchart illustrating an example process for capturing image data using the multispectral, multiaperture imaging system illustrated in FIGS.
[0030] [Figure 14] 13 shows a simplified block diagram of a workflow for processing image data, such as image data captured using the multispectral, multi-aperture imaging system shown in FIGS. 3A-10B and / or using the process shown in FIG. 13.
[0031] [Figure 15]13 illustrates parallax and parallax correction for processing image data, such as image data captured using the multispectral, multi-aperture imaging system shown in FIGS. 3A-10B and / or using the process shown in FIG. 13.
[0032] [Figure 16] Illustrated is a workflow for pixel-by-pixel classification of multispectral image data, such as image data captured using the multispectral, multi-aperture imaging system shown in FIGS. 3A-10B and / or using the process shown in FIG. 13 and then processed according to FIGS. 14 and 15.
[0033] [Figure 17] FIG. 10 shows a simplified block diagram of an example of a distributed computing system including the multispectral multi-aperture imaging system shown in FIGS. 3A to 10B.
[0034] [Figures 18A-18C] 1 illustrates an example embodiment of a portable multispectral and multiaperture imaging system.
[0035] [Figures 19A-19B] 1 illustrates an example embodiment of a portable multispectral and multiaperture imaging system.
[0036] [Figures 20A-20B] This shows an example of a compact multispectral and multiaperture imaging system for USB 3.0, housed in a standard camera housing.
[0037] [Figure 21] An example of a multispectral and multiaperture imaging system that includes an additional light source to improve image alignment is shown.
[0038] [Figure 22]An example of the time course of the progression of healing of a diabetic foot ulcer (DFU) and corresponding measurements of DFU area, volume, and debridement are shown.
[0039] [Figure 23] An example of the time course of healing progression of a non-healing diabetic foot ulcer (DFU) is shown, along with corresponding measurements of DFU area, volume, and debridement.
[0040] [Figure 24] 1 shows a schematic diagram of an example machine learning system for generating a healing prediction based on one or more images taken of a DFU.
[0041] [Figure 25] 1 shows a schematic diagram of an example machine learning system for generating a healing prediction based on one or more images taken of a DFU and one or more patient health measurements.
[0042] [Figure 26] 1 illustrates an example set of wavelength bands used in spectral and / or multispectral imaging for image segmentation and / or generation of predictive parameters in accordance with the techniques of the present invention.
[0043] [Figure 27] 1 is a histogram showing the effect of including clinical variables in an exemplary wound assessment method according to the present technology.
[0044] [Figure 28] 1 shows a schematic diagram of an example of an autoencoder in accordance with a machine learning system and method according to the present technique;
[0045] [Figure 29] 1 shows a schematic diagram of an example of a supervised machine learning algorithm in accordance with the machine learning system and method of the present technology;
[0046] [Figure 30] 1 shows a schematic diagram of an example of an end-to-end machine learning algorithm in accordance with the machine learning system and method of the present technology;
[0047] [Figure 31] 1 is a bar graph illustrating the accuracy exhibited by several example machine learning algorithms in accordance with the present technique.
[0048] [Figure 32] 1 is a bar graph illustrating the accuracy exhibited by several example machine learning algorithms in accordance with the present technique.
[0049] [Figure 33] 1 shows a schematic diagram of an example process for predicting healing and generating a visual display of a conditional probability mapping according to a machine learning system and method in accordance with the present technology.
[0050] [Figure 34] 1 shows a schematic diagram of an example of a conditional probability mapping algorithm that includes one or more feature-wise linear transformation layers (FiLM).
[0051] [Figure 35] 1 illustrates the accuracy demonstrated by several image segmentation approaches for creating conditional cure probability maps in accordance with the present techniques.
[0052] [Figure 36] 1 illustrates an example set of convolution filter kernels used in an example single wavelength analysis method for predicting healing in accordance with machine learning systems and methods according to the present technology.
[0053] [Figure 37]1 shows an example of a true value mask created based on a DFU image for image segmentation according to the machine learning system and method of the present technique.
[0054] [Figure 38] 1 illustrates the accuracy demonstrated by an exemplary segmentation algorithm for wound images in accordance with the machine learning system and method of the present technology. DETAILED DESCRIPTION OF THE INVENTION
[0055] Of the 26 million Americans living with diabetes, approximately 15–25% develop diabetic foot ulcers (DFUs). DFUs limit patients' mobility and reduce their quality of life. Up to 40% of patients with DFUs also develop wound infections, which increase the risk of limb amputation and death. DFUs alone are associated with a 5% first-year mortality rate and a 42% 5-year mortality rate. This is reflected in the high annual risk of major amputation (4.7%) and minor amputation (39.8%). Furthermore, the annual cost of treating a single DFU ranges from approximately $22,000 to $44,000, resulting in a total economic burden of DFUs to the U.S. healthcare system of $9–13 billion per year.
[0056] DFUs with a 30-day percent area reduction (PAR) of >50% are generally considered to heal by 12 weeks with standard wound care. However, using such measurements requires an initial 4-week wound care period to determine whether more effective treatments (e.g., advanced wound care) should be used. When a non-emergency wound, such as a DFU, first appears, the typical clinical response to wound care is to administer standard wound care therapy (e.g., correction of vascular defects, optimization of nutritional status, glycemic control, debridement, dressings, and / or pressure relief) to the patient for approximately 30 days from wound onset and initial assessment. At approximately 30 days, the wound is assessed for healing (e.g., percent area reduction >50%). If the wound has not healed adequately, treatment is supported by one or more advanced wound management therapies. These advanced wound management therapies include growth factors, bioengineered tissue, hyperbaric oxygen, negative pressure, amputation, and recombinant human platelet-derived growth factor (e.g., Regranex). TM gel), bioengineered human dermal substitutes (e.g., Dermagraft TM ), and / or two-layer living skin substitutes (e.g., Apligraf TM However, approximately 60% of DFUs fail to heal adequately even after 30 days of standard wound care therapy. Furthermore, even among fast-healing DFUs, approximately 40% do not heal by 12 weeks, with median healing times estimated at 147 days for toe ulcers, 188 days for midfoot ulcers, and 237 days for heel ulcers.
[0057] If a DFU fails to achieve the desired healing after 30 days of conventional or standard wound care therapy, it may be beneficial to initiate advanced wound care therapy as soon as possible (e.g., within 30 days of initiating wound therapy). However, using conventional assessment methods, physicians typically cannot accurately identify DFUs that do not respond to 30 days of standard wound care therapy. Various strategies aimed at improving DFU treatment are available, but they are prescribed only after the standard wound care therapy has been implemented. Physiological measurement devices, such as transcutaneous oxygen tension measurement, laser Doppler imaging, and indocyanine green video angiography, have also been used in attempts to diagnose the healing potential of DFUs. However, these devices are not suitable for widespread use in the evaluation of DFUs and other wounds due to imprecise results, lack of useful data, poor sensitivity, and prohibitive cost. It is clear that earlier and more accurate methods for predicting the healing of DFUs and other wounds are important to quickly determine the best treatment and reduce the time required for wound closure.
[0058] As outlined herein, the technology of the present invention provides a non-invasive, point-of-care, real-time imaging device capable of diagnosing the healing potential of DFUs, burns, and other wounds. In various embodiments, the systems and methods of the technology of the present invention enable clinicians to determine the healing potential of a wound at the time of wound presentation or initial assessment, or shortly thereafter. In some embodiments, the technology of the present invention can determine the healing potential of individual wound sites (such as DFUs or burns). Based on this predicted healing potential, a decision can be made on whether to administer standard or advanced wound care therapy on or about day 0 of treatment, without delaying the wound for more than four weeks after initial wound presentation. Thus, the technology of the present invention has the potential to shorten healing time and reduce the risk of amputation.
[0059] Examples of spectral and multispectral imaging systems Various spectral and multispectral imaging systems that may be used in accordance with the methods for assessing, predicting, and treating DFUs and other wounds disclosed herein are described below. In some embodiments, images for wound assessment may be captured with a spectral imaging system configured to image light within a single wavelength band. In other embodiments, images may be captured with a spectral imaging system configured to capture two or more wavelength bands. In one particular example, images may be captured with a monochromatic, RGB, and / or infrared imaging device, such as those found in commercially available mobile devices. Yet another embodiment relates to spectral imaging using a multi-aperture system with curved, multi-bandpass filters positioned above each aperture. However, it will be understood that the wound assessment, prediction, and treatment methods according to the present technology are not limited to the specific image acquisition devices disclosed herein and may similarly be implemented using any imaging device capable of acquiring image data in one or more known wavelength bands.
[0060] The present disclosure also relates to techniques for implementing spectral unmixing and image registration to generate a spectral data cube using image information received from such imaging systems. The techniques of the present disclosure, as described below, address various challenges typically encountered in spectral imaging when obtaining image data that accurately reflects wavelength bands from an object being imaged. In some embodiments, the systems and methods described herein can acquire images of a large area of tissue (e.g., 5.9 x 7.9 inches) in a short time (e.g., within 6 seconds), and can achieve this without the need for injection of an imaging contrast agent. In some aspects, for example, the multispectral imaging systems described herein are configured to acquire images of a large area of tissue, e.g., 5.9 x 7.9 inches, in 6 seconds or less, and provide tissue analysis information, such as identification of multiple burn conditions, multiple wound conditions, multiple ulcer conditions, or healing potential; or clinical characteristics, including the cancerous or non-cancerous condition of the imaged tissue, wound depth, wound volume, debridement margin, and the presence of diabetic, non-diabetic, or chronic ulcers, in the absence of an imaging contrast agent. Similarly, in some of the methods described herein, the multispectral imaging system of the present invention can acquire images from a large area of tissue, e.g., 5.9 x 7.9 inches, in less than 6 seconds and output tissue analysis information such as identification of multiple burn conditions, multiple wound conditions, or healing potential; or clinical characteristics including cancerous or non-cancerous condition of the imaged tissue, wound depth, wound volume, debridement margins, and the presence of diabetic, non-diabetic, or chronic ulcers, in the absence of imaging contrast agents.
[0061] One challenge with existing solutions is color distortion or parallax in captured images, which can impair the quality of image data. This can be particularly problematic in applications that rely on the accurate detection and analysis of specific wavelengths of light using optical filters. More specifically, color tinting is a shift in the wavelength of light as it passes through the image sensor. This occurs because the greater the angle of incidence of light on a color filter, the shorter the wavelength of light that passes through it. This effect is typically observed in interference filters, which are fabricated by stacking multiple thin layers of varying refractive index on a transparent substrate. Therefore, longer wavelengths (e.g., red light) are more likely to be blocked at the edges of the image sensor due to the larger angle of incidence of the light, resulting in spatially inconsistent color detection of incident light of the same wavelength as it passes through the image sensor. Without image correction, color tinting manifests as color variations near the edges of a captured image.
[0062] The disclosed technology offers numerous advantages over other commercially available multispectral imaging systems because it is not limited by the lens and / or image sensor configuration and their field of view or aperture size. As those skilled in the art will appreciate, changes to the lenses, image sensors, aperture sizes, or other components of the imaging systems disclosed herein may also include other adjustments to the imaging system. The disclosed technology also offers improvements over other multispectral imaging systems in that it allows components that perform wavelength division, or wavelength division using the entire imaging system (e.g., optical filters), to be separated from components that convert optical energy into a digital output (e.g., image sensors). This can reduce cost, complexity, and / or development time required to reconfigure the imaging system for different multispectral wavelengths. The disclosed technology is also believed to be more robust than other commercially available multispectral imaging systems, in that it can achieve imaging performance comparable to other commercially available multispectral imaging systems in a smaller, lighter form factor. The disclosed techniques also have advantages over other multispectral imaging systems in that they can acquire multispectral images at snapshot, video, or high-speed video rates. Furthermore, the disclosed techniques can multiplex several spectral bands into each aperture, thereby reducing the number of apertures required to acquire a particular number of spectral bands in an imaging dataset, and can provide a more robust implementation of multispectral imaging systems based on multi-aperture technology, due to cost savings from the reduced number of apertures and improved light collection (e.g., commercially available sensor arrays can use larger apertures that fit within fixed sizes and dimensions). Furthermore, the disclosed techniques can provide all of these advantages without compromising resolution or image quality.
[0063] FIG. 1A illustrates an example of a filter 108 positioned along a light path toward an image sensor 110, with light incident on the filter 108 at various angles of incidence. Light rays 102A, 104A, and 106A are shown as rays that pass through the filter 108 and are then refracted by a lens 112 onto the sensor 110. The lens 112 may be replaced with other image-forming optics, such as, but not limited to, a mirror and / or an aperture. In FIG. 1A, the light in each ray is considered broadband, e.g., light having a spectral composition across a wide wavelength range, with only light of specific wavelengths being selectively transmitted by the filter 108. The three light rays 102A, 104A, and 106A are incident on the filter 108 at different angles of incidence. For purposes of illustration, light ray 102A is shown as being substantially normal incident on filter 108, with light ray 104A having a larger angle of incidence than light ray 102A, and light ray 106A having a larger angle of incidence than light ray 104A. Each light ray 102B, 104B, and 106B passing through the filter exhibits a unique spectrum due to the angle-of-incidence-dependent transmission characteristics of filter 108, which are detected by sensor 110. This angle-of-incidence dependency effect causes the bandpass characteristics of filter 108 to shift toward shorter wavelengths as the angle of incidence increases. Furthermore, angle-of-incidence dependency can reduce the transmission efficiency of filter 108 and can change the spectral shape of the bandpass characteristics of filter 108. This combined effect is referred to as angle-of-incidence-dependent spectral transmission. FIG. 1B shows the spectrum of each of the rays in FIG. 1A detected by a spectrometer, hypothetically located at sensor 110, illustrating the spectral shift of the bandpass characteristic of filter 108 as the angle of incidence increases. Curves 102C, 104C, and 106C show the center wavelengths of the bandpass characteristics decreasing, thus representing a decrease in the wavelength of light emitted from the optical system in this example. As shown in this figure, the spectral shape and transmittance peak of the bandpass characteristics also change with the angle of incidence. In certain consumer applications, image processing can be performed to remove the visible effects of this angle-dependent spectral transmission.However, such post-processing techniques cannot recover the exact wavelengths of light that actually hit the filter 108. Therefore, the resulting image data may not be usable for certain high precision applications.
[0064] Another challenge faced by certain existing spectral imaging systems is the time required to capture a complete spectral image data set. This is discussed in connection with Figures 2A and 2B. A spectral image sensor acquires the spectral irradiance I(x, y, λ) of a particular scene and collects a three-dimensional (3D) data set, commonly referred to as a data cube. Figure 2A shows an example of a spectral image data cube 120. As shown in this figure, the data cube 120 represents three-dimensional image data, where two spatial dimensions (x and y) correspond to the two-dimensional (2D) surface of the image sensor and the spectral dimension (λ) corresponds to a particular wavelength range. The size of the data cube 120 is N x N y N λ and N x and N y is the number of sample points in each spatial dimension (x, y), and N λ is the number of sample points along the spectral axis λ. Because the data cube has a higher dimensionality than currently available 2D detector arrays (e.g., image sensors), typical spectral imaging systems simultaneously measure all sample points of the data cube 120 by taking time-sequential 2D slices (i.e., planes) of the data cube 120 (referred to herein as a "scanning" imaging system) or by dividing the data cube 120 in the computational process to obtain multiple two-dimensional data sets that can be reassembled into the data cube 120 (referred to herein as a "snapshot" imaging system).
[0065] FIG. 2B shows an example of how a particular scanning spectral imaging technique generates data cube 120. Specifically, FIG. 2B shows portions 132, 134, and 136 of data cube 120 that can be collected during a single detector integration time. A point-scanning spectrometer, for example, can image portion 132, which extends across the entire spectral plane λ, at a single spatial location (x, y). A point-scanning spectrometer can construct data cube 120 by performing multiple integrations for each (x, y) location in the spatial dimensions. A filter wheel imaging system, for example, can image portion 134, which extends across the entire spatial dimensions x and y, but is located in only a single spectral plane λ. A wavelength-scanning imaging system, such as a filter wheel imaging system, can construct data cube 120 by performing multiple integrations across the spectral plane λ. A line-scanning spectrometer, for example, can image portion 136, which extends across the entire spectral dimension λ and one of the spatial dimensions (x or y), but is located at only a single point in the other spatial dimension (y or x). A line-scan spectrometer can construct a data cube 120 by performing multiple integrations over one spatial dimension (y or x) position.
[0066] In applications where both the object and the imaging system are stationary (or relatively stationary over the exposure time), the aforementioned scanning imaging systems have the advantage of providing a high-resolution data cube 120. Line-scan and wavelength-scanning imaging systems may benefit from using the entire image sensor for each spectral or spatial image. However, motion of the imaging system and / or object between exposures may introduce artifacts into the resulting image data. For example, the same (x,y) location in the data cube 120 may actually represent a different physical location on the imaged object in the spectral dimension λ. Such artifacts may introduce errors in subsequent analysis and / or require alignment (e.g., adjusting the spectral dimension λ so that a particular (x,y) location corresponds to the same physical location on the object).
[0067] In contrast, a snapshot imaging system 140 can capture the entire datacube 120 in a single integration time or exposure, thereby avoiding the impact of motion on image quality. Figure 2C shows an example of an image sensor 142 and an optical filter array 144 (e.g., a color filter array (CFA)) that can be used to create a snapshot imaging system. In this example, the color filter array (CFA) 144 has a repeating pattern of color filter units 146 arranged on the surface of the image sensor 142. This method of obtaining spectral information can also be referred to as a multispectral filter array (MSFA) or a spectrally resolved detector array (SRDA). In the example shown, the color filter unit 146 includes various color filters arranged in a 5x5 pattern, thereby generating 25 spectral channels in the resulting image data. The CFA can use the various color filters to split incident light into bands for each color filter and direct the split light to photoreceptors of each color on the image sensor. In this way, for a given color 148, only one of the 25 photoreceptors actually detects a signal representing light of that color wavelength. Thus, using such a snapshot imaging system 140, a single exposure can produce 25 color channels, but the amount of measurement data for each color channel is less than the total output of the image sensor 142. In some embodiments, the CFA may include one or both of a filter array (MSFA) and a spectrally resolved detector array (SRDA), and / or may include conventional Bayer filters, CMYK filters, or other absorptive or interferometric filters. One type of interferometric filter is a thin-film filter array arranged in a grid, with each grid element corresponding to one or more sensor elements. Another type of interferometric filter is a Fabry-Perot filter.Nano-etched Fabry-Perot interference filters typically exhibit bandpass characteristics with a full width at half maximum (FWHM) of approximately 20-50 nm. However, they can be advantageously used in some embodiments because of their low roll-off at the transition from the center wavelength of the filter's passband to its stopband. Furthermore, such filters exhibit a low optical density (OD) in the stopband, thereby enhancing sensitivity to light outside the passband. This combined effect allows these particular filters to be sensitive to spectral regions that would be blocked by the high roll-off of similar high-OD interference filters with similar FWHMs, constructed from multiple thin-film layers deposited by coating techniques such as evaporation or ion beam sputtering. In embodiments using dye-based CMYK filters or RGB (Bayer) filters, low spectral roll-off and a large FWHM at the passband of each individual filter are preferred, allowing for unique spectral transmittance for each wavelength across the observed spectrum.
[0068] Therefore, the data cube 120 obtained by a snapshot imaging system has one of the following two characteristics that can be problematic for high-precision imaging applications: First, the data cube 120 obtained by a snapshot imaging system has a size N smaller than the (x, y) dimension of the detector array. x and N y The second problem is that the data cube 120 obtained by a snapshot imaging system has a resolution of N due to the small size of N due to numerical interpolation for a particular (x,y) position. x and N y The dimensions of the (x,y) detector array may be the same as the (x,y) dimensions of the detector array. However, if interpolation was used in generating the data cube, this means that any particular value in the data cube is not an actual measurement of the wavelength of light incident on the sensor, but an estimate of that measurement based on surrounding values.
[0069] Another existing component used in multispectral imaging with a single exposure is the multispectral imaging beamsplitter. In such imaging systems, a beamsplitter cube splits the incident light into different color bands, each of which is observed by multiple independent image sensors. While the beamsplitter design can be modified to adjust the spectral band being measured, it is difficult to split the incident light into more than four beams without compromising the system's performance. Therefore, four spectral channels appears to be the practical limit of this approach. A closely related method uses thin-film filters to split the light instead of bulky beamsplitter cubes / prisms, but the cumulative transmittance attenuation and spatial limitations of multiple successive filters still limit this approach to approximately six spectral channels.
[0070] In some embodiments, the aforementioned challenges, among others, can be addressed by the multi-aperture spectral imaging system and associated image data processing techniques disclosed herein, which includes a multi-bandpass filter (preferably a curved multi-bandpass filter) that selectively transmits light incident through each aperture. This particular configuration achieves all of the design goals of fast imaging speed, high-resolution images, and high fidelity of the detected wavelength. Thus, the optical design and associated image data processing techniques disclosed herein can be used in portable spectral imaging systems and / or for imaging moving subjects, while generating data cubes suitable for high-precision applications (e.g., clinical tissue analysis, biometric identification, and episodic clinical indications). Such high-precision applications can include diagnosing the stage (0-3) of pre-metastatic malignant melanoma, classifying the severity of wounds or burns in skin tissue, or histologically diagnosing the severity of diabetic foot ulcers. Thus, the snapshot spectral acquisition and small form factor described in some embodiments enable the present invention to be used in clinical settings dealing with transient events, including diagnosing various types of retinopathies (e.g., non-proliferative diabetic retinopathy, proliferative diabetic retinopathy, and age-related macular degeneration) and imaging pediatric patients with high levels of motion. Therefore, those skilled in the art will appreciate that the use of the multi-aperture systems with flat or curved multiple bandpass filters disclosed herein represents a significant technological advance over conventional spectral imaging implementations. More specifically, the multi-aperture systems of the present invention may enable the collection of 3D spatial images of or relating to the curvature, depth, volume, and / or area of an object based on the calculated perspective difference between each aperture. However, the multi-aperture methods disclosed herein are not limited to any particular filters and may include flat and / or thin filters based on interference or absorption filtering.The invention disclosed herein can be modified to include flat filters in the image space of an imaging system when the angle of incidence to an appropriate lens or aperture is small or within acceptable limits. Filters may be placed at the entrance and / or exit pupil of an imaging lens or aperture stop, and such configurations would be readily apparent to one skilled in the art of optical engineering.
[0071] Various aspects of the present disclosure are described below with reference to specific examples and embodiments. These examples and embodiments are for illustrative purposes only and are not intended to limit the present disclosure. While the examples and embodiments described herein focus on particular calculations and algorithms for illustrative purposes, those skilled in the art will appreciate that these examples are for illustrative purposes only and are not intended to limit the present disclosure. For example, while certain examples are provided with respect to multispectral imaging, the multi-aperture imaging systems and associated filters disclosed herein can be configured to perform hyperspectral imaging in other implementations. Furthermore, while specific examples are presented as achieving advantages in portable applications and / or applications involving moving objects, it will be appreciated that the imaging system designs and associated processing techniques disclosed herein can produce highly accurate data cubes suitable for use with fixed imaging systems and / or for analyzing relatively stationary objects.
[0072] Overview of Electromagnetic Spectral Range and Image Sensors Reference is made herein to particular colors or portions of the electromagnetic spectrum, and these wavelengths are hereinafter described as defined by ISO 21348, "Definitions of Irradiance Spectral Types." As discussed in more detail below, wavelength ranges of particular colors may be passed collectively through particular filters in certain imaging applications.
[0073] Electromagnetic radiation in the wavelength range of 760 nm to 380 nm, or about 760 nm to about 380 nm, is typically considered to be the "visible" spectrum, i.e., the portion of the spectrum that is perceptible by the color receptors of the human eye. Within the visible spectrum, red light is typically considered to have a wavelength of 700 nm or about 700 nm, or in the range of 760 nm to 610 nm or about 760 nm to about 610 nm. Orange light is typically considered to have a wavelength of 600 nm or about 600 nm, or in the range of 610 nm to 591 nm or about 610 nm to about 591 nm. Yellow light is typically considered to have a wavelength of 580 nm or about 580 nm, or in the range of 591 nm to 570 nm or about 591 nm to about 570 nm. Green light is generally considered to have a wavelength of 550 nm or about 550 nm, or in the range of 570 nm to 500 nm or about 570 nm to 500 nm. Blue light is generally considered to have a wavelength of 475 nm or about 475 nm, or in the range of 500 nm to 450 nm or about 500 nm to 450 nm. Purple (magenta) light is generally considered to have a wavelength of 400 nm or about 400 nm, or in the range of 450 nm to 360 nm or about 450 nm to 360 nm.
[0074] Looking beyond the visible spectrum, infrared (IR) refers to electromagnetic radiation with wavelengths longer than visible light and generally invisible to the human eye. IR has wavelengths extending from the nominal lower limit of the red visible spectrum at about 760 nm or 1 mm to about 1 mm or 1 mm. Within this range, near-infrared (NIR) refers to the portion of the spectrum adjacent to the red range, ranging in wavelength from about 760 nm to about 1400 nm or 760 nm to 1400 nm.
[0075] Ultraviolet (UV) light refers to a portion of electromagnetic radiation with wavelengths shorter than visible light and generally invisible to the human eye. UV has wavelengths ranging from the nominal lower limit of the visible spectrum at about 40 nm or violet at 40 nm to about 400 nm. Within this range, near ultraviolet (NUV) refers to the portion of the spectrum adjacent to the violet range and has wavelengths ranging from about 400 nm to about 300 nm or 400 nm to 300 nm; middle ultraviolet (MUV) has wavelengths ranging from about 300 nm to about 200 nm or 300 nm to 200 nm; and far ultraviolet (FUV) has wavelengths ranging from about 200 nm to about 122 nm or 200 nm to 122 nm.
[0076] The image sensors described herein can be configured to detect any of the above-mentioned ranges of electromagnetic radiation, depending on the particular wavelength range appropriate for a particular application. Typical silicon charge-coupled device (CCD) or complementary metal-oxide semiconductor (CMOS) sensors have spectral sensitivity spanning the entire visible spectrum, as well as most of the near-infrared (IR) spectrum and parts of the UV spectrum. In some implementations, back-illuminated or front-illuminated CCD or CMOS arrays can additionally or alternatively be used. For applications requiring high signal-to-noise ratios and scientific-grade measurements, some implementations can additionally or alternatively use scientific complementary metal-oxide semiconductor (sCMOS) cameras or electronically multiplied CCD cameras (EMCCDs). In other implementations, sensors known to operate in specific color ranges (e.g., short-wave infrared (SWIR), mid-wave infrared (MWIR), or long-wave infrared (LWIR)) and corresponding optical filter arrays can additionally or alternatively be used, depending on the intended application. These sensors and optical filter arrays may additionally or alternatively comprise cameras based on detector materials such as indium gallium arsenide (InGaAs) or indium antimony (InSb), or cameras based on microbolometer arrays.
[0077] The image sensors used in the multispectral imaging techniques disclosed herein may be used in conjunction with optical filter arrays, such as color filter arrays (CFAs). Some types of CFAs can split incident light into red (R), green (G), and blue (B) visible regions and direct the split visible light to dedicated red, green, or blue photodiode photoreceptors on the image sensor. A common example of a CFA is the Bayer pattern, which is a color filter array in which RGB color filters are arranged in a specific pattern on a rectangular grid of photodetectors. The Bayer pattern is composed of 50% green, 25% red, and 25% blue, with rows of repeated red and green color filters alternating with rows of repeated blue and green color filters. Some types of CFAs (e.g., CFAs for RGB-NIR sensors) can also split NIR light and direct the split NIR light to dedicated photodiode photoreceptors on the image sensor.
[0078] The wavelength range of the filter components of the CFA may determine the wavelength range represented by each image channel in the captured image. Thus, in various embodiments, the red channel of the image may correspond to the red wavelength range of the color filter, may include portions of yellow and orange light, and may range from about 570 nm to about 760 nm or from 570 nm to 760 nm. In various embodiments, the green channel of the image may correspond to the green wavelength range of the color filter, may include portions of yellow light, and may range from about 570 nm to about 480 nm or from 570 nm to 480 nm. In various embodiments, the blue channel of the image may correspond to the blue wavelength range of the color filter, may include portions of violet light, and may range from about 490 nm to about 400 nm or from 490 nm to 400 nm. Those skilled in the art will appreciate that the exact start and end wavelengths (or portions of the electromagnetic spectrum) defining the colors (e.g., red, green, and blue) of a CFA may vary depending on the implementation of the CFA.
[0079] Furthermore, typical visible-light CFAs transmit light outside the visible spectrum. Therefore, many image sensors limit their sensitivity to infrared light by placing a thin-film infrared-reflective filter in front of the sensor that blocks infrared wavelengths but passes visible light. However, some imaging systems disclosed herein may omit the thin-film infrared-reflective filter and instead pass infrared light. Thus, red, green, and / or blue channels may be used to collect infrared wavelength bands. In some implementations, a blue channel may be used to collect a specific NUV wavelength band. Because the red, green, and blue channels exhibit different spectral responses at each wavelength in a stacked spectral image, each channel having a unique transmission efficiency, a known transmission profile may be used to provide a unique weighting response for each spectral band before unmixing. For example, this weighting response may include the known transmission responses of the red, blue, and green channels in the infrared and ultraviolet wavelength regions, allowing each channel to be used to collect bands from these wavelength regions.
[0080] As described in further detail below, additional color filters can be placed before the CFA along the optical path toward the image sensor to selectively extract specific bands of light incident on the image sensor. The color filters disclosed herein may be a combination of (thin film) dichroic and / or absorptive filters, or a single dichroic and / or absorptive filter. Some of the color filters disclosed herein may be bandpass filters that pass frequencies within a specific region (in the passband) but block (attenuate) frequencies outside that region (in the block region). Some of the color filters disclosed herein may be multi-bandpass filters that pass multiple, discrete wavelength regions. These "frequency bands" may have narrower passbands, greater attenuation in the block region, and steeper spectral roll-off than the broad color range of the CFA filter. Roll-off is defined as the steepness of the spectral response at the transition from the filter's passband to its block region. For example, the color filters disclosed herein can cover passbands from about 20 nm to about 40 nm, or from 20 nm to 40 nm. The particular configuration of the color filter can determine the wavelength bands that are actually incident on the sensor, which can improve the accuracy of the imaging techniques disclosed herein. The color filters described herein can be configured to selectively block or pass specific bands of electromagnetic radiation within the above ranges, depending on the particular wavelength bands that are appropriate for a particular application.
[0081] The term "pixel" is used herein to describe the output generated by an element of a two-dimensional detector array. In contrast, a photodiode, i.e., a single light-sensitive element, of a two-dimensional detector array behaves as a transducer that can convert photons into electrons via the photoelectric effect and then convert the electrons into a signal that can be used to determine a pixel value. An element of a data cube can be referred to as a "voxel" (e.g., an element of volume). A "spectral vector" refers to a vector that represents spectral data (e.g., the spectrum of light received from a particular point in object space) at a particular (x,y) location of the data cube. A single horizontal plane of a data cube (e.g., an image representing one spectral dimension) is referred to herein as an "image channel." In certain embodiments described herein, spectral video information may be captured, and the resulting data dimensions may be N x N y N λ N t (where N t is the number of frames captured during the video sequence).
[0082] Overview of an example of a multi-aperture imaging system with curved multi-bandpass filters FIG. 3A shows a schematic diagram of an example multi-aperture imaging system 200 with a curved multi-bandpass filter according to the present disclosure. The illustrated schematic diagram includes a first image sensor region 225A (photodiodes PD1-PD3) and a second image sensor region 225B (photodiodes PD4-PD6). The photodiodes PD1-PD6 may be, for example, photodiodes formed on a semiconductor substrate (e.g., a CMOS image sensor). Typically, each photodiode PD1-PD6 may be a single unit made of some material, semiconductor, sensor element, or other device capable of converting incident light into electrical current. This diagram shows only a small portion of the entire multi-aperture imaging system for purposes of illustrating the structure and operation of the multi-aperture imaging system; it will be appreciated that in an implementation, the image sensor region may include hundreds or even thousands of photodiodes (and corresponding color filters). The first image sensor region 225A and the second image sensor region 225B may be implemented as separate sensors or as separate regions on the same image sensor, depending on the implementation. While FIG. 3A shows two apertures and corresponding optical paths and sensor areas, it will be appreciated that the optical design principles illustrated in FIG. 3A can be extended to designs including three or more apertures and corresponding optical paths and sensor areas, depending on the implementation.
[0083] The multi-aperture imaging system 200 includes a first opening 210A that provides a first optical path toward a first sensor area 225A and a second opening 210B that provides a first optical path toward a second sensor area 225B. These apertures may be adjustable to increase or decrease the brightness of the light reflected in an image, or they may be adjusted to change the exposure time of a particular image so that the brightness of the light incident on the image sensor area remains constant. These apertures may be located anywhere along the optical axis of the multi-aperture imaging system as would be reasonable for a person skilled in the art of optical design. The optical axes of the optical components located along the first optical path are indicated by dashed lines 230A and 230B, respectively, although it will be appreciated that these dashed lines do not represent the physical structure of the multi-aperture imaging system 200. Optical axis 230A and optical axis 230B are separated by a distance D, which can result in parallax between the images captured by first sensor area 225A and second sensor area 225B. "Parallax" refers to the distance between two corresponding points on the left and right sides (or top and bottom) of a stereoscopic image, where the same physical point in object space appears at different locations in each image. Processing techniques that correct for and utilize this parallax are described in more detail below.
[0084] Optical axis 230A and optical axis 230B pass through center C of the corresponding aperture. Other optical components may also be positioned around these optical axes (e.g., the rotational symmetry points of the optical components may be positioned along the optical axes). For example, first curved multi-bandpass filter 205A and first imaging lens 215A may be positioned around first optical axis 230A, and second curved multi-bandpass filter 205B and second imaging lens 215B may be positioned around second optical axis 230B.
[0085] As used herein, the terms "on" and "above," when used with respect to the placement of optical elements, refer to the location of a particular structure (e.g., a color filter or lens) through which light entering imaging system 200 from object space passes before reaching (or incident on) another structure. To illustrate, along a first optical path, curved multi-bandpass filter 205A is located above aperture 210A, which is located above imaging lens 215A, which is located above CFA 220A, which is located above first image sensor area 225A. Thus, light entering from object space (e.g., the physical space being imaged) first passes through first curved multi-bandpass filter 205A, then aperture 210A, then imaging lens 215A, then CFA 220A, and finally incident on first image sensor area 225A. The second optical path (e.g., the optical path including curved multi-bandpass filter 205B, aperture 210B, imaging lens 215B, CFA 220B, and second image sensor area 225B) may be similarly arranged. In another implementation, apertures 210A, 210B and / or imaging lenses 215A, 215B may be positioned above curved multi-bandpass filters 205A, 205B. Furthermore, another implementation may not use a physical aperture and may rely on a clear aperture in an optical component to control the brightness of the light imaged onto sensor area 225A and sensor area 225B. Thus, lenses 215A, 215B may be positioned above apertures 210A, 210B and curved multi-bandpass filters 205A, 205B. In this implementation, apertures 210A and 210B may be located above lenses 215A and 215B, or apertures 210A and 210B may be located below lenses 215A and 215B, as may be determined necessary by those skilled in the art of optical design.
[0086] The first CFA 220A, located above the first sensor area 225A, and the second CFA 220B, located above the second sensor area 225B, can function as filters that selectively pass specific wavelengths, splitting incident visible light into red, green, and blue regions (denoted by the symbols R, G, and B, respectively). The light is "split" by allowing only the selected wavelengths to pass through the color filters of the first CFA 220A and the second CFA 220B. The split light is received by dedicated red, green, or blue diodes on the image sensor, respectively. While red, blue, and green filters are commonly used, in other embodiments, different color filters can be used depending on the color channels required for the captured image data. For example, filters that selectively pass ultraviolet, infrared, or near-infrared light can be used, similar to an RGB-IR CFA.
[0087] As shown, each filter of the CFA is positioned above a single photodiode PD1-PD6. FIG. 3A also shows an example of a microlens (designated ML) that can be formed or otherwise positioned on each color filter to focus incident light onto the active detector area. Alternative implementations may have multiple photodiodes (e.g., a collection of two, four, or more adjacent photodiodes) below a single filter. In the example shown, photodiodes PD1 and PD4 are positioned below a red filter and therefore output pixel information for a red channel; photodiodes PD2 and PD5 are positioned below a green filter and therefore output pixel information for a green channel; and photodiodes PD3 and PD6 are positioned below a blue filter and therefore output pixel information for a blue channel. Additionally, as described in more detail below, the particular color channel output by a given photodiode can be an even narrower frequency band based on the active light source and / or the particular frequency band passed through the multi-bandpass filters 205A, 205B, thereby allowing a given photodiode to output different image channel information with different exposures.
[0088] The imaging lenses 215A, 215B can be shaped to focus an image of the object scene onto the sensor areas 225A, 225B. Each imaging lens 215A, 215B is not limited to a single convex lens as shown in FIG. 3A , but can be composed of as many optical elements and surfaces as needed to form the image. Various types of commercially available or custom-designed imaging lenses or lens assemblies can be used. The optical elements or lens assemblies can be formed or coupled to each other so that they are housed in series or stacked using an opto-mechanical barrel with a retaining ring or bezel. In some embodiments, the optical elements or lens assemblies can include one or more combined lens groups, which can include two or more optical components cemented or otherwise combined. In various embodiments, the multi-bandpass filters described herein may be placed in front of a lens assembly of a multispectral imaging system, in front of a single lens of a multispectral imaging system, after a lens assembly of a multispectral imaging system, after a single lens of a multispectral imaging system, inside a lens assembly of a multispectral imaging system, inside a combined lens group of a multispectral imaging system, directly on the surface of a single lens of a multispectral imaging system, or directly on the surface of an element of a lens assembly of a multispectral imaging system. Furthermore, apertures 210A and 210B may be eliminated, and lenses 215A and 215B may be of the type typically used in digital single-lens reflex (DSLR) or mirrorless camera photography. Furthermore, lenses 215A and 215B may be of the type used in machine vision, using a C-mount or S-mount thread for attachment.Focus adjustment can be achieved, for example, by moving the imaging lenses 215A, 215B relative to the sensor areas 225A, 225B, or by moving the sensor areas 225A, 225B relative to the imaging lenses 215A, 215B, based on manual focus, contrast-based autofocus, or other suitable autofocus techniques.
[0089] Each multi-bandpass filter 205A, 205B can be configured to selectively pass multiple light beams in a narrow frequency band, e.g., in some embodiments, multiple light beams in a 10-50 nm frequency band (or in other embodiments, multiple light beams in a wider or narrower frequency band). As shown in FIG. 3A, both multi-bandpass filters 205A, 205B can pass a frequency band λc (a "common frequency band"). In implementations with three or more optical paths, each multi-bandpass filter can pass this common frequency band. In this manner, each sensor area captures image information in the same frequency band (a "common channel"). As described in more detail below, this image information obtained in the common channel can be used to align a set of images captured by each sensor area. Some implementations may have one common frequency band and a corresponding common channel, or multiple common frequency bands and corresponding common channels.
[0090] Each of the multiple bandpass filters 205A, 205B can be configured to selectively pass one or more specific frequency bands in addition to the common frequency band λc. In this manner, the imaging system 200 can increase the number of different spectral channels collectively imaged by the sensor areas 205A, 205B beyond the number of spectral channels imaged by a single sensor area. This aspect is advantageous in that the specific frequency bands λc can be selectively passed. u1 and a multi-bandpass filter 205A that passes the specific frequency band λ u2where λ is the wavelength of the light emitted by the multi-bandpass filter 205B. u1 and λ u2 indicates different frequency bands. While shown as passing two frequency bands in this figure, each multi-bandpass filter of the present disclosure can pass more than one frequency band at a time. For example, as described below in connection with FIGS. 11A and 11B, in some implementations, each multi-bandpass filter can pass four frequency bands. In various embodiments, more frequency bands may be passed. For example, an implementation using four cameras may include a multi-bandpass filter configured to pass eight frequency bands. In some embodiments, the number of frequency bands may be, for example, four, five, six, seven, eight, nine, ten, twelve, fifteen, sixteen, or more.
[0091] The multiple bandpass filters 205A, 205B have curvatures selected to reduce the angle-of-incidence-dependent spectral transmission of their respective sensor regions 225A, 225B. As a result, when receiving narrowband illumination from object space, each photodiode (e.g., a color filter covering the sensor region that passes that wavelength) located across the entire surface of the sensor region 225A, 225B that is sensitive to that wavelength will receive substantially the same wavelength of light, without the wavelength shift seen in photodiodes located near the edges of the sensor, as discussed above in connection with FIG. 1A. This configuration can produce more accurate spectral image data than would be possible using flat filters.
[0092] Figure 3B shows an example of an optical design for the optical components that make up one optical path of the multi-aperture imaging system shown in Figure 3A. More specifically, Figure 3B shows a custom achromatic doublet 240 that can be used as multiple bandpass filters 205A, 205B. This custom achromatic doublet 240 passes light through a housing 250 to an image sensor 225. The housing 250 may include apertures 210A, 210B and imaging lenses 215A, 215B, as described above.
[0093] The achromatic doublet 240 is configured to correct optical aberrations by incorporating the required surfaces for the multi-bandpass filter coatings 205A, 205B. The illustrated achromatic doublet 240 includes two lenses, which can be made from glass or other optical materials with different dispersions and refractive indices. In other implementations, three or more lenses may be used. These achromatic doublet lens designs incorporate the multi-bandpass filter coatings 205A, 205B on a curved front surface 242, and incorporate curved single-lens optical surfaces onto which the filter coatings 205A, 205B are deposited to cancel optical aberrations. The combined effect of the curved front surface 242 and curved back surface 244 of the achromatic doublet 240 limits the refractive or light-gathering power, allowing these lenses alone, housed in the housing 250, to constitute the primary light-gathering element. The achromatic doublet 240 can thus contribute to improving the accuracy of image data captured by the imaging system 200. These individual lenses can be mounted adjacent to one another, e.g., bonded or cemented, so that the aberrations of one lens are canceled by the aberrations of the other lens. The curved front surface 242 or the curved back surface 244 of the achromatic doublet 240 can be coated with a multi-bandpass filter coating 205A, 205B. Other doublet designs may be implemented in the systems described herein.
[0094] Further variations on the optical designs described herein may be implemented. For example, in some embodiments, a single lens or other optical singlet, such as the convex (positive) or concave (negative) lens shown in FIG. 3A, may be included in the optical path instead of the doublet 240 shown in FIG. 3B. FIG. 3C illustrates an example implementation in which a flat filter 252 is positioned between the lens housing 250 and the sensor 225. The achromatic doublet 240 of FIG. 3C, incorporating a flat filter 252 with multi-band transmission characteristics, corrects optical aberrations without significantly contributing to the optical power of each lens housed in the housing 250. FIG. 3D illustrates another example implementation in which a multi-bandpass coating is implemented by applying a multi-bandpass coating 254 to the front surface of the lens assembly housed within the housing 250. In this manner, the multi-bandpass coating 254 may be applied to any curved surface of any optical element housed within the housing 250.
[0095] FIGS. 4A-4E illustrate one embodiment of a multispectral, multi-aperture imaging system 300 incorporating the optical design described in connection with FIGS. 3A and 3B. More specifically, FIG. 4A illustrates a perspective view of imaging system 300, including a semi-transparent housing 305 to reveal the internal components. The actual housing 305 may be larger or smaller than the illustrated housing 305, depending, for example, on the desired embedded computing resource. FIG. 4B illustrates a front view of imaging system 300. FIG. 4C illustrates a cutaway side view of imaging system 300 taken along line CC in FIG. 4B. FIG. 4D illustrates a bottom view of imaging system 300, showing processing board 335. The following describes FIGS. 4A-4D.
[0096] The housing 305 of the imaging system 300 may be housed within another housing. For example, in a portable implementation, the imaging system 300 may be housed within a housing that may include one or more molded handles to allow the imaging system 300 to be held securely. An example of a portable implementation is shown in more detail in FIGS. 18A-18C and 19A-19B. The top surface of the housing 305 includes four openings 320A-320D. A different multi-bandpass filter 325A-325D is positioned above each opening 320A-320D and held in place by a filter cap 330A-330B. The multi-bandpass filters 325A-325D may be curved or non-curved, and as described herein, each multi-bandpass filter passes a common frequency band and at least one unique frequency band, thereby enabling high-precision multispectral imaging across more spectral channels than would typically be captured by an image sensor covered with a color filter array. The image sensors, imaging lenses, and color filters described above are disposed within camera housings 345A-345D. In some embodiments, the image sensors, imaging lenses, and color filters described above may be housed within a single camera housing, as shown, for example, in FIGS. 20A-20B. While the illustrated implementation uses separate sensors (e.g., one sensor housed within each of camera housings 345A-345D), it will be appreciated that other implementations may use a single image sensor disposed across the entire area exposed through openings 320A-320D. In this embodiment, camera housings 345A-345D are secured to imaging system housing 305 using support 340, although other support materials may be used in various implementations.
[0097] The top surface of the housing 305 supports an optional illumination board 310, which is covered with an optical diffusing element 315. The illumination board 310 is described in more detail in connection with FIG. 4E below. The diffusing element 315, which may be constructed of glass, plastic, or other optical materials, diffuses the light emitted from the illumination board 310, ensuring that the object space receives substantially uniform spatial illumination. Uniform illumination of the target object can provide a substantially uniform amount of illumination across the object's surface at each wavelength, which can be beneficial in certain imaging applications, such as clinical analysis of tissue by imaging. In some embodiments, the imaging systems disclosed herein may utilize ambient light instead of or in addition to the light from the optional illumination board.
[0098] Because the illumination board 310 generates heat during use, the imaging system 300 includes a heat sink 350 with a number of heat dissipation fins 355. The heat dissipation fins 355 may extend into the spaces between each of the camera housings 345A-345D, and the top of the heat sink 350 can conduct heat from the illumination board 310 to the fins 355. The heat sink 350 may be made from a suitable thermally conductive material. The heat sink 350 may also assist in dissipating heat from other components, thereby eliminating the need for a fan in some implementations of the imaging system.
[0099] A processing board 335, which communicates with cameras 345A-345D, is secured within housing 305 by multiple supports 365. Processing board 335 can control the operation of imaging system 300. While not shown, imaging system 300 can also be configured with one or more memories, such as memory for storing data generated using the imaging system, and / or a computer-executable instruction module for system control. Processing board 335 can be configured in various ways depending on the system design objectives. For example, a processing board can be configured to control the activation of specific LEDs on illumination board 310 (e.g., via a computer-executable instruction module). In some implementations, a highly stable synchronous buck LED driver can be used, which allows software-based analog current control of the LEDs and LED fault detection. In some implementations, image data analysis functionality can be further provided on processing board 335 or a separate processing board (e.g., via a computer-executable instruction module). Although not shown, imaging system 300 may include a data interconnect between the sensor and processing board 335 so that processing board 335 can receive data from the sensor and process that data, and may include a data interconnect between illumination board 310 and processing board 335 so that the processing board can activate specific LEDs on illumination board 310.
[0100] FIG. 4E illustrates an example of an illumination board 310 that may be included in imaging system 300, with the illumination board 310 separated from other components. The illumination board 310 has four arms extending from a central portion, with three columns of LEDs arranged along each arm. The LEDs in adjacent columns are laterally offset from each other to space adjacent LEDs apart. Each LED column includes multiple rows of LEDs of various colors. Four green LEDs 371 are arranged in the center, with one green LED located at each corner of the center. Starting from the innermost row (e.g., the row closest to the center), each column includes a row of two crimson LEDs 372 (for a total of eight crimson LEDs). Continuing radially outward, each arm has a center column with a row containing one amber LED 374, two rows on either side with two short-blinking blue LEDs 376 (for a total of eight short-blinking blue LEDs), another center column with one amber LED 374 (for a total of eight amber LEDs), two rows on either side with one non-PPG NIR LED 373 and one red LED 375 (for a total of four non-PPG NIR LEDs and four red LEDs), and one center column with one PPG NIR LED 377 (for a total of four PPG NIR LEDs). "PPG" LEDs refer to LEDs that are activated in multiple sequential exposures to capture photoplethysmography (PPG) information representative of pulsatile blood flow in living tissue. It will be appreciated that various other colors and / or arrangements of LEDs may be used in alternative embodiments on the illumination board.
[0101] Figure 5 illustrates another embodiment of a multispectral, multi-aperture imaging system 300 with the optical design described in connection with Figures 3A and 3B. Similar to the design of imaging system 300, imaging system 400 includes four optical paths, shown here as apertures 420A-420D with multi-bandpass filter lens groups 425A-425D secured to housing 405 by retaining rings 430A-430D. Imaging system 400 further includes an illumination board 410 secured to the front of housing 405 between retaining rings 430A-430D, and a diffuser 415 disposed on illumination board 410 to aid in spatially uniform illumination of the target object.
[0102] The illumination board 410 of the imaging system 400 has LEDs arranged in four arms of a cross, with each arm having two closely spaced rows of LEDs. This configuration may make the illumination board 410 more compact than the illumination board 310 described above and suitable for use with imaging systems having smaller form factor requirements. In this example configuration, each arm has one green LED and one blue LED in the outermost row, and, moving inward, two rows of yellow LEDs, one row of amber LEDs, one row of red LEDs and one crimson LED, and one row of amber LEDs and one NIR LED. Thus, in this implementation, the LEDs are arranged so that the LEDs emitting longer wavelength light are located in the center of the illumination board 410 and the LEDs emitting shorter wavelength light are located at the edges of the illumination board 410.
[0103] 6A-6C illustrate another embodiment of a multispectral, multi-aperture imaging system 500 incorporating the optical design described in connection with FIGS. 3A and 3B. More specifically, FIG. 6A illustrates a perspective view of imaging system 500, FIG. 6B illustrates a front view of imaging system 500, and FIG. 6C illustrates a cutaway side view of imaging system 500 taken along line CC in FIG. 6B. Imaging system 500 includes similar components to those previously described in connection with imaging system 300 (e.g., housing 505, illumination board 510, diffuser 515, multiple bandpass filters 525A-525D secured above apertures by retaining rings 530A-530D, etc.), but in a shorter form factor (e.g., in one embodiment, fewer embedded computer components and / or smaller dimensions of the embedded computer components). Imaging system 500 further includes a direct camera-to-frame mount 540 for providing a rigid and secure camera alignment position.
[0104] 7A-7B illustrate another embodiment of a multispectral, multi-aperture imaging system 600. FIGS. 7A-7B illustrate another possible arrangement of ambient light sources 610A-610C in the multi-aperture imaging system 600. As shown, four lens assemblies with multi-bandpass filters 625A-625D, each with the optical design described in connection with FIGS. 3A-3D, can be arranged in a rectangular or square configuration to receive light from four cameras 630A-630D (including image sensors). Three rectangular light-emitting elements 610A-610C can be positioned parallel to each other outside and between each lens assembly with multi-bandpass filters 625A-625D. These rectangular light-emitting elements can be broad-spectrum light-emitting panels, each containing LEDs emitting light in different frequency bands.
[0105] 8A-8B illustrate another embodiment of a multispectral, multi-aperture imaging system 700. FIGS. 8A-8B illustrate another possible arrangement of ambient light sources 710A-710D in the multi-aperture imaging system 700. As shown, four lens assemblies with multi-bandpass filters 725A-725D employing the optical design described in connection with FIGS. 3A-3D can be arranged in a rectangular or square configuration to receive light from four cameras 730A-730D (including image sensors). The four cameras 730A-730D are shown as an example of a closer configuration, which may minimize the perspective difference between the lenses. Four rectangular light emitting elements 710A-710D can be arranged in a square configuration surrounding the lens assemblies with multi-bandpass filters 725A-725D. These rectangular light emitting elements may be broad-spectrum light emitting panels, each containing LEDs emitting light in a different frequency band.
[0106] Figures 9A-9C illustrate another embodiment of a multispectral, multi-aperture imaging system 800. The imaging system 800 includes a frame 805 connected to a lens assembly front frame 830, which includes an aperture 820 and a support structure for a micro-video lens 825, which may include multiple bandpass filters using the optical design described with respect to Figures 3A-3D. The micro-video lens 825 receives light from four cameras 845 (including imaging lenses and image sensor areas) mounted on a lens assembly rear frame 840. Four linear rows of LEDs 811 are arranged along the four sides of the lens assembly front frame 830, with a diffusing element 815 for each LED on each side. Figures 9B and 9C illustrate possible dimensions, in inches, for the multi-aperture imaging system 800.
[0107] FIG. 10A illustrates another embodiment of a multispectral, multi-aperture imaging system 900 incorporating the optical design described in connection with FIGS. 3A-3D. The imaging system 900 can be implemented as a set of multiple bandpass filters 905 that can be attached to a multi-aperture camera 915 of a mobile device 910. For example, a particular mobile device 910, such as a smartphone, can be equipped with a stereoscopic imaging system having two apertures leading to two image sensor areas. The multi-aperture spectral imaging techniques disclosed herein can be implemented in such a device by providing a suitable set of multiple bandpass filters 905 that pass multiple narrow frequency bands of light to the sensor areas. The set of multiple bandpass filters 905 can also be equipped with a light source (e.g., an LED array and a diffuser) that irradiates the object space with light in these frequency bands.
[0108] The imaging system 900 may further include a mobile application that configures the mobile terminal to generate and process the multispectral datacube (e.g., in a clinical tissue classification application, a biometric application, a material analysis application, or other application). Alternatively, the mobile application may configure the mobile terminal 910 to transmit the multispectral datacube over a network to a remote processing system, and then receive and display the analysis results. An example of a user interface 910 for such an application is shown in FIG. 10B.
[0109] 11A-11B show an example of a set of frequency bands that can pass through four filters when implemented in the multispectral, multi-aperture imaging system shown in FIGS. 3A-10B, and these set of frequency bands are incident on an image sensor, for example, equipped with a Bayer CFA (or another RGB CFA or RGB-IR CFA). These transmission curves also include the effect of quantum efficiency of the sensors used in this example. As shown, the set of four cameras captures eight unique channels, or eight unique frequency bands. Each filter passes two common frequency bands (the two leftmost peaks) and two other frequency bands for each camera. In this implementation, the first and third cameras receive light in the first shared NIR frequency band (the rightmost peak), and the second and fourth cameras receive light in the second shared NIR frequency band (the second peak from the right). Each camera receives one unique frequency band, ranging from approximately 550 nm to approximately 800 nm or from 550 nm to 800 nm. Thus, these cameras can capture eight unique spectral channels using a compact configuration. Graph 1010 in FIG. 11B shows the spectral irradiance of the LED board shown in FIG. 4E, which may be used to illuminate the four cameras shown in FIG. 11A.
[0110] In this implementation, eight frequency bands were selected based on generating spectral channels suitable for clinical tissue classification, but may be optimized for signal-to-noise ratio (SNR) and frame rate while limiting the number of LEDs (which transmit heat within the imaging system). Because tissue (e.g., animal tissue, such as human tissue) exhibits higher contrast in blue wavelengths than in green or red wavelengths, the eight frequency bands include a common frequency band of blue light (the left-most peak in graph 1000) that passes through all four filters. More specifically, as shown in graph 1000, human tissue exhibits the highest contrast when imaged with a frequency band centered around 420 nm. Because the channel corresponding to this common frequency band is used for disparity correction, higher contrast allows for more accurate correction. For example, when performing disparity correction, the image processor can use local or global methods to find a set of disparities that maximizes a figure of merit indicating similarity between local image patches or images. Alternatively, the image processor can use a similar method to minimize a figure of merit indicating dissimilarity. These figures of merit may be based on entropy, correlation, absolute difference, or deep learning methods. Global disparity calculations can be performed iteratively and terminated when the figure of merit stabilizes. Local methods can calculate disparity pixel-by-pixel, using a fixed patch in one image as the figure of merit input and evaluating different disparity values for multiple patches extracted from other images during testing. All of these methods may constrain the range of disparities considered. This constraint may be based, for example, on knowledge of the depth and distance of the object. It may also be based on the range of expected tilts of the object. The calculated disparity may be constrained by projective geometry, such as an epipolar constraint. Disparities at various resolutions can be calculated using the output of disparities calculated at lower resolutions as the initial value or constraint for the disparity calculation at the next level of resolution.For example, a disparity calculated at a 4-pixel resolution in a first calculation can be used to set a constraint of ±4 pixels in subsequent disparity calculations at higher resolutions. Any algorithm that calculates from disparity benefits from higher contrast, especially when that contrast source is relevant for all viewpoints. Typically, a common frequency band can be selected to accommodate the highest contrast imaging of the materials expected to be imaged in a particular application.
[0111] Because color separation between adjacent channels may not be perfect after image capture, this implementation also includes an additional common frequency band—a green frequency band adjacent to the blue frequency band—that passes through all of the filters shown in graph 1000. The reason for using such a frequency band is that blue color filter pixels have a wide spectral passband and are therefore sensitive to the green spectral region. This is typically referred to as spectral overlap and can be considered intentional crosstalk between adjacent RGB pixels. This spectral overlap allows the color camera's spectral sensitivity to approach that of the human retina, making the resulting color space qualitatively similar to human vision. Therefore, the common green channel across each filter allows for the separation of the portion of the signal generated by the blue photodiode that truly corresponds to the received blue light. This separation can be achieved using a spectral unmixing algorithm that factors in the transmittance of the multi-bandpass filters (shown as the solid black lines in the T legend) and the corresponding transmittance of the CFA color filters (shown as the dashed red, green, and blue lines in the Q legend, respectively). It will be appreciated that in some implementations, red light may be used as the common frequency band, in which case a second common channel may not be required.
[0112] 12 shows a high-level block diagram of an example compact imaging system 1100 with high-resolution spectral imaging capabilities, the imaging system 1100 comprising a set of components including a processor 1120 coupled to a multi-aperture spectral camera 1160 and a light source 1165. A working memory 1105, a storage device 1110, an electronic display 1125, and a memory 1130 are also in communication with the processor 1120. As described herein, the imaging system 1100 may be able to capture more image channels by placing multiple, different bandpass filters over the apertures of the multi-aperture spectral camera 1160 than would be possible with different color filters placed on the CFA of the image sensor.
[0113] The compact imaging system 1100 may be a device such as a mobile phone, a digital camera, a tablet computer, or a personal digital assistant. The compact imaging system 1100 may also be a fixed device such as a desktop computer or a videoconferencing station that uses an internal or external camera for capturing images. The compact imaging system 1100 may be a combination of an image capture device and a separate processing device that receives image data from the image capture device. A user may have access to multiple applications on the compact imaging system 1100. These applications may include traditional photography applications, applications for still and video capture, dynamic color correction applications, and brightness and tint correction applications.
[0114] The image capture system 1100 includes a multi-aperture spectral camera 1160 for capturing images. The multi-aperture spectral camera 1160 may be, for example, any of the devices shown in FIGS. 3A-10B. The multi-aperture spectral camera 1160 may be connected to the processor 1120 so that images captured in different spectral channels from different sensor regions can be transmitted to the image processor 1120. Additionally, as described in more detail below, the processor may control one or more light sources 1165 to emit light of specific wavelengths during specific exposures. The image processor 1120 may be configured to perform various processes on the received captured images so as to output a high-quality, parallax-corrected multispectral data cube.
[0115] The processor 1120 may be a general-purpose processing unit or a processor specifically designed for imaging applications. As shown, the processor 1120 is coupled to a memory 1130 and a working memory 1105. In the illustrated embodiment, the memory 1130 stores an acquisition control module 1135, a data cube creation module 1140, a data cube analysis module 1145, and an operating system 1150. These modules contain instructions that configure the processor to perform various image processing and device management tasks. The processor 1120 uses the working memory 1105 to store working sets of processor instructions contained in each module of the memory 1130. Alternatively, the processor 1120 uses the working memory 1105 to store dynamic data created during operation of the image acquisition system 1100.
[0116] As described above, the processor 1120 is configured with several modules stored in the memory 1130. In some implementations, the imaging control module 1135 includes instructions for configuring the processor 1120 to adjust the focus position of the multi-aperture spectral camera 1160. The imaging control module 1135 further includes instructions for configuring the processor 1120 to capture images, such as multispectral images captured in different spectral channels and PPG images captured in the same spectral channel (e.g., the NIR channel), with the multi-aperture spectral camera 1160. Non-contact PPG imaging typically uses near-infrared (NIR) wavelengths for illumination, taking advantage of the increased photon penetration into tissue at these wavelengths. Thus, the imaging control module 1135, the multi-aperture spectral camera 1160, and the processor 1120 with the working memory 1105 are one means for capturing a series of spectral images and / or a sequence of images.
[0117] The data cube creation module 1140 includes instructions that configure the processor 1120 to generate a multispectral data cube based on the intensity signals received from each photodiode in different sensor regions. For example, the data cube creation module 1140 can estimate the disparity between the same region of the imaged object based on the spectral channels corresponding to the common frequency bands passed by all the multiple bandpass filters, and use this disparity to align all the spectral images captured in all channels with each other (e.g., alignment can be performed so that the same point on a particular object is at substantially the same pixel location (x, y) in all spectral channels). The aligned images can be combined to form a multispectral data cube, and the disparity information can be used to measure the depth of the separately imaged objects, for example, the difference between the depth of normal tissue and the depth of the deepest point of a wound site. In some embodiments, the data cube creation module 1140 can further perform spectral unmixing to identify which portions of the photodiode intensity signals correspond to which frequency bands passed by the filters, for example, based on a spectral unmixing algorithm that factors in the filter transmittance and the quantum efficiency of the sensor.
[0118] The data cube analysis module 1145 can implement various techniques to analyze the multispectral data cube generated by the data cube creation module 1140 for various applications. For example, some implementations of the data cube analysis module 1145 can provide the multispectral data cube (which may include depth information) to a machine learning model trained to classify each pixel according to a specific condition. This specific condition can be a clinical condition for tissue imaging, such as burn status (e.g., first-degree burn, second-degree burn, third-degree burn, or normal tissue type), wound status (e.g., hemostasis, inflammation, proliferation, remodeling, or normal skin type), healing potential (e.g., a score reflecting the likelihood that the tissue will heal from the wound state with or without a specific treatment), perfusion status, cancer status, or other tissue condition associated with the wound. The data cube analysis module 1145 can also analyze the multispectral data cube for biometric and / or material analysis purposes.
[0119] The operating system module 1150 configures the processor 1120 to manage the memory and processing resources of the image capture system 1100. For example, the operating system module 1150 may include device drivers that manage hardware resources such as the electronic display 1125, the storage device 1110, the multi-aperture spectral camera 1160, the light source 1165, etc. Thus, in some embodiments, the instructions included in the aforementioned image processing module may not interact directly with these hardware resources, but instead may interact through standard subroutines or APIs included in the operating system component 1150. The instructions in the operating system 1150 may then interact directly with these hardware components.
[0120] Additionally, the processor 1120 may be configured to control the display 1125 to display the captured images and / or the analysis results of the multispectral datacube (e.g., classified images) to a user. The display 1125 may be external to the imaging device including the multi-aperture spectral camera 1160 or may be part of the imaging device. The display 1125 may also be configured to provide a viewfinder to a user before capturing an image. Additionally, the display 1125 may include an LCD screen or an LED screen and may implement touch-sensitive technology.
[0121] The processor 1120 may write data, such as captured images, multispectral data cubes, and data representing analysis results of the data cubes, to the storage module 1110. While the storage module 1110 is depicted as a conventional disk drive in this figure, those skilled in the art will appreciate that the storage module 1110 may be configured as any type of storage media device. For example, the storage module 1110 may include a disk drive, such as a floppy disk drive, a hard disk drive, an optical disk drive, or a magneto-optical disk drive, or solid-state memory, such as flash memory, RAM, ROM, and / or EEPROM. The storage module 1110 may further include multiple memory units, any one of which may be configured to be located internal to the image capture device 1100 or external to the image capture system 1100. For example, the storage module 1110 may include a ROM memory containing system program instructions stored within the image capture system 1100. Additionally, the storage module 1110 may include a memory card or high-speed memory removable from the camera and configured to store captured images.
[0122] 12 shows a system including separate components, such as a processor, an image sensor, and a memory, those skilled in the art will recognize that these separate components can be combined in various ways to achieve particular design goals. For example, in another embodiment, memory components may be combined with processor components to reduce cost and improve performance.
[0123] Furthermore, while FIG. 12 depicts two memory components, a memory component 1130 containing several modules and a separate memory 1105 containing working memory, those skilled in the art will recognize that other memory structures may be utilized in some embodiments. For example, a design may utilize ROM or static RAM memory to store processor instructions implementing each module contained in memory 1130. Alternatively, processor instructions may be read upon system startup from disk storage integrated into image capture system 1100 or connected via an external device port. The processor instructions may then be loaded into RAM to facilitate execution of the instructions by the processor. For example, working memory 1105 may be RAM memory, and instructions may be loaded into working memory 1105 prior to execution by processor 1120.
[0124] Overview of an example of image processing technology FIG. 13 is a flowchart illustrating an example process 1200 for capturing image data using the multispectral, multi-aperture imaging system illustrated in FIGS. 3A-10B and 12. FIG. 13 illustrates four example exposures that can be used to generate the multispectral datacube described herein: a visible light exposure 1205, an additional visible light exposure 1210, a non-visible light exposure 1215, and an ambient light exposure 1220. It will be appreciated that these exposures may be captured in any order, and that some of these exposures may be omitted from or added to certain workflows described below. While process 1200 will be described below with reference to the frequency bands illustrated in FIGS. 11A and 11B, a similar workflow may be implemented using image data generated based on a different set of frequency bands. Additionally, in various embodiments, image acquisition and / or parallax correction may be enhanced by further implementing flat-field correction according to various known flat-field correction techniques.
[0125] To obtain the visible light exposure 1205, control signals can be sent to the illumination board to illuminate the LEDs corresponding to the first five peaks (the five peaks on the left corresponding to visible light in graph 1000 of FIG. 11A ). It may be necessary to allow the light output waveforms of the specific LEDs to stabilize simultaneously (e.g., for 10 milliseconds). The imaging control module 1135 can then initiate an exposure of the four cameras, which can continue for, e.g., about 30 milliseconds. The imaging control module 1135 then stops the exposure and acquires data from the sensor areas (e.g., by transferring raw intensity signals from the photodiodes to working memory 1105 and / or data store 1110). This data may include a common spectral channel used for parallax correction, as described herein.
[0126] To improve the signal-to-noise ratio, in some implementations, an additional visible light exposure 1210 can be taken using the same process as described for visible light exposure 1205. Taking two identical or nearly identical exposures improves the signal-to-noise ratio, allowing for more accurate analysis of the image data. However, in implementations where the signal-to-noise ratio of a single image is acceptable, this second exposure may be omitted. Also, in some implementations, taking two exposures in a common spectral channel may allow for more accurate parallax correction.
[0127] In some implementations, an additional non-visible light exposure 1215 corresponding to NIR or IR light can be captured. For example, the capture control module 1135 can activate two NIR LEDs corresponding to the two NIR channels shown in FIG. 11A . It may be necessary to simultaneously stabilize the light output waveforms of certain LEDs (e.g., for 10 milliseconds). The capture control module 1135 can then initiate an exposure of the four cameras, which can last for, e.g., about 30 milliseconds. The capture control module 1135 then stops the exposure and acquires data from the sensor areas (e.g., by transferring raw intensity signals from the photodiodes to the working memory 1105 and / or data store 1110). Because the shape or position of the object relative to exposures 1205 and 1210 can be assumed to be unchanged, a pre-calculated disparity value can be used to align each NIR channel, eliminating the need for a common frequency band that passes through all sensor areas.
[0128] In some implementations, multiple exposures can be used to capture sequential images to generate PPG data indicative of deformation of a tissue site due to pulsatile blood flow. In some implementations, these PPG exposures can be captured at non-visible wavelengths. While combining PPG data with multispectral data can improve the accuracy of certain medical imaging analyses, capturing PPG data can increase the time required for the image capture process. In some implementations, the increased time required for the image capture process can result in movement of the portable imaging device and / or the object, which can introduce errors. Therefore, in certain implementations, capturing PPG data can be omitted.
[0129] In some implementations, an ambient light exposure 1220 can be captured. In this exposure, all LEDs can be turned off and an image can be captured using ambient lighting (e.g., sunlight or light from other sources). The capture control module 1135 can then initiate an exposure of the four cameras and continue this exposure for a desired period of time (e.g., approximately 30 milliseconds). The capture control module 1135 then stops the exposure and acquires data from the sensor area (e.g., by transferring raw intensity signals from the photodiodes to the working memory 1105 and / or data store 1110). The effects of ambient light can be removed from the multispectral data cube by subtracting the intensity values of the ambient light exposure 1220 from the numerical values of the visible light exposure 1205 (or the visible light exposure 1205 corrected for signal-to-noise ratio using the second exposure 1210) and also from the numerical values of the invisible light exposure 1215. Such processing can improve the accuracy of downstream analysis by separating the portions of the resulting signal that represent light emitted by the light source and light reflected from the object / tissue site. In some implementations, this step may be omitted if sufficient analytical accuracy is achieved using only visible light exposures 1205, 1210 and non-visible light exposures 1215.
[0130] It will be appreciated that the specific exposure times listed above are examples in one implementation, and that in other implementations the exposure times may vary depending on the image sensor, the intensity of the light source, and the object being imaged.
[0131] Figure 14 shows a simplified block diagram of a workflow 1300 for processing image data, such as image data captured using the multispectral, multi-aperture imaging system shown in Figures 3A-10B and 12 and / or using process 1200 shown in Figure 13. Although workflow 1300 shows the output of two RGB sensor areas 1301A, 1301B, the number of sensor areas in workflow 1300 can be increased, and sensor areas can correspond to different CFA color channels.
[0132] The RGB sensor outputs from the two sensor regions 1301A and 1301B are stored in 2D sensor output modules 1305A and 1305B, respectively. The numerical values of these sensor regions are sent to nonlinear mapping modules 1310A and 1310B, which can perform parallax correction by determining the disparity between images captured using a common channel and then applying this determined disparity across all channels to align all spectral images with each other.
[0133] The outputs of these nonlinear mapping modules 1310A and 1310B are then provided to a depth calculation module 1335, which can calculate the depth of a particular region of interest in the image data. For example, depth may represent the distance between the object and the image sensor. In some implementations, multiple depth values can be calculated and compared to measure the depth of the object relative to some other reference than the image sensor. For example, the deepest depth of the wound bed can be measured, as can the depth (deepest, shallowest, or average) of the normal tissue surrounding the wound bed. The deepest depth of the wound can be determined by subtracting the normal tissue depth from the wound bed depth. This depth comparison can also be performed at other locations in the wound bed (e.g., at all or some of the predetermined sampling points), thereby constructing a 3D map of the wound depth at various points (shown in FIG. 14 as z(x,y) where z is the depth value). The larger the disparity, the more computationally intensive the algorithms for calculating depth as described above, but in some embodiments, the larger the disparity, the better the depth calculations may be.
[0134] The outputs of these nonlinear mapping modules 1310A, 1310B are then also provided to a linear equation module 1320, which can process the detected values as a set of linear equations for spectral unmixing. In one implementation, the Moore-Penrose pseudoinverse equation can be used as a function of at least the sensor's quantum efficiency and the filter's transmittance values to calculate the actual spectral values (e.g., the intensity of light of a particular wavelength incident on each image point (x,y)). Spectral unmixing can be used in implementations requiring high accuracy, such as clinical diagnostics and other biological applications. Spectral unmixing can also be applied to provide estimates of photon flux and signal-to-noise ratio.
[0135] Based on the parallax-corrected spectral channel images and spectral unmixing, the workflow 1300 can generate a spectral data cube 1325, such as in the form shown in the figure of F(x,y,λ), where F denotes the light intensity at a particular image location (x,y) at a particular wavelength or frequency band λ.
[0136] FIG. 15 illustrates parallax and parallax correction for processing image data, such as image data captured using the multispectral, multi-aperture imaging system shown in FIGS. 3A-10B and 12 and / or using the process shown in FIG. 13. A first set of images 1410 shows image data of the same physical location on the same object captured by four different sensor regions. As shown, the object locations are not in the same locations in these raw images based on the (x,y) coordinate frame of the photodiode grid of the image sensor region. A second set of images 1420 shows the same object location after parallax correction, where the images share the same (x,y) location in the coordinate frame of the aligned images. It will be appreciated that such alignment may include cropping certain data from edge regions of the images that do not completely overlap each other.
[0137] FIG. 16 illustrates a workflow 1500 for pixel-by-pixel classification of multispectral image data, such as image data captured using the multispectral, multi-aperture imaging system shown in FIGS. 3A-10B and 12 and / or using the process shown in FIG. 13 and then processed according to FIGS. 14 and 15 .
[0138] In block 1510, a multispectral, multi-aperture imaging system 1513 can capture image data representing physical points 1512 on an object 1511. In this example, the object 1511 includes tissue from a patient having a wound. The wound may include a burn, a diabetic ulcer (e.g., a diabetic foot ulcer), a non-diabetic ulcer (e.g., a pressure ulcer or a slow-healing wound), a chronic ulcer, a post-surgical incision, an amputation site (before or after amputation), a cancerous lesion, or damaged tissue. When PPG information is included, the imaging systems disclosed herein provide a method for assessing conditions involving changes in tissue blood flow and pulse rate, including tissue perfusion; cardiovascular health; injuries such as ulcers; peripheral arterial disease; and respiratory health.
[0139] In block 1520, data captured by the multispectral, multi-aperture imaging system 1513 can be processed to generate a multispectral datacube 1525 having multiple different wavelengths 1523, which may include multiple different images (PPG data 1522) of the same wavelength corresponding to different times. For example, the image processor 1120, via the datacube creation module 1140, can be configured to generate the multispectral datacube 1525 according to the workflow 1300. Some implementations may associate depth values at various points in the spatial dimensions, as previously described.
[0140] In block 1530, the multispectral datacube 1525 can be analyzed as input data 1525 to a machine learning model 1532 to generate a classification mapping 1535 of the imaged tissue. The classification mapping allows each pixel in the image data (each pixel representing a specific point on the imaged object 1511, after registration) to be assigned a specific tissue classification or a specific healability score. Different classifications or scores can be indicated in the classified image output using visually distinct colors or patterns. Thus, even if many images of the object 1511 are taken, the output can be a single image of the object (e.g., a regular RGB image) overlaid with multiple visual representations of the pixel-by-pixel classifications.
[0141] In some implementations, the machine learning model 1532 may be an artificial neural network. An artificial neural network is a computational entity inspired by biological neural networks but adapted for implementation by a computer device. While dependencies between inputs and outputs are difficult to detect, artificial neural networks can be used to model complex relationships between inputs and outputs and to find patterns in data. A neural network typically includes an input layer, one or more intermediate ("hidden") layers, and an output layer, each of which contains many nodes. The number of nodes in each layer may vary. A neural network with two or more hidden layers is considered "deep." Nodes in each layer are connected to all or some of the nodes in the next layer, and the weights of these connections are typically learned from data during a training process, e.g., using backpropagation to adjust the neural network's parameters to generate expected outputs from corresponding inputs in labeled training data. Thus, an artificial neural network can be thought of as an adaptive system configured to change its structure (e.g., connections and / or weights) during training based on the information flowing through it, with the hidden layer weights encoding meaningful patterns in the data.
[0142] A fully connected neural network is a neural network in which each node in the input layer is connected to each node in the next layer (the first hidden layer), each node in the first hidden layer is connected to each node in the next hidden layer, and so on until each node in the final hidden layer is connected to each node in the output layer.
[0143] A CNN is a type of artificial neural network that, like the previously described artificial neural networks, is composed of multiple nodes with learnable weights. Unlike the previously described artificial neural networks, each layer of a CNN may have nodes arranged in three dimensions, with width, height, and depth corresponding to the pixel values (e.g., width and height) of a 2x2 array of each video frame and the number of video frames (e.g., depth) in the sequence. Nodes in each layer may be locally connected to a small region (called a receptive field) of the previous layer, which also has width and height. Weights in hidden layers may take the form of convolutional filters applied to the receptive field. In some embodiments, the convolutional filter may be two-dimensional, thus allowing repeated convolution (or image convolution) using the same filter for each frame of the input volume or a specified subset of frames. In other embodiments, the convolutional filter may be three-dimensional, thus spanning the full depth of multiple nodes in the input volume. Each node in each convolutional layer of a CNN shares weights, replicating the convolutional filter of a given layer across the entire width and height of the input volume (e.g., for every frame) can reduce the total number of trainable weights and improve the applicability of the CNN to datasets outside of the training data. Values in certain layers may be pooled to reduce the number of calculations in the next layer (e.g., values representing specific pixels may be passed on to the next layer, while other values are discarded). Furthermore, as the CNN grows in depth, a pooling mask may be reintroduced into the discarded values, returning the number of data points relative to the original dimension. Multiple layers, some of which may be fully connected, can be stacked to form a CNN structure.
[0144] During training, an artificial neural network can be applied to the paired data in the training data to vary parameters and predict the output of a particular pair given the input. For example, the training data can include a multispectral data cube (input) and a labeled classification mapping (expected output), e.g., by a clinician designating wound areas corresponding to particular clinical conditions. The labeled classification mapping can include a healed label (1) or an unhealed label (0) assigned some time after initial wound imaging to confirm actual healing. Other implementations of the machine learning model 1532 can be trained to make other types of predictions, such as predicting the likelihood that a wound will heal and reduce by a certain percentage within a certain time period (e.g., at least a 50% reduction in wound area within 30 days), or predict the state of the wound, such as hemostasis, inflammation, pathogen colonization, proliferation, remodeling, or normal skin. In some implementations, patient measurements may be further incorporated into the input data to further improve classification accuracy, or the training data may be segmented based on patient measurements and used on other patients with the same measurements to train the machine learning model 1532 on other examples. The patient measurements may include textual information or medical history or aspects thereof that describe characteristics of the patient or the patient's health, such as the area of the wound, lesion, or ulcer; the patient's BMI; the patient's diabetic status; the presence of peripheral vascular disease or chronic inflammation in the patient; the number of other wounds the patient currently has or has had; whether the patient is receiving or has recently received immunosuppressants (e.g., chemotherapy) or other medications that positively or negatively affect wound healing rates; HbA1c; stage IV chronic renal failure; whether the patient has type 1 or type 2 diabetes; chronic anemia; asthma; medication use; smoking status; diabetic neuropathy; deep vein thrombosis; previous myocardial infarction; transient ischemic attack; and sleep apnea; and any combination thereof.These measurements can be converted into vector representations with appropriate processing, such as vector embeddings of words, vectors of binary values indicating whether a patient meets a particular measurement (e.g., whether or not they have type 1 diabetes), or numerical values indicating the degree of each measurement in the patient.
[0145] In block 1540, a classification map 1535 can be output to the user. In this example, the classification map 1535 uses a first color 1541 to indicate pixels classified according to a first state and a second color 1542 to indicate pixels classified according to a second state. The classification and resulting classification map 1535 may exclude background pixels based on, for example, object recognition, background color identification, and / or depth values. As shown, some implementations of the multispectral, multi-aperture imaging system 1513 can project the classification map 1535 onto the tissue site. This aspect may be particularly beneficial if the classification map includes a visual indication of recommended resection margins and / or resection depth.
[0146] Such methods and systems may assist clinicians and surgeons in the process of managing skin wounds, such as making decisions about burn excision, amputation level, lesion removal, and wound triage. The embodiments described herein can be used to identify and / or classify the severity of pressure ulcers, hyperemia, limb deterioration, Raynaud's phenomenon, scleroderma, chronic wounds, abrasions, lacerations, bleeding, breakthrough injuries, puncture wounds, penetration wounds, skin cancer (basal cell carcinoma, squamous cell carcinoma, malignant melanoma, or actinic keratosis), or any type of tissue change in which the tissue's characteristics and properties differ from normal. The devices described herein can also be used to monitor normal tissue; assist and improve wound treatment methods (e.g., allowing for faster and more sophisticated approaches to determining debridement margins); and evaluate the progress of recovery from wounds or diseases (especially after treatment has been applied). Some embodiments described herein provide devices capable of identifying normal tissue adjacent to damaged tissue, determining the margins and / or depth of resection, monitoring the healing process after implantation of a prosthetic device such as a left ventricular assist device, assessing the viability of tissue or regenerative cell grafts, or monitoring post-operative recovery (particularly after reconstructive surgery). Additionally, embodiments described herein can be used to assess wound changes or the development of normal tissue following wound healing, particularly following the administration of therapeutic agents such as steroids, hepatocyte growth factor, fibroblast growth factor, antibiotics, regenerative cells (e.g., isolated or enriched cell populations including stem cells, endothelial cells, and / or endothelial progenitor cells), etc.
[0147] Overview of an example distributed computing environment Figure 17 shows a simplified block diagram of an example of a distributed computing system 1600 including a multispectral, multi-aperture imaging system 1605, which may be any of the multispectral, multi-aperture imaging systems shown in Figures 3A-10B and 12. As shown, a data cube analysis server 1615 may include one or more computers, perhaps arranged as a server cluster or server farm, whose memory and processors may be located within a single computer or may be distributed across many computers (including computers located remotely from each other).
[0148] The multispectral multi-aperture imaging system 1605 may include networking hardware (e.g., wireless internet, satellite communication, Bluetooth, or other communication) for communicating with user devices 1620 and the data cube analysis server 1615 via the network 1610. For example, in some implementations, the processor of the multispectral multi-aperture imaging system 1605 may be configured to control image capture and then transmit the raw data to the data cube analysis server 1615. In other implementations, the processor of the multispectral multi-aperture imaging system 1605 may be configured to control image capture, perform spectral unmixing and parallax correction to generate a multispectral data cube, and then transmit the multispectral data cube to the data cube analysis server 1615. In some implementations, all processing and analysis may be performed locally on the multispectral multi-aperture imaging system 1605, and the multispectral data cube and resulting analysis results may be transmitted to the data cube analysis server 1615 for performing aggregate analysis and / or for use in training or retraining machine learning models. Thus, the data cube analysis server 1615 may provide the latest machine learning models to the multispectral and multiaperture imaging system 1605. The processing load for the final results of analyzing the multispectral data cube may be shared in various ways between the multiaperture imaging system 1605 and the data cube analysis server 1615 depending on the processing power of the multiaperture imaging system 1605.
[0149] The network 1610 may include any suitable network, such as an intranet, the Internet, a cellular network, a local area network, other similar networks, or combinations thereof. The user device 1620 may include any network-equipped computing device, such as a desktop computer, a laptop, a smartphone, a tablet, an e-reader, a video game console, etc. For example, when tissue classification is required, the results (e.g., classified images) measured by the multi-aperture imaging system 1605 and the data cube analysis server 1615 may be transmitted to a patient's or physician's designated user device, a hospital information system storing the patient's electronic medical record, and / or a centralized public health database (e.g., maintained by the U.S. Centers for Disease Control and Prevention).
[0150] An example of the implementation result background: Burn morbidity and mortality pose a challenge for wounded combatants and their medical personnel. Historically, burn injury rates among combat casualties ranged from 5% to 20%, with approximately 20% of these casualties requiring complex burn surgery, such as at the U.S. Army Institute of Surgical Research (ISR) Burn Center. Because burn surgery requires specialized training, it is typically performed by ISR staff rather than U.S. Army hospital staff. The limited number of burn specialists significantly complicates the logistics of providing medical care to burned soldiers. Therefore, utilizing novel, objective methods for detecting burn depth preoperatively and intraoperatively could significantly increase the number of medical personnel (including those outside of the ISR) available to provide medical care to burned patients during ongoing combat. This increased availability of medical personnel could expand the scope of more complex burn care for advancing medical roles for burned combatants.
[0151] To begin to address this need, we developed a novel cart-based imaging device that uses multispectral imaging (MSI) and artificial intelligence (AI) algorithms to aid in the preoperative assessment of burn healability. The device can acquire images of a large area of tissue (e.g., 5.9 x 7.9 square inches) in a short time (e.g., within 6 seconds, 5 seconds, 4 seconds, 3 seconds, 2 seconds, or 1 second) without the need for injection of imaging contrast agents. This study of the general public demonstrated that the accuracy of the device in assessing burn healability was superior (e.g., reaching 70-80%) to the clinical judgment of burn specialists.
[0152] method: Civilian patients with burns of varying severity were imaged within 72 hours of burn injury and at several time points up to 7 days after burn injury. The exact burn severity in each image was determined using healing assessments over a 3-week period or punch biopsies. The accuracy of the device in identifying and distinguishing healing from non-healing burn tissue in first-, second-, and third-degree burns was analyzed on a pixel-by-pixel basis.
[0153] result: Data was collected from 38 civilian patients, totaling 58 burns and 393 images. The AI algorithm achieved a sensitivity of 87.5% and a specificity of 90.7% in predicting non-healing burn tissue.
[0154] Conclusion: The accuracy of our novel device and its AI algorithm in determining burn wound healing potential was shown to be superior to the clinical judgment of burn specialists. Future studies will focus on redesigning the device to enable portability and evaluating its use in intraoperative settings. Design modifications to enable portability include reducing the device's size to that of a handheld system, increasing the field of view, shortening the acquisition time for a single snapshot, and evaluating the device's use in intraoperative settings using a porcine model. These developments were performed using a benchtop multispectral imaging (MSI) subsystem that demonstrated comparable functionality in basic imaging studies.
[0155] Additional light source for image alignment In various embodiments, one or more additional light sources may be used in conjunction with any of the embodiments disclosed herein to improve the accuracy of image registration. FIG. 21 illustrates an example embodiment of a multi-aperture spectral imager 2100 including a projector 2105. In some embodiments, the projector 2105 or other suitable light source may be, for example, one of the multiple light sources 1165 described above with respect to FIG. 12. In embodiments including an additional light source such as the alignment projector 2105, the method may further include additional exposures. The additional light source, such as the projector 2105, may project one or more dots, stripes, grids, random spots, or other suitable spatial patterns in a single spectral band, multiple spectral bands, or broadbands into the field of view of the imager 2100 that are individually or cumulatively visible through all cameras of the imager 2100. For example, the projector 2105 may project shared or common channel light, broadband illumination, or cumulatively visible illumination that can be used to verify the accuracy of image registration calculated based on the common-band approach described above. As used herein, "cumulatively visible illumination" refers to a selected plurality of wavelengths whose pattern is converted by each image sensor of a multispectral imaging system. For example, cumulatively visible illumination can include a plurality of wavelengths whose pattern is converted by each channel, even if the plurality of wavelengths does not include a wavelength common to all channels. In some embodiments, the type of pattern projected by projector 2105 may be selected based on the number of apertures into which the pattern is imaged.For example, if a pattern is visible through only one aperture, a relatively high density pattern may be preferable (e.g., having a relatively narrow autocorrelation of about 1-10 pixels, about 20 pixels, less than about 50 pixels, less than about 100 pixels, etc.), whereas if the pattern is imaged through multiple apertures, a less dense pattern or a pattern with a less narrow autocorrelation may be useful. In some embodiments, an additional exposure taken with the projected spatial pattern is included in the parallax calculation to improve alignment accuracy over embodiments in which the exposure is taken without the projected spatial pattern. In some embodiments, an additional light source can project a stripe pattern of a single spectral band, multiple spectral bands, or a broadband (e.g., shared or common channel, or broadband illumination) into the field of view of the imager, which can be individually or cumulatively visible through all cameras, to improve image alignment based on the phase of the stripe pattern. In some embodiments, the additional light source can improve image registration by projecting multiple unique spatial arrangements of dots, grids, and / or specks in a single spectral band, multiple spectral bands, or broadband (e.g., shared or common channels, or broadband illumination) into the imager's field of view, which can be individually or cumulatively visible through all cameras. In some embodiments, the method further includes an additional sensor with a single aperture or multiple apertures that can detect the shape of a single or multiple objects within the field of view. For example, the additional sensor may use LIDAR, light field, or ultrasound technology to further improve the accuracy of image registration using the aforementioned common-band approach. The additional sensor may be a single aperture or multi-aperture sensor capable of detecting light field information, or may be capable of detecting other signals such as ultrasound or pulsed laser.
[0156] Implementing machine learning for wound assessment, healing prediction, and treatment
[0003] Exemplary embodiments of machine learning systems and methods for wound assessment, healing prediction, and treatment are described below. The various imaging devices, systems, methods, techniques, and algorithms described herein may all be applied to the field of wound imaging and wound analysis. The implementations described below may include acquiring one or more images of a wound in one or more known wavelength bands, and, based on the one or more images, performing one or more of the following: segmenting the image into wound and non-wound areas; predicting the percentage reduction in wound area after a predetermined period of time; predicting the likelihood of healing of individual areas of the wound after a predetermined period of time; displaying a visual indication related to such segmentation or prediction; indicating whether a standard or advanced wound care therapy should be selected; and other functions.
[0157] In various embodiments, a wound assessment system or clinician can determine the appropriate level of wound care therapy based on the results of the machine learning algorithms disclosed herein. For example, if the output of a wound healing prediction system indicates that the imaged wound will close more than 50% within 30 days, the wound healing prediction system can request or notify a healthcare professional or patient to apply a standard care therapy. If the output of a wound healing prediction system indicates that the imaged wound will not close more than 50% within 30 days, the wound healing prediction system can request or notify a healthcare professional or patient to use one or more advanced wound care therapies.
[0158] In the treatment of existing wounds, wounds such as diabetic foot ulcers (DFUs) may receive one or more standard wound care regimens, such as the standard of care (SOC) defined by the Centers for Medicare and Medicaid, during the initial 30 days of treatment. The standard of care (SOC) may include one or more of the following: optimization of nutritional status; debridement by any means to remove necrotic tissue; maintenance of a clean, moist granulation tissue bed with appropriate moist dressings; treatment necessary for resolution of any infections that may be present; addressing deficiencies in intravascular perfusion to the limb associated with the DFU; reducing pressure load on the DFU; and appropriate glycemic control. During the first 30 days of standard of care (SOC), measurable signs of DFU healing are defined as a decrease in DFU size (wound surface area or wound volume), a decrease in the amount of exudate from the DFU, and a decrease in the amount of necrotic tissue within the DFU. An example of the progression of DFU healing is shown in Figure 22.
[0159] Advanced wound care (AWC) is typically administered when healing is not observed after the first 30 days of standard of care (SOC). The Centers for Medicare and Medicaid have not published a definition or outline of advanced wound care (AWC), but it can be considered any treatment that does not fall within the standard of care (SOC) as defined above. Advanced wound care (AWC) is an area of intense research and innovation, with new options being introduced almost continuously for clinical use. Therefore, coverage for AWC is determined on a case-by-case basis, and some patients may not receive reimbursement for treatments considered to be advanced wound care (AWC). Based on this understanding, advanced wound care (AWC) may include, but is not limited to, one or more of the following: hyperbaric oxygen therapy; negative pressure wound therapy; bioengineered skin substitutes; synthetic growth factors; extracellular matrix proteins; matrix metalloproteinase modulators; and electrical stimulation therapy. An example of the healing progression of a non-healing DFU is shown in FIG.
[0160] In various embodiments, the wound assessment and / or wound healing prediction described herein may be based solely on one or more images of the wound, or on one or more images of the wound in combination with patient health data (e.g., one or more health measurements, clinical characteristics, etc.). The techniques described herein may take a single image or a set of multispectral images (MSIs) of a patient's tissue site, such as an ulcer or other wound, process the images using the machine learning systems described herein, and output one or more healing prediction parameters. A variety of healing parameters may be predicted using the techniques of the present invention. Some examples of healing predictive parameters include: (1) A binary yes / no value regarding whether the ulcer heals to a reduction in ulcer area of more than 50% (or another desired threshold (%) according to clinical standards) within 30 days (or another desired time period according to clinical standards); (2) the percentage of the likelihood that the ulcer will heal to a reduction in ulcer area of more than 50% (or another desired threshold (%) according to clinical standards) within 30 days (or another desired time period according to clinical standards); and (3) A prediction of the actual ulcer area expected to decrease within 30 days (or another desired time period according to clinical standards) due to ulcer healing. These include, but are not limited to: In another example, a system using the present technology may provide binary yes / no values or percentages indicating the likelihood of healing for smaller portions of a wound, such as individual pixels or subsets of pixels in an image of the wound, and these yes / no values or percentages indicating the likelihood of healing may indicate whether the individual portions of the wound are tissue that is likely to heal after a specified period of time, or tissue that is likely not to heal after a specified period of time.
[0161] FIG. 24 illustrates one example of an approach for providing such a healing prediction. As shown, a multispectral imaging sensor is used to capture images of a wound at different wavelengths, at different times, or simultaneously, to provide input and output values to a neural network, such as an autoencoder neural network (a type of artificial neural network, described in more detail below). This type of neural network can reduce the number of features represented by the input, and herein, the number of values (e.g., numerical values) representing pixel values in a single or multiple input images. These features can be provided to, for example, a fully connected feedforward artificial neural network or a machine learning classifier, such as the system illustrated in FIG. 25, to output a healing prediction for the imaged ulcer or other wound.
[0162] FIG. 25 illustrates an example of another approach for providing such a healing prediction. As shown, an image (or a set of multispectral images taken at different wavelengths at different times or simultaneously using a multispectral imaging sensor) is provided as input to a neural network, such as a convolutional neural network ("CNN"). The CNN operates on a two-dimensional ("2D") array of pixel values (e.g., numbers along the height and width of the image sensor used to capture the image data) and provides an output that is a one-dimensional ("1D") representation of the image. The numbers can indicate a classification of each pixel in the input image, for example, a classification of each pixel according to one or more physiological states related to an ulcer or other wound.
[0163] As shown in FIG. 25 , the patient measurement data repository can store other types of information about patients, referred to herein as patient measurements, clinical variables, or health measurements. Patient measurements may include textual information describing patient characteristics, such as ulcer area; the patient's body mass index (BMI); the number of other wounds the patient currently has or has had; diabetic status; whether the patient is currently receiving or has recently received immunosuppressants (e.g., chemotherapy) or other medications that positively or negatively affect wound healing rates; HbA1c; stage IV chronic renal failure; type 1 or type 2 diabetes; chronic anemia; asthma; medication use; smoking status; diabetic neuropathy; deep vein thrombosis; previous myocardial infarction; transient ischemic attack; and sleep apnea; and any combination thereof. However, various other measurements may also be used. Various exemplary measurements are shown in Table 1 below. [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4]
[0164] These measurements can be converted into a vector representation through appropriate processing, such as vector embeddings of words, vectors of binary values indicating whether a patient has a particular measurement (e.g., whether or not the patient has type 1 diabetes), or numerical values indicating the extent of each measurement in the patient. In various embodiments, any one of these patient measurements, or a combination of all or some of these patient measurements, can be used to improve the accuracy of the healing prediction parameters produced by the systems and methods of the present technology. In an exemplary study, image data obtained at the initial clinic visit for treatment of a DFU and analyzed alone without considering clinical variables was found to accurately predict the percent area reduction of the DFU with approximately 67% accuracy. Predictions based solely on patient history were approximately 76% accurate, using the most important features: wound area, BMI, number of previous wounds, HbA1c, stage IV chronic renal failure, type 1 or type 2 diabetes, chronic anemia, asthma, medication use, smoking status, diabetic neuropathy, deep vein thrombosis, previous myocardial infarction, transient ischemic attack, and sleep apnea. When these medical variables were combined with imaging data, the accuracy of predictions increased to approximately 78%.
[0165] In one exemplary embodiment shown in Figure 25, image data represented in one dimension (1D) can be concatenated with patient measurements represented as vectors. This concatenated value can then be provided as an input to a fully connected neural network, which outputs healing prediction parameters.
[0166] The system shown in Figure 25 can be considered a single machine learning system with multiple machine learning models and a vector-represented patient measurement generator. In some embodiments, the entire machine learning system can be trained in an end-to-end manner, such that the CNN and fully connected network can adjust parameters via backpropagation, and healing prediction parameters can be generated from input images by adding the vector-represented patient measurements to the values passed through the CNN and fully connected network.
[0167] An example of a machine learning model Artificial neural networks are artificial in the sense that they are computational entities inspired by biological neural networks but adapted for implementation by computer devices. While dependencies between inputs and outputs are difficult to detect, artificial neural networks can be used to model complex relationships between inputs and outputs and to find patterns in data. Neural networks typically include an input layer, one or more intermediate ("hidden") layers, and an output layer, each of which contains many nodes. The number of nodes in each layer can vary. Neural networks with two or more hidden layers are considered "deep." Nodes in each layer are connected to all or some of the nodes in the next layer, and the weights of these connections are typically learned based on training data during a training process, e.g., using backpropagation to adjust the neural network's parameters to generate expected outputs from corresponding inputs in labeled training data. Thus, an artificial neural network may be an adaptive system configured to change its structure (e.g., connections and / or weights) during training based on information flowing through the artificial neural network, and the weights in the hidden layer can be thought of as encoding meaningful patterns in the data.
[0168] A fully connected neural network is a neural network in which each node in the input layer is connected to each node in the next layer (the first hidden layer), each node in the first hidden layer is connected to each node in the next hidden layer, and so on until each node in the final hidden layer is connected to each node in the output layer.
[0169] An autoencoder is a neural network that includes an encoder and a decoder. A particular autoencoder's goal is to compress input data in an encoder and then restore the encoded data in a decoder, resulting in an output that is a good, complete reconstruction of the original input data. An example of an autoencoder neural network described herein, such as the autoencoder neural network shown in FIG. 24, can use pixel values from an image of a wound (e.g., structured in vector or matrix format) as input to its input layer. One or more subsequent layers, or "encoder layers," encode the input information by reducing its dimensionality (e.g., by representing the input values using fewer dimensions than the n dimensions of the original input values). One or more subsequent hidden layers ("decoder layers") decode the encoded information to generate a feature vector as output at the output layer. An example of a training process for an autoencoder neural network can be unsupervised learning, which learns hidden layer parameters that generate output data consistent with the provided input data. Therefore, the number of nodes in the input layer typically matches the number of nodes in the output layer. Dimensionality reduction allows an autoencoder neural network to learn the most important features of an input image, where the autoencoder's middle layer (or another middle layer) represents "feature-reduced" input values. In some examples, an autoencoder neural network can reduce an image of, say, about 1 million pixels (where each pixel value can be considered a separate feature of the image) to a feature set of about 50 numeric values. The dimensionality reduction method of representing images can also be used in other machine learning models, such as the classifier shown in FIG. 25, a suitable CNN, or other neural network, to output healing prediction parameters.
[0170] A CNN is a type of artificial neural network that, like the previously described artificial neural networks, is composed of multiple nodes with learnable weights between each node. Unlike the previously described artificial neural networks, each layer of a CNN may have nodes arranged in three dimensions, with width, height, and depth corresponding to the pixel values (e.g., width and height) of a 2x2 array of each image frame and the number of image frames (e.g., depth) in an image sequence. In some embodiments, nodes in each layer may be locally connected to a small region (called a receptive field) of the previous layer, which has a width and height. The weights in the hidden layer may take the form of convolutional filters applied to the receptive field. In some embodiments, the convolutional filter may be two-dimensional, thereby allowing the convolution (or convolutional transformation of an image) to be repeated using the same filter for each frame of the input volume or for a specified subset of frames. In other embodiments, the convolutional filter may be three-dimensional, thus spanning the full depth of multiple nodes in the input volume. Each node in each convolutional layer of a CNN shares weights, replicating a given layer's convolutional filter across the entire width and height of the input volume (e.g., for every frame) can reduce the total number of trainable weights and improve the applicability of the CNN to datasets outside the training data. Values in a particular layer may be pooled to reduce the number of calculations in the next layer (e.g., values representing specific pixels may be passed on to the next layer, while other values are discarded). Furthermore, as the CNN's depth increases, a pooling mask may be reintroduced into the discarded values, returning the number of data points relative to the original dimensions. Multiple layers, some of which may be fully connected, can be stacked to form a CNN structure. During training, an artificial neural network can be applied to pairs of data in its training data to vary its parameters and predict the output of a particular pair given an input.
[0171] Artificial intelligence refers to a computer-based system capable of performing tasks normally considered to require human intelligence. The artificial intelligence system disclosed herein is capable of performing image analysis (and other data analysis) that would otherwise require the skill and intelligence of a human physician. Advantageously, the artificial intelligence system disclosed herein can make such predictions at the patient's first visit, rather than requiring the patient to wait 30 days for a wound healing assessment.
[0172] The ability to learn is an important aspect of intelligence, because systems without this ability typically cannot gain knowledge from experience. Machine learning is a field of computer science that imparts learning capabilities to computers without explicit programming, enabling, for example, artificial intelligence systems to learn complex tasks or adapt to changes in their environment. The machine learning system disclosed herein can be trained using large amounts of labeled training data to determine the healing potential of wounds. Through such machine learning, the artificial intelligence system disclosed herein can learn new relationships between the appearance of wounds (captured as image data, such as by MSI) and the healing potential of those wounds.
[0173] The artificial intelligence and machine learning systems disclosed herein include computer hardware, one or more memories, and one or more processors, such as those described with respect to the various imaging systems and devices described herein. Any of the machine learning systems and / or methods according to the present techniques may be implemented on or in communication with the processors and / or memories of the various imaging systems and devices disclosed herein.
[0174] An example of the implementation of multispectral imaging of diabetic foot ulcers (DFUs) In an exemplary application of the machine learning systems and methods disclosed herein, following imaging on day 0, a machine learning algorithm according to the previous description was used to predict the percent area reduction (PAR) of the imaged wounds at day 30. To make this prediction, the machine learning algorithm was trained to take multispectral imaging (MSI) data and clinical variables as inputs and output a scalar value representing the predicted percent area reduction. After 30 days, individual wounds were evaluated to measure the true percent area reduction. The predicted percent area reduction was compared to the measured true percent area reduction in a 30-day assessment of wound healing. The performance of the algorithm was measured using the coefficient of determination (R 2 ) was used to score the score.
[0175] The machine learning algorithm in this application was a bagged ensemble of decision tree classifiers adapted to use data from the DFU image database. Other suitable ensembles of classifiers, such as the XGBoost algorithm, may be implemented similarly. The DFU image database contained 29 individual images of diabetic foot ulcers from 15 subjects. The true percent area reduction at 30 days for each image was known. The algorithm was trained using leave-one-out cross-validation (LOOCV). R 2 The score was calculated after combining the prediction results for the test image obtained from each LOOCV segmentation.
[0176] The MSI data consisted of eight 2D channels, each representing diffusely reflected light from tissue at a specific wavelength filter. Each channel had a 15 cm × 20 cm field of view and a resolution of 1044 × 1408 pixels. The eight wavelength bands included 420 nm ± 20 nm; 525 nm ± 35 nm; 581 nm ± 20 nm; 620 nm ± 20 nm; 660 nm ± 20 nm; 726 nm ± 41 nm; 820 nm ± 20 nm; and 855 nm ± 30 nm, where "±" indicates the half-width of each spectral channel. The eight wavelength bands are shown in Figure 26. Quantitative features were calculated from each channel: the mean, median, and standard deviation of all pixel values.
[0177] Additionally, clinical variables obtained from each subject included age, degree of chronic kidney disease, length of DFU on day 0, and width of DFU on day 0.
[0178] For MSI data cubes with 1 to 8 channels, we created separate algorithms using features extracted from every possible combination of the 8 channels (wavelength bands) (a total of C1(8) + C2(8) + ... + C8(8) = 255 different feature sets). 2 The values were calculated and ordered from smallest to largest. 2 The 95% confidence intervals for R values were calculated from the predicted results of the algorithms trained on each feature set. To determine whether a particular feature set provided an improvement above chance, feature sets were identified that did not include 0.0 within the 95% confidence interval of the results of the algorithm trained on that feature set. The same analysis was performed 255 more times, including all clinical variables for each feature set. Additionally, to determine whether the clinical variables influenced the algorithm's performance, a t-test was used to compare the R values obtained from the 255 algorithms trained using the clinical variables. 2The mean values of were compared with 255 algorithms trained without including clinical variables. The results of the analysis are shown in Tables 2 and 3 below. Table 2 shows the performance of the feature set that does not include clinical variables and includes only image data. [Table 2]
[0179] As shown in Table 2, the best-performing feature set, which did not include clinical features, included only three of the eight channels available in the MSI data. The top five feature sets all included the 726 nm wavelength band. The bottom five feature sets each included only one wavelength band. Although the 726 nm wavelength band was included in each of the top five feature sets, the poorest performance was observed when only this 726 nm wavelength band was used. Table 3 below shows the performance of feature sets that included image data and clinical variables such as age, degree of chronic kidney disease, DFU length on day 0, and DFU width on day 0. [Table 3]
[0180] Among the feature sets that included clinical variables, the best-performing feature set included all eight channels available in the MSI data. All five of the top-performing feature sets included the 855 nm wavelength band. Histograms for the models with and without clinical variables are shown in Figure 27, with vertical lines indicating the mean of each distribution.
[0181] By comparing the importance of clinical features, the R between all feature sets that did not include clinical variables 2 The average value of R between all feature sets, including clinical variables, is 2 The R obtained from a model trained on a feature set that did not include clinical variables was compared to the mean R 2The mean R obtained from models trained on feature sets including clinical variables was 0.31. 2 The mean value of was found to be 0.32. The difference between the means was calculated using a t-test, with a p-value of 0.0443. Therefore, the model trained with clinical variables was found to be significantly more accurate than the model trained without clinical features.
[0182] Extracting features from image data While the application example described extracting pixel mean values, standard deviations, and medians, various other features may be extracted from image data and used to create healing prediction parameters. Feature types include local features, semi-local features, and global features. Local features may indicate the texture of an image patch, while global features may include contour representations, shape descriptors, and texture features. Global texture features and local texture features provide different information about a particular image because they provide different support for texture calculation. In some cases, global features can generalize the entire object with a single vector. On the other hand, local features are calculated at multiple points in the image, allowing for robust analysis even in partially occluded or low-quality images. However, specialized classification algorithms may be required to handle variations in the number of feature vectors per image.
[0183] Local features may include, for example, scale-invariant feature transform (SIFT), speeded-up robust features (SURF), features from accelerated segment test (FAST), binary robust invariant scalable keypoints (BRISK), Harris corner detection operator, binary robust independent elementary features (BRIEF), oriented FAST and rotated BRIEF (ORB), and KAZE features. Semi-local features may include, for example, edges, splines, lines, and moments within a small window. Global features may include, for example, color, Gabor features, wavelet features, Fourier features, texture features (e.g., first, second, and higher moments), neural network features from 1D, 2D, and 3D convolutional or hidden layers, and principal component analysis (PCA).
[0184] Application example in RGB-based DFU imaging As a further example for deriving healing prediction parameters, a similar MSI method may be used based on RGB data obtained from, for example, a photographic digital camera. In this case, the algorithm may obtain data from the RGB image, and may also obtain data from the subject's medical history or other clinical variables, and output a healing prediction parameter such as a conditional probability indicating whether a DFU will respond to a standard wound care regimen over a 30-day period. In some embodiments, this conditional probability is the probability that the DFU being treated is a non-healing ulcer, given that x is the input data for the model characterized by parameters θ, and is given by the following equation: P モデル (y=“non-healing”│x;θ)
[0185] The scoring method for the RGB data may be similar to the scoring method in the MSI application described above. In one example, predicted non-healing areas can be compared with true non-healing areas measured by a 30-day healing assessment of wounds such as DFUs. The results of this comparison indicate the performance of the algorithm. The method for performing such a comparison may be based on the clinical outcome of the output image.
[0186] In this application example, four outcomes can be obtained for each healing prediction parameter generated by the healing prediction algorithm. A true positive (TP) outcome indicates a wound area reduction of less than 50% (e.g., the DFU is a non-healing ulcer) and the algorithm predicts a wound area reduction of less than 50% (e.g., the machine learning device outputs a prediction of non-healing). A true negative (TN) outcome indicates a wound area reduction of at least 50% (e.g., the DFU is a healing ulcer) and the algorithm predicts a wound area reduction of at least 50% (e.g., the machine learning device outputs a prediction of healing). A false positive (FP) outcome indicates a wound area reduction of at least 50%, but the algorithm predicts a wound area reduction of less than 50%. A false negative (FN) outcome indicates a wound area reduction of less than 50%, but the algorithm predicts a wound area reduction of at least 50%. After prediction and evaluation of actual cure, these results can be summarized using the performance measures of accuracy, sensitivity and specificity shown in Table 4 below. [Table 4]
[0187] In a retrospective study, an image database of DFUs was obtained containing 149 individual images of diabetic foot ulcers from 82 subjects. 69% of the DFUs in this dataset were considered "healed" because they had reached the target wound area reduction of 50% by day 30. The mean wound surface area was 3.7 cm. 2 and the median wound area was 0.6 cm 2 It was.
[0188] Color photographic images (RGB images) were used as input data for the developed model. The RGB images consisted of three 2D channels, each representing light diffusely reflected from tissue at wavelengths commonly used with conventional color camera sensors. These images were taken by clinicians using a handheld digital camera. The imager, working distance, and field of view (FOV) selection varied among the images. Before training the algorithm, each image was manually cropped to center the ulcer in the field of view (FOV). After cropping, the images were interpolated to a size of 3 channels x 256 pixels x 256 pixels. The aspect ratio of the original images was not maintained during this interpolation process. However, if desired, the aspect ratio could be maintained during these preprocessing steps. A clinical dataset (e.g., clinical variables and health measurements) was also obtained from each subject, including medical history, previous wounds, and blood test results.
[0189] To perform this analysis, we developed two algorithms with the initial goal of identifying new representations for image data that can be combined with patient health measurements in traditional machine learning classification approaches. Many methods are available for creating these image representations, including principal component analysis (PCA) and scale-invariant feature transform (SIFT). In one example, we performed machine learning to compress images using a separately trained unsupervised approach to predict DFU healing. In a second example, we used an end-to-end supervised approach to predict DFU healing.
[0190] The unsupervised feature extraction approach described above used an autoencoder algorithm consistent with the method shown in Figure 24. An example of an autoencoder is shown schematically in Figure 28. This autoencoder included an encoder module and a decoder module. The encoder module was a 16-layer VGG convolutional network. The 16th layer displayed a compressed image representation. The decoder module was also a 16-layer VGG network with an added upsampling function and removed the pooling function. JPEG0007793018000010.jpg18170
[0191] The autoencoder was pre-trained using PASCAL visual object classes (VOC) data and fine-tuned using the DFU images in this example dataset. Each individual image, containing 3 channels x 256 pixels x 256 pixels (total pixels: 65,536), was aggregated into a single vector per element and compressed into multiple vectors representing 50 data points. After training, the same encoder-decoder algorithm was used for all images in the dataset.
[0192] After extracting the image compression vector, the image compression vector was used as input to a second supervised machine learning algorithm. Combinations of image features and patient features were tested using various machine learning algorithms, including logistic regression, k-nearest neighbors, support vector machines, and various decision tree models. An example of a supervised machine learning algorithm using the image compression vector and patient clinical variables as inputs to predict DFU healing is shown schematically in FIG. 29. The machine learning algorithm may be any of a variety of known machine learning algorithms, such as multilayer perceptron, quadratic discriminant analysis, naive Bayes, or an ensemble of such algorithms.
[0193] As an alternative to the unsupervised feature extraction approach described above, we investigated an end-to-end machine learning approach, as outlined in Figure 30. In this end-to-end approach, we modified a 16-layer VGG CNN by concatenating patient health measurement data with image vectors in the first fully connected layer. In this way, the encoder module and the subsequent machine learning algorithm can be trained simultaneously. Other methods involving global variables (e.g., patient health measurements or clinical variables) have also been proposed to improve CNN performance or repurpose CNNs. The most commonly used method is the feature-wise linear modulation (FiLM) generative model. The supervised machine learning algorithm was trained using k-fold cross-validation. Results for each image were calculated as one of the following: true positive, true negative, false positive, and false negative. These results were summarized using the performance measures listed in Table 4.
[0194] The prediction accuracy of the unsupervised feature extraction (autoencoder) and machine learning approaches shown in Figures 28 and 29 was determined using seven machine learning algorithms and three input feature combinations shown in Figure 31. Each algorithm was trained using three-fold cross-validation and reported as an average accuracy (±95% confidence interval). Of the algorithms trained using this approach, only two algorithms were able to exceed the accuracy baseline. The accuracy baseline was defined as when a simple classifier predicted all DFUs as healing ulcers. The two algorithms that exceeded the baseline were logistic regression and support vector machines, which included a combination of image data and patient data. Patient health measures important for predicting DFU healing and used in these models included wound area; body mass index (BMI); number of previous wounds; hemoglobin A1c (HbA1c); renal failure; type 1 or type 2 diabetes; anemia; asthma; medication use; smoking status; diabetic neuropathy; deep vein thrombosis (DVT); previous myocardial infarction (MI); and combinations thereof.
[0195] The results using the end-to-end machine learning approach shown in Figure 30 showed significantly better performance than the baseline, as shown in Figure 32. However, compared to the unsupervised approach, the performance of this approach was not significantly better, but the average accuracy was higher than the other methods tested.
[0196] Predicting healing of a subset of wound areas In a further exemplary embodiment, the systems and methods of the present technology not only generate a single healing probability for the entire wound, but can also predict the area of tissue in each individual wound that will not heal after 30 days with standard wound care. To achieve this output, a machine learning algorithm was trained to take MSI or RGB data as input and generate healing prediction parameters for a portion of the wound (e.g., individual pixels in an image of the wound, or a subset thereof). The present technology can be further trained to output a visual display, such as an image highlighting areas of ulcerated tissue that are predicted not to heal within 30 days.
[0197] FIG. 33 illustrates an example process for predicting healing and generating a visual display. As described elsewhere herein, a spectral data cube is obtained using the method illustrated in FIG. 33. This data cube is sent to machine learning software for processing. The machine learning software can implement all or part of pre-processing, a machine learning wound assessment model, and post-processing. The machine learning module outputs a conditional probability map that is processed by a post-processing module (e.g., probability thresholding) to generate a visually outputtable result in the form of a classified image for the user. As shown as an image output to the user in FIG. 33, the system of the present invention can display an image of the wound to the user, in which pixels in areas that are likely to heal and pixels in areas that are not likely to heal are displayed with different visual displays.
[0198] The process shown in Figure 33 was applied to a set of DFU images, and the predicted non-healed areas were compared with the true non-healed areas measured by a 30-day healing assessment of the DFU. The results of this comparison demonstrate the performance of the algorithm. The method used to perform this comparison was based on the clinical outcomes of the output images. The DFU image database contained 28 individual images of diabetic foot ulcers from 19 subjects. For each image, the true area of the wound that had not healed after 30 days of standard wound care was known. The algorithm was trained using leave-one-out cross-validation (LOOCV). The result for each image was calculated as one of the following: true positive, true negative, false positive, or false negative. These results were summarized using the performance measures listed in Table 4.
[0199] A convolutional neural network was used to create a conditional probability map for each input image. This algorithm includes an input layer, a convolutional layer, a deconvolutional layer, and an output layer. MSI or RGB data is typically input to the convolutional layer. A convolutional layer typically consists of a convolutional stage (e.g., an affine transform) whose output is used as input to a detection stage (e.g., a nonlinear transform such as a rectified linear unit [ReLU]). The result may be used for the next convolutional and detection stages. The result may be downsampled using a pooling function or used directly as the result of the convolutional layer. The result from the convolutional layer is provided as input to the next layer. A deconvolutional layer typically begins with an inverse pooling layer, followed by a convolutional and detection stage. These layers are typically configured in the following order: input layer, convolutional layer, deconvolutional layer. This configuration is often considered to be an encoder layer first, followed by a decoder layer. The output layer typically consists of multiple fully connected neural networks applied to each vector in any of the dimensions of the tensor output from the previous layer. The results from these fully connected neural networks are aggregated into a matrix called a conditional probability map.
[0200] Each entry in the conditional probability map represents a region in the original DFU image. Each region can be a one-to-one mapping with each pixel in the input MSI image, or an n-to-1 mapping (where n represents an aggregate of several pixels in the original image). The conditional probability value in this map indicates the probability that the tissue in that region of the image will not respond to standard wound care. This result is obtained by segmenting the pixels in the original image, and the predicted non-healing regions are segmented from the predicted healing regions.
[0201] The results from a particular layer in a convolutional neural network can be modified using information from another source. In this example, this modification can be achieved using clinical data (e.g., patient health measurements or clinical variables as described herein) from the subject's medical history or treatment plan as the source of the information. Thus, the results from the convolutional neural network can be adjusted depending on the magnitude of variables other than imaging data. To implement this, a feature-wise linear transformation layer (FiLM) can be incorporated into the network architecture, as shown in FIG. 34. The FiLM layer is a machine learning algorithm trained to learn the parameters of an affine transformation, which is applied to one of multiple layers in the convolutional neural network. The input to this machine learning algorithm is a numeric vector, in this case, a numeric vector representing clinically meaningful patient history in the form of patient health measurements or clinical variables. The training of this machine learning algorithm can be performed simultaneously with the training of the convolutional neural network. One or more FiLM layers using different inputs and machine learning algorithms can be applied to various layers of the convolutional neural network.
[0202] The input data for conditional probability mapping included multispectral imaging (MSI) data and color photographic images (RGB images). The MSI data consisted of eight 2D channels, each representing diffusely reflected light from tissue at a specific wavelength filter. Each channel had a field of view of 15 cm × 20 cm and a resolution of 1044 × 1408 pixels. The eight wavelengths included 420 nm ± 20 nm, 525 nm ± 35 nm, 581 nm ± 20 nm, 620 nm ± 20 nm, 660 nm ± 20 nm, 726 nm ± 41 nm, 820 nm ± 20 nm, and 855 nm ± 30 nm, as shown in Figure 26. The RGB images consisted of three 2D channels, each representing diffusely reflected light from tissue at wavelengths used by conventional color camera sensors. The field of view for each channel was 15 cm x 20 cm, with a resolution of 1044 pixels x 1408 pixels.
[0203] To segment images based on cure probability, they used a CNN architecture called SegNet. According to the developers of SegNet, this model uses RGB images as input and outputs a conditional probability map. They further modified the input layer to accept eight-channel MSI images. They also modified the SegNet architecture to include a FiLM layer.
[0204] To demonstrate the feasibility of segmenting DFU images into healed and non-healed regions, we developed various deep learning models, each using different inputs. These models used two types of input features: only MSI data or only RGB images. In addition to varying input features, we also varied the training aspects of the algorithm. Some of these modifications included pretraining the deep learning model with the PASCAL visual object classes (VOC) dataset, pretraining the deep learning model with a database of images of different types of tissue wounds, prespecifying the kernels in the input layer using a filter bank, early stopping, random image augmentation during algorithm training, and averaging the results of random image augmentation during inference to generate a single aggregated conditional probability map.
[0205] The top two deep learning models using one of the two input features were identified as performing better than chance. Replacing RGB data with MSI data improved results; the number of image-based errors decreased from 9 to 7. However, both the MSI and RGB data methods proved viable for generating conditional probability maps of DFU healing potential.
[0206] In addition to evaluating whether the SegNet architecture can achieve the desired accuracy in segmenting wound images, we also evaluated whether images of other types of wounds, though unexpected, would be suitable for use in a training system to segment DFU images or other wound images based on the conditional healing probability mapping. As previously mentioned, the SegNet CNN architecture can be suitable for segmenting DFU images when pre-trained using DFU image data as training data. However, in some cases, for certain types of wounds, an adequate training image set consisting of a large number of images may not be available. Figure 35 shows an example color image of a DFU (A) and four examples of how different segmentation algorithms segmented the DFU into areas predicted to heal and areas predicted not to heal. In image (A), taken on the first day of assessment, the area of the wound identified as non-healing in the assessment performed four weeks later is indicated by a dashed line. Images (B)–(E), corresponding to different segmentation algorithms, show the predicted non-healing tissue in color. As shown in image (E), the SegNet algorithm, pre-trained on a database of burn images rather than a database of DFU images, was able to highly accurately predict non-healing tissue regions, closely matching the dashed outlines in image (A) corresponding to areas determined to be non-healing by actual examination. In contrast, a naive Bayesian linear model trained on DFU image data (image (B)), a logistic regression model trained on DFU image data (image (C)), and a SegNet pre-trained using PACAL VOC data (image (D)) all performed less well, showing significantly larger and less accurately outlined non-healing tissue regions in images (B)-(D).
[0207] An example of single wavelength analysis of a DFU image In another implementation, it was found that assessment of wound percent area reduction (PAR) at 30 days and / or segmentation in the form of conditional probability maps can also be performed based on single wavelength band image data, without using MSI or RGB image data. To implement this method, a machine learning algorithm was trained to take features extracted from single wavelength band images as input and output a scalar value indicating the predicted percent area reduction.
[0208] All images were obtained from subjects in accordance with an Institutional Review Board (IRB)-approved clinical trial protocol. The dataset included 28 individual images of diabetic foot ulcers from 17 subjects. All subjects were imaged at their first visit for wound treatment. All wounds were at least 1.0 cm wide in their largest dimension. Only subjects who were prescribed standard wound care therapy were included in the study. DFU healing was assessed by clinicians during routine follow-up visits to determine the true percent area reduction (PAR) after 30 days of treatment. During this healing assessment, images of the wound were collected and compared to images taken on day 0 to determine the true percent area reduction (PAR).
[0209] Various machine learning algorithms, such as classifier ensembles, may be used. Two machine learning algorithms for regression were used in this analysis. One algorithm was a bagged ensemble of decision tree classifiers (bagged decision trees), and the other was a random forest ensemble. The features used to train the machine learning regression models were all obtained from pre-treatment DFU images taken at the first visit for treatment of the DFUs included in this study.
[0210] Eight grayscale images of each DFU were acquired at characteristic wavelengths in the visible or near-infrared spectrum. Each image had a field of view of approximately 15 cm x 20 cm and a resolution of 1044 x 1408 pixels. The eight characteristic wavelengths were selected using an optical bandpass filter set with wavelength bands of 420 nm ± 20 nm, 525 nm ± 35 nm, 581 nm ± 20 nm, 620 nm ± 20 nm, 660 nm ± 20 nm, 726 nm ± 41 nm, 820 nm ± 20 nm, and 855 nm ± 30 nm, as shown in Figure 26.
[0211] The raw image, 1044 pixels x 1408 pixels, contained a reflectance intensity value for each pixel. Quantitative features were calculated based on the reflectance intensity values, including the first and second moments (e.g., mean and standard deviation) of the reflectance intensity values. In addition, the median was also calculated.
[0212] After performing these calculations, the filter sets may be individually applied to the raw image to create multiple transformed images. In a specific example, a total of 512 filters with dimensions of 7 pixels by 7 pixels or another suitable kernel size may be used. FIG. 36 shows an example of a set of 512 filter kernels with dimensions of 7 pixels by 7 pixels that may be used in an example implementation. This example (but not limited to) filter set may be obtained through training a convolutional neural network (CNN) for segmentation of a DFU. The 512 filters shown in FIG. 36 were obtained from the first kernel set of the CNN's input layer. These filters were "trained" by updating their weights to avoid large deviations from the Gabor filters included in the filter bank.
[0213] Filters may also be applied to the raw images by convolution. From the 512 images resulting from applying these convolution filters, a single three-dimensional matrix with dimensions of 512 channels by 1044 pixels by 1408 pixels may be constructed. Additional features may also be calculated from this three-dimensional matrix. For example, in some embodiments, the mean, median, and standard deviation of the intensity values of the three-dimensional matrix may be calculated as additional features for input into a machine learning algorithm.
[0214] In addition to the six features mentioned above (e.g., the mean, median, and standard deviation of pixel values of the raw image, and the mean, median, and standard deviation of pixel values of a three-dimensional matrix constructed by applying a convolution filter to the raw image), additional features and / or linear or non-linear combinations of such features may be further included as desired. For example, the product of two features or the ratio of two features may be used as a new input feature to the machine learning algorithm. In one example, the product of the mean and median may be used as an additional input feature.
[0215] The algorithm was trained using leave-one-out cross-validation (LOOCV). One DFU image was set as the test set, and the remaining DFU images were used as the training set. After training, the model was used to predict the percent area reduction of the set of DFU images. After the prediction was complete, the set of DFU images was added back to the full set, and the process was repeated with the other set of DFU images. The LOOCV method was repeated until each DFU image had been used once as the holdout set (test set). After collecting the test set results for each cross-validation fold, the overall performance of the model was calculated.
[0216] The predicted percent area reduction for each DFU image was compared to the measured true percent area reduction for a 30-day assessment of healing of the DFU. The performance of the algorithm was evaluated using the coefficient of determination (R 2) was used to score the results. 2 The usefulness of individual wavelengths was determined by using the R 2 The R value is an index showing the proportion of variance in the DFU area reduction rate (%) explained by features extracted from DFU images. 2 The value is defined by the following formula:
number
number
number
[0217] The goal of this study was to determine whether using each of the eight different wavelengths independently in a regression model would yield significantly better results than chance. To determine whether there was an improvement above chance for a particular feature set, the R of the algorithm trained on that feature set was calculated. 2A feature set was identified that did not include zero within the 95% confidence interval of the original values. To perform this evaluation, we conducted eight independent experiments in which machine learning models were trained using the original six features: the mean, median, and standard deviation of the raw images; and the mean, median, and standard deviation of a three-dimensional matrix created by transforming the raw images using a convolutional filter. Random forest and bagged decision tree models were trained. We reported the algorithms that performed well in cross-validation. The results of these eight models were reexamined to determine whether the lower bound of the 95% confidence interval (CI) was above zero. If the lower bound of the 95% CI was not above zero, additional features created by nonlinear combinations of the original six features were employed.
[0218] Using the initial six features, seven of the eight wavelengths tested could be used to create a regression model that explained a large portion of the variance in the percent area reduction from the DFU dataset. These seven wavelengths, ranked from most effective to least effective, were 660 nm; 620 nm; 726 nm; 855 nm; 525 nm; 581 nm; and 420 nm. When the product of the mean and median values of the three-dimensional matrix was included as an additional feature, the last wavelength listed in the table below (820 nm) was found to be effective. The results of this testing are summarized in Table 5. [Table 5]
[0219] Thus, it has been demonstrated that the imaging and analysis systems and methods described herein appear to be capable of accurately generating one or more healing prediction parameters even when based on images from a single wavelength band. In some embodiments, the use of a single wavelength band may facilitate the calculation of one or more aggregate quantitative features derived from the images, such as the mean, median, or standard deviation of the raw image data, and / or the mean, median, or standard deviation of a set of multiple images or a three-dimensional matrix generated by applying one or more filters to the raw image data.
[0220] An example of a segmentation system and method for wound images As previously described, the machine learning techniques described herein can be used to reliably predict parameters associated with wound healing, such as overall wound healing (e.g., percent area reduction) and / or healing associated with a portion of a wound (e.g., probability of healing associated with an individual pixel or subset of pixels in a wound image), by analyzing spectral images containing reflectance data at individual wavelengths or multiple wavelengths. Furthermore, some of the methods disclosed herein can predict wound healing parameters based at least in part on aggregated quantitative features, e.g., statistical quantities such as mean, standard deviation, median, etc., calculated based on a subset of pixels in a wound image that are determined to correspond to wound tissue regions, i.e., "wound pixels," as opposed to calluses, normal skin, background, or other non-wound tissue regions. Therefore, to improve or optimize the accuracy of such predictions based on a set of wound pixels, it is preferable to accurately select a subset of wound pixels in a wound image.
[0221] Segmentation of an image (e.g., an image of a DFU) into wound and non-wound pixels has traditionally been performed manually, for example, by a physician or other clinician inspecting the image and selecting a set of wound pixels based on the image. However, such manual segmentation is time-consuming, inefficient, and prone to human error. For example, formulas used to calculate area and volume lack the accuracy and precision required to measure the raised portions of a wound. Furthermore, determining the true borders of a wound and classifying tissue within the wound (e.g., epithelial proliferation) require advanced capabilities. Because changes in measurements of a wound are often crucial information for determining the effectiveness of treatment, errors in initial wound measurements can lead to inaccurate treatment decisions.
[0222] To address these issues, systems and methods consistent with the present technology are suitable for automated wound edge detection and tissue type identification in wound regions. In some embodiments, systems and methods consistent with the present technology can be configured to automatically segment a wound image into at least wound pixels and non-wound pixels, with aggregated quantitative features calculated based on a subset of the wound pixels to a desired degree of accuracy. Furthermore, it may be desirable to implement a system or method capable of segmenting a wound image into wound pixels and non-wound pixels and / or further segmenting these wound or non-wound pixels into one or more subclasses without requiring the development of additional healing prediction parameters.
[0223] Color photographs of wounds can be used to develop a dataset of diabetic foot ulcer images. A variety of color camera systems can be used to acquire this data. In one implementation, a total of 349 images were used. A trained physician or other clinician can use a software program to identify and label the wound, callus, normal skin, background, and / or other types of pixels in each wound image. These labeled images, known as a truth mask, may contain multiple colors corresponding to the number of labeling types in the image. Figure 37 shows an example image of a DFU (left) and its corresponding truth mask (right). The example truth mask shown in Figure 37 includes purple areas corresponding to background pixels, yellow areas corresponding to callus pixels, and cyan areas corresponding to wound pixels.
[0224] Based on a set of ground truth images, a convolutional neural network (CNN) can be used to automatically segment the aforementioned tissue types. In some embodiments, the structure of this algorithm may be a shallow U-net with multiple convolutional layers. In one implementation, 31 convolutional layers were used to obtain the desired segmentation results. However, various other algorithms for image segmentation may be applied to obtain the desired output.
[0225] In one example of segmentation implementation, the DFU image database was randomly divided into three sets, and a training set of 269 images was used to train the algorithm, a test set of 40 images was used to select hyperparameters, and a validation set of 40 images was used for validation. The algorithm was trained using the steepest descent method, and the accuracy of the test set was monitored. When the accuracy of the test set reached a maximum, the training of the algorithm was stopped. The validation set was then used to judge the results of the algorithm.
[0226] The results from each image in the validation set obtained by the U-net algorithm were compared to its corresponding true mask. This comparison was performed by comparing pixels. For each of the three tissue types, the comparison results were summarized using the following classification: True positives (TP) included the total number of pixels in the true mask where the target tissue type was present, but the machine learning model predicted the presence of the same tissue type. True negatives (TN) included the total number of pixels in the true mask where the target tissue type was not present, but the machine learning model predicted the presence of the same tissue type. False positives (FP) included the total number of pixels in the true mask where the target tissue type was not present, but the machine learning model predicted the presence of the same tissue type. False negatives (FN) included the total number of pixels in the true mask where the target tissue type was present, but the machine learning model predicted the presence of the same tissue type. These results were summarized using the following measures:
[0227]
number
[0228]
number
[0229]
number
[0230] In some embodiments, the algorithm may be trained for multiple epochs, with the number of epochs determined to optimize accuracy. In one implementation described herein, the image segmentation algorithm was trained for 80 epochs. Training was monitored and the best accuracy on the test dataset was found to be achieved after 73 epochs.
[0231] The performance of the U-net segmentation algorithm was measured by accuracy and found to be better than chance. Furthermore, when predicting only one tissue type using a simple classifier, U-net outperformed all three available simple classifiers. Regardless of the overfitting issue, the performance of this machine learning model on the validation set was shown to be viable based on the metrics used to summarize the results.
[0232] Figure 38 shows three example wound image segmentation results using the U-net segmentation algorithm in combination with the method described herein. In each of the three examples, an image of a DFU is shown in the right column, and the automated image segmentation output generated by a U-net segmentation algorithm trained as described herein is shown in the middle column. The manually constructed truth mask corresponding to each DFU image is shown in the left column of Figure 38. These images visually demonstrate the high accuracy of segmentation using the method described herein.
[0233] term All of the methods and tasks described herein may be performed by a computer system and may be fully automated. In some cases, the computer system includes multiple separate computers or computing devices (e.g., physical servers, workstations, storage arrays, cloud computing resources, etc.) communicating and interoperating over a network to perform the functions described herein. Each such computing device typically includes a processor (or multiple processors) that executes program instructions or program modules stored in memory or other non-transitory computer-readable storage media or non-transitory computer-readable storage devices (e.g., solid-state storage devices, disk drives, etc.). Various functions disclosed herein may be embodied in such program instructions or implemented in the form of application-specific circuitry (e.g., ASICs or FPGAs) in the computer system. When a computer system includes multiple computing devices, these devices may be co-located or located at different locations. Results from the methods and tasks disclosed herein may be persistently stored by converting physical storage devices, such as solid-state memory chips or magnetic disks, into various forms. In some embodiments, the computer system may be a cloud-based computing system in which processing resources are shared by multiple separate business entities or other users.
[0234] The processes disclosed herein may be initiated in response to an event, such as a user- or system administrator-initiated request, a predetermined schedule, a dynamically determined schedule, or some other event. Once initiated, executable program instructions stored on one or more non-transitory computer-readable media (e.g., hard drives, flash memory, removable media, etc.) may be loaded into memory (e.g., RAM) of a server or other computing device. The executable instructions may then be executed by a hardware-based computer processor included in the computing device. In some embodiments, the processes disclosed herein, or portions thereof, may be implemented on multiple computing devices and / or multiple processors, either serially or in parallel.
[0235] Depending on the embodiment, certain acts, events, or functions of any of the processes or algorithms described herein may occur in a different order, may be added to one another, may be combined with one another, or may be omitted altogether (e.g., not all of the acts and events described herein are required to execute an algorithm described herein). Furthermore, in certain embodiments, multiple acts or events may occur simultaneously rather than sequentially, for example, using multithreaded processing, interrupt processing, or multiple processors or processor cores, or other parallel structures.
[0236] The various illustrative logic blocks, modules, routines, and algorithmic steps described in connection with the embodiments disclosed herein can be implemented as electronic devices (e.g., ASIC or FPGA devices), computer software running on computer hardware, or a combination thereof. Furthermore, the various illustrative logic blocks and modules described in connection with the embodiments disclosed herein can be implemented by devices such as processor devices, digital signal processors (“DSPs”), application specific integrated circuits (“ASICs”), field programmable gate arrays (“FPGAs”) or other programmable logic devices, discrete gate or transfer logic, discrete hardware components, or combinations of these designed to perform the functions described herein. A processor device may be a microprocessor, but in alternative aspects, a processor device may be a controller, microcontroller, or state machine, combinations thereof, or the like. A processor device may include electrical circuitry configured to process computer-executable instructions. In another embodiment, a processor device includes an FPGA or other programmable device that performs logical operations without processing computer-executable instructions. The processor unit may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, a combination of multiple microprocessors, a combination of a DSP core and one or more microprocessors, or other such configurations. While described herein primarily with respect to digital technology, the processor unit may also include primarily analog components. For example, all or some of the rendering techniques described herein may be implemented using analog circuitry, or a combination of analog and digital circuitry.A computing environment can include any type of computer system, including, but not limited to, microprocessor-based computer systems, mainframe computers, digital signal processors, portable computing devices, device controllers, and computer engines within appliances, to name a few.
[0237] Components of the methods, processes, routines, or algorithms described in connection with the embodiments disclosed herein may be embodied directly in hardware, in software modules executed by a processor unit, or in a combination of the hardware and software modules. The software modules may be included in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or other form of non-transitory computer-readable storage medium. Exemplary storage media may be coupled to the processor unit such that the processor unit can read information from, or write information to, such storage medium. In this aspect, the storage medium may be integral to the processor unit. The processor unit and the storage medium may be included in an ASIC. The ASIC may be included in a user terminal. In this aspect, the processor unit and the storage medium may reside as discrete components in the user terminal.
[0238] As used herein, conditional terms such as "can," "may," "could," "may," "for example," and the like generally mean that a particular embodiment includes a particular feature, component, or step, and that another embodiment does not include that feature, component, or step, unless otherwise specified or interpreted differently based on the context in which the terms are used. Thus, such conditional terms generally do not imply that one or more embodiments somehow require that feature, component, or step, or that one or more embodiments necessarily include decision logic, regardless of whether a particular feature, component, or step is included or performed in a particular embodiment, in the presence or absence of other inputs or promptings. Additionally, terms such as "including," "comprising," and "having" are synonymous and are used in an open-ended, inclusive sense and do not exclude additional components, features, acts, operations, etc. Furthermore, the term "or" is also used in an inclusive (rather than exclusive) sense; for example, when used to list elements, "or" can mean one, some, or all of the listed elements.
[0239] Unless otherwise indicated, it is understood that disjunctive language such as "at least one of X, Y, or Z" is generally used herein to indicate that a particular item, term, etc. may be X, Y, or Z, or any combination thereof (e.g., X, Y, or Z). Thus, such disjunctive language is generally not intended, and should not be intended, to imply that at least one of X, at least one of Y, and at least one of Z are each required to be present in a particular embodiment.
[0240] While the foregoing detailed description has shown, described, and pointed out novel features applied in various embodiments, it will be understood that the form and details of the foregoing apparatus or algorithms may be variously omitted, substituted, and changed without departing from the scope of the present disclosure. Since some of the features described herein can be used or practiced separately from other features, it will be readily understood that certain embodiments described herein may be embodied in forms that do not provide all of the features and advantages described herein. All changes that come within the meaning and range of equivalency of the claims are embraced within their scope.
Claims
1. A system for assessing or predicting wound healing, comprising: and an application executable by a mobile computing device, the mobile computing device including one or more processors in communication with at least one light detecting element, the one or more processors configured to perform the following operations according to processor-executable instructions included in the application, the operations comprising: receiving at least one signal from the light detecting element indicative of light of at least a first wavelength reflected from a tissue region comprising a wound or portion thereof; generating an image having a plurality of pixels representative of the tissue region based on the at least one signal; measuring a reflected intensity value of each pixel at a first wavelength for at least a subset of the plurality of pixels based on the signal; determining one or more quantitative features of the subset of pixels based on the reflectance intensity value of each pixel in the subset; and using one or more machine learning algorithms to generate at least one scalar value corresponding to a predicted amount of healing of the wound or portion thereof after a predetermined period of time from generation of the image based on the one or more quantitative features of the subset of the plurality of pixels. system.
2. The system of claim 1 , wherein the operations further comprise determining the predicted amount of healing of the wound or portion thereof after a predetermined period of time.
3. The system of claim 1 , wherein the predicted amount of healing is a predicted percentage reduction in area of the wound or portion thereof.
4. The system of claim 1 , wherein the predetermined period is 30 days.
5. 2. The system of claim 1, wherein the operations further include identifying at least one patient health measure corresponding to a patient having the tissue region, and wherein the at least one scalar value is generated based on the one or more quantitative features of the subset of the plurality of pixels and the at least one patient health measure.
6. 6. The system of claim 5, wherein the at least one patient health measurement value includes at least one variable selected from a patient demographic variable, a compliance variable, an endocrine system variable, a cardiovascular system variable, a musculoskeletal system variable, a nutritional status variable, an infectious disease variable, a renal variable, an obstetric-gynecological system variable, a medication use variable, a variable related to other diseases, or a clinical test value.
7. 6. The system of claim 5, wherein the at least one patient health measure comprises at least one characteristic selected from the group consisting of a patient's age, a degree of the patient's chronic kidney disease, a length of the wound or portion thereof on the date the image was created, and a width of the wound or portion thereof on the date the image was created.
8. 2. The system of claim 1, wherein the first wavelength is within a range of 620 nm ± 20 nm, 660 nm ± 20 nm, or 420 nm ± 20 nm, and the one or more machine learning algorithms include a random forest ensemble.
9. 10. The system of claim 1, wherein the first wavelength is within a range of 726 nm±41 nm, 855 nm±30 nm, 525 nm±35 nm, 581 nm±20 nm, or 820 nm±20 nm, and the one or more machine learning algorithms comprise an ensemble of classifiers.
10. The operation further comprises automatically segmenting pixels in the image into wound pixels and non-wound pixels; and The system of claim 1 , further comprising selecting a subset of the plurality of pixels that includes pixels of the wound.
11. 11. The system of claim 10, wherein the plurality of pixels are automatically segmented using a segmentation algorithm including at least one of a U-Net including multiple convolutional layers and a SegNet including multiple convolutional layers.
12. 2. The system of claim 1, wherein the one or more quantitative features of the subset of the plurality of pixels are selected from the group consisting of a mean value of the reflectance intensity values of pixels of the subset, a standard deviation of the reflectance intensity values of pixels of the subset, and a median value of the reflectance intensity values of pixels of the subset.
13. The operation further comprises: applying multiple filter kernels individually to the image by convolution to create multiple transformed images; constructing a three-dimensional matrix from the plurality of transformed images; and determining one or more quantitative characteristics of the three-dimensional matrix; The system of claim 1 , wherein the at least one scalar value is generated based on the one or more quantitative features of the subset of the plurality of pixels and based on the one or more quantitative features of the three-dimensional matrix.
14. 14. The system of claim 13, wherein the one or more quantitative features of the three-dimensional matrix are selected from the group consisting of a mean value of the numerical values of the three-dimensional matrix, a standard deviation of the numerical values of the three-dimensional matrix, a median value of the three-dimensional matrix, and a product of the mean and median value of the three-dimensional matrix.
15. 15. The system of claim 14, wherein the at least one scalar value is generated based on a mean value of the reflectance intensity values of the pixels of the subset, a standard deviation of the reflectance intensity values of the pixels of the subset, a median value of the reflectance intensity values of the pixels of the subset, a mean value of the numerical values of the three-dimensional matrix, a standard deviation of the numerical values of the three-dimensional matrix, and a median value of the three-dimensional matrix.
16. The operation further comprises: receiving a second signal from the at least one light detecting element indicative of light at a second wavelength reflected from the tissue region; measuring a reflected intensity value of each pixel at a second wavelength for at least a subset of the plurality of pixels based on the second signal; and determining one or more additional quantitative features of the subset of the plurality of pixels based on the reflected intensity value of each pixel at a second wavelength; The system of claim 1 , wherein the at least one scalar value is generated based at least in part on the one or more additional quantitative features of the subset of the plurality of pixels.
17. The system of claim 1 , wherein the at least one scalar value is generated on a date when a predetermined period begins.
18. The operation further comprises: receiving an indication of an actual amount of healing of the wound or portion thereof after the predetermined period of time, the indication being determined based on measuring one or more dimensions of the wound or portion thereof after the predetermined period of time has elapsed since determining the predicted amount of healing of the wound or portion thereof; and 10. The system of claim 1, further comprising updating at least one of the one or more machine learning algorithms by providing images of at least the wound or portion thereof and the actual amount of healing as training data.
19. 10. The system of claim 1, wherein the operations further comprise selecting between a standard wound care therapy and an advanced wound care therapy based at least in part on a predicted amount of healing of the wound or portion thereof before the predetermined period of time has elapsed.
20. The choice between standard wound care therapy and advanced wound care therapy is If the predicted healing amount indicates that the wound or portion thereof will be more than 50% healed or closed after 30 days, prescribing one or more standard therapies selected from improving nutritional status, debridement to remove necrotic tissue, maintaining granulation tissue with a dressing, treatment needed to resolve any infection that may be present, addressing lack of intravascular perfusion to the limb with the wound or portion thereof, reducing pressure load on the wound or portion thereof, or glycemic control; and and when the predicted healing amount indicates that the wound or portion thereof will not heal or close by more than 50% after 30 days, prescribing one or more advanced therapies selected from hyperbaric oxygen therapy, negative pressure wound therapy, bioengineered skin substitutes, synthetic growth factors, extracellular matrix proteins, matrix metalloproteinase modulators, and electrical stimulation therapy.
20. The system of claim 19.
21. The system of claim 1 , wherein the mobile computing device comprises a smartphone or a tablet.
22. 1. A method for assessing or predicting wound healing, comprising: Under the control of one or more processors of the mobile computing device; receiving, from a light detecting element of the mobile computing device, at least one signal indicative of light of at least a first wavelength reflected from a tissue region including the wound or a portion thereof; generating an image having a plurality of pixels representative of the tissue region based on the at least one signal; measuring a reflected intensity value of each pixel at the first wavelength for at least a subset of the plurality of pixels based on the signal; determining one or more quantitative features of the subset of pixels based on the reflectance intensity value of each pixel in the subset; and using one or more machine learning algorithms to generate at least one scalar value corresponding to a predicted amount of healing of the wound or portion thereof after a predetermined period of time from generation of the image based on the one or more quantitative features of the subset of the plurality of pixels. method.
23. 23. The method of claim 22, further comprising determining the predicted amount of healing of the wound or portion thereof after a predetermined period of time.
24. 23. The method of claim 22, wherein the healing amount is the predicted percent area reduction of the wound or portion thereof.
25. 23. The method of claim 22, wherein the mobile computing device comprises a smartphone or a tablet.
Citation Information
Patent Citations
Apparatus and method for imaging and monitoring based on fluorescence
JP2011521237A
Collecting and analyzing data for diagnostic purposes
JP2017524935A
Reflection-mode multi-spectral time-resolved optical imaging method and apparatus for tissue classification
JP2018502677A
Methods and systems for evaluating tissue healing
JP2018534965A
Oxyvu-1 hyperspectral tissue oxygenation (HTO) measurement system
US20140257113A1