Automated classifier for diagnosing acute otitis media

A DR-RNN-based system for AOM diagnosis on tympanic membrane images addresses low accuracy in primary care by using a large training set and identifying key features, enhancing diagnostic reliability and explainability.

WO2025184335A1PCT designated stage Publication Date: 2025-09-04UNIV OF PITTSBURGH OF THE COMMONWEALTH SYST OF HIGHER EDUCATION
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/017574
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-28
Filing Date
2025-02-27
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Existing diagnostic methods for acute otitis media (AOM) in primary care settings suffer from low accuracy due to the use of non-generalizable training data and inclusion of ideal or obstructed images, limiting the effectiveness of neural networks in clinical applications.

Method used

A system utilizing a deep residual-recurrent neural network (DR-RNN) trained on tympanic membrane images to predict features such as color, position, translucency, distinct erythema, and air-fluid interface, determining AOM diagnosis with confidence and identifying predominant contributing features, supported by a quality filter for image suitability.

Benefits of technology

Enhances diagnostic accuracy for AOM by leveraging a large training set from primary care clinics, improving the reliability of neural network models in non-ideal conditions, and providing explainable AI for feature contributions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025017574_04092025_PF_FP_ABST
    Figure US2025017574_04092025_PF_FP_ABST
Patent Text Reader

Abstract

A system for assessing AOM includes a computing device having a processor apparatus, wherein the processor apparatus implements a diagnostic classifier component that comprises a. trained neural network, the processor apparatus being structured and configured to receive tympanic membrane image data representing one or more images of a tympanic membrane of the patient, provide the tympanic membrane image data to the diagnostic classifier component, and process the tympanic membrane image data in the diagnostic classifier component to determine: (i) a plurality of tympanic membrane features from the tympanic membrane image data, (ii) a diagnosis of whether the patient has AOM based on the plurality of tympanic membrane features, (iii) a. confidence in the diagnosis, and (iv) an identification one or more of the tympanic membrane features determined to be predominant contributing features that led to the diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

AUTOMATED CLASSIFIER FOR DIAGNOSING ACUTE OTITIS MEDIA CROSS REFERENCE TO RELATED APPLICATIONS:

[0001] This application claims priority to U.S. Provisional Patent Application Serial No. 63 / 558,812, filed on February 28, 2024, and titled “Automated Classifier for Diagnosing Acute Otitis Media,” the disclosure of which is incorporated herein by reference. FIELD OF THE INVENTION

[0002] The disclosed concept relates generally to the diagnosis of acute otitis media (AOM) in patients, and, in particular, to a system and method for diagnosing AOM based on tympanic membrane images using an automated classifier in the form of a trained neural network, such as, without limitation, a deep residual-recurrent neural network (DR-RNN). BACKGROUND OF THE INVENTION

[0003] Otitis media is a general term for middle-ear inflammation and may be classified clinically as either acute otitis media (AOM) or otitis media with effusion (OME). AOM represents a bacterial superinfection of the middle ear fluid. OME, on the other hand, is a sterile effusion that tends to subside spontaneously.

[0004] AOM is the second most frequently diagnosed illness in children in the United States following the common cold and is the most cited indication for antimicrobials. Despite the high prevalence of AOM, accuracy of diagnosis has been consistently 75% or lower across primary care and pediatric practitioners. Methods used to enhance accuracy and facilitate the diagnosis of AOM have evolved over time. Training programs, such as the “Enhancing Proficiency in Otitis Media” curriculum, were developed to improve practitioners’ skills in diagnosing otitis media. Other clinical tools have included pneumatic otoscopy, tympanometry, smartphone-based otoscope attachments, serum biomarkers, and novel imaging technologies such as a light field otoscope and optical coherence tomography. Despite these efforts, diagnostic accuracy remains low and, accordingly, further innovation is warranted.

[0005] Recently, efforts toward improving diagnostic accuracy of AOM have focused on developing artificial intelligence algorithms. Multiple research teams have used deep learning to train neural networks to recognize the presence of AOM and, in some instances, other ear-related diagnoses. Some previously developed neural networks have limited clinical applications due to training with ideal, non-obstructed images. Other training data sets werecollected in specialty clinics or intraoperatively on sedated patients, which limits generalizability. Although most cases of AOM are diagnosed in a primary care setting, only one previous model was developed using data collected from a primary care clinic. The Kuruvilla et al. model (Kuruvilla A, Shaikh N, Hoberman A, Kovacevic J., Automated diagnosis of otitis media: vocabulary and grammar, Int J Biomed Imaging.2013:2013: 327515. doi: 10.1155 / 2013 / 327515) is based on a relatively small sample size, which limits its clinical application. There are no studies that use a large training set collected from a primary care setting and include nonideal or partially obstructed images. SUMMARY OF THE INVENTION

[0006] In one embodiment, a system for assessing acute otitis media (AOM) in a patient is provided that includes a computing device having a processor apparatus, wherein the processor apparatus implements a diagnostic classifier component that comprises a trained neural network, the processor apparatus being structured and configured to receive tympanic membrane image data representing one or more images of a tympanic membrane of the patient, provide the tympanic membrane image data to the diagnostic classifier component, and process the tympanic membrane image data in the diagnostic classifier component to determine: (i) a plurality of tympanic membrane features from the tympanic membrane image data, (ii) a diagnosis of whether the patient has AOM based on the plurality of tympanic membrane features, (iii) a confidence in the diagnosis, and (iv) an identification one or more of the tympanic membrane features determined to be predominant contributing features that led to the diagnosis.

[0007] In another embodiment, a method for assessing acute otitis media (AOM) in a patient is provided that includes receiving tympanic membrane image data representing one or more images of a tympanic membrane of the patient, providing the tympanic membrane image data to a diagnostic classifier component that comprises a trained neural network, and processing the tympanic membrane image data in the diagnostic classifier component to determine: (i) a plurality of tympanic membrane features from the tympanic membrane image data, (ii) a diagnosis of whether the patient has AOM based on the plurality of tympanic membrane features, (iii) a confidence in the diagnosis, and (iv) an identification one or more of the tympanic membrane features determined to be predominant contributing features that led to the diagnosis.

[0008] In yet another embodiment, a method for assessing whether a number of images of a tympanic membrane of a patient are suitable for use in a system configured forassessing acute otitis media (AOM) in the patient using a trained neural network is provided. The method includes receiving the tympanic membrane image data representing the number of images of the tympanic membrane of the patient, providing the tympanic membrane image data to a trained image quality classifier component comprising a trained frame-level, binary quality filter, wherein the trained image quality classifier component has been previously trained using a plurality of frames from a plurality of training images and expert annotation data indicating suitability of the frames for use in the system, and processing the tympanic membrane image data in the trained image quality classifier component to determine whether the tympanic membrane image data is suitable for use in the system.

[0009] In still another embodiment, an apparatus for assessing whether a number of images of a tympanic membrane of a patient are suitable for use in a system configured for assessing acute otitis media (AOM) in the patient using a trained neural network is provided. The apparatus includes a processor apparatus implementing a trained image quality classifier component comprising a trained frame-level, binary quality filter, wherein the trained image quality classifier component has been previously trained using a plurality of frames from a plurality of training images and expert annotation data indicating suitability of the frames for use in the system. The processor apparatus is structured and configured to receive the tympanic membrane image data representing the number of images of the tympanic membrane of the patient, provide the tympanic membrane image data to the trained image quality classifier component, and process the tympanic membrane image data in the trained image quality classifier component to determine whether the tympanic membrane image data is suitable for use in the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] A full understanding of the invention can be gained from the following description of the preferred embodiments when read in conjunction with the accompanying drawings in which:

[0011] FIG.1 is a block diagram of a diagnostic system for determining features of the tympanic membrane (TM) of a patient based on tympanic membrane image data is that is captured from the patient (e.g., in video form or as one or more still images captured using a device such as an otoscope based imaging device) and determining and outputting certain AOM diagnostic information associated with the patient based on the determined features according to an exemplary embodiment of the disclosed concept;

[0012] FIG.2 is a block diagram of an otoscopic video capture device according toone exemplary embodiment of the disclosed concept;

[0013] FIG.3 is a block diagram of a computing device for implementing the diagnostic system of FIG.1 according to one exemplary embodiment of the disclosed concept;

[0014] FIG.4 is a schematic diagram of an example neural network that is implemented as part of diagnostic classifier component for implementing the diagnostic system of FIG.1 according to an exemplary embodiment of the disclosed concept;

[0015] FIG.5 is diagram showing an exemplary API raw output according to an exemplary embodiment of the disclosed concept;

[0016] FIG.6 is diagram showing an exemplary API raw output according to an alternative exemplary embodiment of the disclosed concept. DETAILED DESCRIPTION OF THE EXEMPLARY EMBODIMENTS

[0017] As used herein, the singular form of “a”, “an”, and “the” include plural references unless the context clearly dictates otherwise.

[0018] As used herein, the statement that two or more parts or components are “coupled” shall mean that the parts are joined or operate together either directly or indirectly, i.e., through one or more intermediate parts or components, so long as a link occurs.

[0019] As used herein, the term “number” shall mean one or an integer greater than one (i.e., a plurality).

[0020] As used herein, the terms “component” and “system” are intended to refer to a computer related entity, either hardware, a combination of hardware and software, software, or software in execution. For example, a component can be, but is not limited to being, a process running on a processor, a processor, an object, an executable, a thread of execution, a program, and / or a computer. By way of illustration, both an application running on a server and the server can be a component. One or more components can reside within a process and / or thread of execution, and a component can be localized on one computer and / or distributed between two or more computers.

[0021] Directional phrases used herein, such as, for example and without limitation, top, bottom, left, right, upper, lower, front, back, and derivatives thereof, relate to the orientation of the elements shown in the drawings and are not limiting upon the claims unless expressly recited therein.

[0022] The disclosed concept will now be described, for purposes of explanation, inconnection with numerous specific details in order to provide a thorough understanding of the disclosed concept. It will be evident, however, that the disclosed concept can be practiced without these specific details without departing from the spirit and scope of this innovation.

[0023] As described herein, the disclosed concept provides a system and method for determining a number of predetermined features of the tympanic membrane (TM) of a patient based on a number of tympanic membrane images that is / are captured from the patient (e.g., in the form of video or a number of still images) and determining certain diagnostic information associated with the patient (as described below) based on the predicted features. The number of tympanic membrane images may, for example and without limitation, be captured using any suitable imaging device, such as, without limitation, a digital otoscope having video recording capability (e.g., a video-otoscope or an otoendoscope) or by a smartphone having an otoscope adapter attached thereto.

[0024] More specifically, data representing the number of tympanic membrane images that is / are captured from a patient is fed into a diagnostic classifier component comprising a trained neural network. In the non-limiting exemplary embodiment, the trained neural network is a deep residual-recurrent neural network (DR-RNN) as described in greater detail herein. The diagnostic classifier component is configured to perform the following based on the received tympanic membrane image data: (i) predict a plurality of predetermined features (e.g., color, position, translucency, distinct erythema, and air-fluid interface) of the tympanic membrane of the patient (including a predicted probability for each feature and an impact for each feature as described herein), (ii) determine a diagnosis (i.e., AOM v. no AOM) associated with the patient based on the feature determinations / predictions, (iii) determine a confidence in the determined diagnosis, and (iv) determine the predominant contributing features (e.g., the two or three main features) that led to that diagnosis, including the percentage impact that such features had on reaching the diagnosis (based on the “impact” determined for each feature as described herein). In the exemplary embodiment, the predominant contributing features are a subset of all of the features, and the number of features included in the subset may be selected by the user. For example, the user may determine that the diagnostic classifier component should output only the two (or three) features that have the largest percentage impact on reaching the diagnosis as the predominant contributing features. In another exemplary embodiment, the determined predominant contributing features are those features that have a determined impact that is greater than a predetermined level, such as, without limitation 20%. Thus, the disclosed concept includes an explainable AI component that is configured to extract information fromthe last two portions of the neural network (the perceptron portion 105 and the fully connected neural network portion 95 as described herein) and based on that information, provide information about the features that were the most influential for each diagnosis. As noted above, the five predetermined features that are considered in the non-limiting exemplary embodiment are color, position, translucency, distinct erythema, and air-fluid interface (each of which is described in more detail below). The impact of each feature (with confidence) in determining a diagnosis of AOM vs. no AOM is determined, totaling 100%.

[0025] In the exemplary embodiment, the neural network is previously trained to predict each feature using the tympanic membrane image data as the input in a binary fashion (i.e., present (1) or not present (0)). More specifically, each feature is “dichotomized” in advance and training data for training the neutral network is created by having expert clinicians annotate / label tympanic image data by indicating whether each dichotomized feature is present or not present in images comprising the by indicating whether each dichotomized feature is present or not present in images comprising the training data. The expert clinicians also label each image in the training data with an overall diagnosis (AOM or No AOM). That training data is then used to train and test the neural network before it is deployed for diagnostic purposes as described herein.

[0026] The color feature is dichotomized such that certain colors in the tympanic membrane (e.g., white, pale yellow) are in one category while certain other colors in the tympanic membrane (e.g., gray, pink, amber, blue) are in another category. During diagnosis, the color feature is thus determined by the diagnostic classifier component to be present or absent based on a color-based prediction performed by the trained neural network.

[0027] The position feature is dichotomized such that an outwardly bulging tympanic membrane is in one category, while a neutral or retracted tympanic membrane is in another category. During diagnosis, the position feature is thus determined by the diagnostic classifier component to be present or absent based on a bulginess-based prediction performed by the trained neural network.

[0028] The translucency feature is dichotomized such that an opaque tympanic membrane (indicative of a possible infection) is in one category, while a translucent tympanic membrane is in another category. During diagnosis, the translucency feature is thus determined by the diagnostic classifier component to be present or absent based on a translucency-based prediction performed by the trained neural network.

[0029] The distinct erythema feature is dichotomized such that the presence of distinct erythema as evidenced by marked redness in the tympanic membrane is in onecategory, while the absence of distinct erythema is in another category. During diagnosis, the distinct erythema feature is thus determined by the diagnostic classifier component to be present or absent based on a distinct erythema / marked redness-based prediction performed by the trained neural network.

[0030] The air-fluid interface feature is dichotomized such that the presence of an interface (i.e. line) between air (e.g., a bubble) and fluid in or around the tympanic membrane is in one category, while the absence of an interface is in another category. During diagnosis, the air-fluid interface feature is thus determined by the diagnostic classifier component to be present or absent based on an interface-based prediction performed by the trained neural network.

[0031] According to the disclosed concept, the model learns to predict features and diagnoses at the same time (i.e., combined learning). In the exemplary embodiment, the combined learning is implemented via weighted loss functions for both feature determinations and diagnosis predictions. As described in detail herein, the DR-RNN of the exemplary embodiment has, among other layers, a diagnosis layer (the final decision layer) that provides supervision to the earlier layers to predict features more accurately. In the exemplary embodiment, as noted above, for each of the number of tympanic membrane images that is / are captured for a patient, the diagnostic classifier component including the DR-RNN model: (i) predicts a plurality of predetermined features (e.g., color, position, translucency, distinct erythema, and air-fluid interface) of the tympanic membrane of the patient as described herein, (ii) determines a diagnosis (i.e., AOM v. no AOM) associated with the patient based on the feature determinations / predictions, (iii) determines a confidence in the determined diagnosis, and (iv) determines the predominant contributing features (e.g., the two main features) that led to that diagnosis, including the percentage impact that such features had on reaching the diagnosis. In the exemplary embodiment, if the probability is 50% or greater, the model diagnoses AOM. A cutoff of 50% is used because in well trained models, most cases will have been assigned a probability of close to 100% and most controls a probability of close to 0%.

[0032] According to an aspect of the disclosed concept, the neural network (e.g., DR- RNN) is previously trained to predict both features of the tympanic membrane and the diagnosis of AOM vs no AOM using tympanic membrane image data obtained from a patient cohort that have been annotated by validated otoscopists using the feature facts and assumptions described above. In one particular exemplary implementation during development of the disclosed concept, this was done by developing a training library ofotoscopic assessments on children presenting for well or sick visits to outpatient pediatric offices. A video clip was taken of each child’s TM using an endoscope (Storz Hopkins) or an otoscope (Hillrom Macroview Plus) connected to a smartphone with an appropriate adapter. Saved videos were reviewed by two validated otoscopists who assigned a final diagnosis. Disagreements were resolved by discussion. Videos for which experts could not arrive at a diagnosis (because of near complete occlusion of the TM with cerumen or because the video was completely out of focus) were excluded. Expert consensus was used as the reference standard in this exemplary implementation because myringotomy and tympanocentesis are invasive and, accordingly, not practical for use in a large cohort of unselected children.

[0033] According to an optional further aspect of the disclosed concept, a software application may be provided to assist users (e.g., care providers) with the capture and annotation of the TM images. In the exemplary embodiment, the application is a medical grade smartphone application that uses the main camera on the smartphone to capture video or still image(s) of the TM. The application may also be configured to operate on a digital otoscope in the event such a device is used for the image capture. Using the application, the user can adjust focus and brightness to obtain the best possible image for further processing. In addition, voice recognition software is embedded to allow users to take video clips of the TM using voice commands (e.g., capture and stop). After recording the video or still images, the application allows cropping to facilitate further processing. Optionally, the application also allows users to record their impressions regarding the appearance of the TM and their presumptive diagnosis.

[0034] In addition, according to another aspect of the disclosed concept, a quality filter is provided to prompt users that a video segment or a number of still images that are acquired may not be adequate for diagnostic purposes. In the exemplary embodiment, the quality filter is a previously trained image quality classifier. Also in the exemplary embodiment, the quality filter is previously trained using a plurality of frames from non- cropped video recordings. In particular, the frames of the videos are sampled (e.g., 10.0%; 1 frame from each consecutive 10 frames, reducing 30 frames per second to 3 frames per second at equal intervals), and multidimensional data is extracted from the frames. The data are then reduced to two dimensions, and frames are allocated into clusters (e.g., 100 clusters) using a k-means algorithm. Based on the frames in each cluster, each cluster is expert annotated as accept or reject. The data as just described is then used to train a frame-level, binary quality filter (e.g., a deep learning image classification model) with an output of accept or reject. In the exemplary embodiment, videos in which 70% or more of the framesare rejected generate a prompt that encourages users to obtain a new recording. As described elsewhere herein, the quality filter may be implemented locally on the image capture device being used, or remotely in server, such as a cloud-based server, that is in communication with the image capture device being used.

[0035] FIG.1 is a block diagram of a diagnostic system 5 for determining features of the tympanic membrane (TM) (e.g., color, position, translucency, distinct erythema, and air- fluid interface) of a patient based on tympanic membrane image data is that is captured from the patient and determining and outputting certain AOM diagnostic information associated with the patient based on the determined features according to an exemplary embodiment of the disclosed concept. In the exemplary embodiment, the AOM diagnostic information includes: (i) the determined diagnosis for the patient (i.e., AOM v. no AOM), (ii) a confidence in the diagnosis, and (iii) the predominant contributing features (e.g., the two main features) that led to that diagnosis, including the percentage impact that such features had on reaching the diagnosis. As seen in FIG.1, the non-limiting exemplary diagnostic system 5 includes an otoscopic video capture device 10 that is structured to be able to capture video from within the auditory canal of a patient, and in particular images of the TM of the patient. For example, and without limitation, otoscopic video capture device 10 may be a video-otoscope, an otoendoscope, or a smartphone having an otoscope adapter attached thereto. It will be appreciated, however, that these examples are not meant to be limiting and that diagnostic system 5 may use any suitable image capture device for capturing still or video images of a patient’s tympanic membrane. Diagnostic system 5 further includes a computing device 15. Computing device 15 is structured to receive video data from otoscopic video capture device 10 by, for example, a wired or wireless connection, including over a suitable network. In one embodiment, computing device 15 may be a remote device such as, without limitation, a cloud-based server computer. In another embodiment, computing device may be a local device such as, for example and without limitation, a PC, a laptop computer, a tablet computer, a smartphone, or any other suitable device structured to perform the functionality described herein. Computing device 15 is structured and configured to receive the video data from otoscopic video capture device 10 and process the video data as described in detail herein to determine features of the TM of the patient and / to determine and output the AOM diagnostic information for the patient.

[0036] FIG.2 is a block diagram of otoscopic video capture device 10 according to one non-limiting exemplary embodiment. As seen in FIG.2, the exemplary otoscopic video capture device 10 includes an input apparatus 20 (e.g., a keyboard or touchscreen), an outputapparatus 25 (e.g., an LCD or other display for visual output and / or a speaker for audio output), a video camera 30, a light / speculum assembly 35 (e.g., an otoscope adapter for a smartphone) coupled to the video camera 30, and a processor apparatus 40. A user is able to provide input into processor apparatus 40 using input apparatus 20, and processor apparatus 40 provides output signals to output apparatus 25 to enable output apparatus 25 to provide visual and / or audio information to the user as described herein. Processor apparatus 40 comprises a processor portion and a memory portion. The processor portion may be, for example and without limitation, a microprocessor (μP), a microcontroller, or some other suitable processing device, that interfaces with the memory portion. The memory portion can be any one or more of a variety of types of internal and / or external storage media such as, without limitation, RAM, ROM, EPROM(s), EEPROM(s), FLASH, and the like that provide a storage register, i.e., a machine readable medium, for data storage such as in the fashion of an internal storage area of a computer, and can be volatile memory or nonvolatile memory. The memory portion of processor apparatus 40 has stored therein a number of routines that are executable by the processor portion of processor apparatus 40. As seen in FIG.2, one or more of the routines implement (by way of computer / processor executable instructions) a video capture component 45 that comprises the software application for assisting users with the capture and annotation of videos described elsewhere herein. In addition, in this exemplary embodiment, one or more of the routines implement (by way of computer / processor executable instructions) an image quality classifier component 50 that comprises the quality filter described elsewhere herein. In this exemplary embodiment, the image quality classifier functionality is thus provided locally as part of the device that is capturing the images (e.g., as part of a digital otoscope or smartphone). It will be understood, however, that this is not meant to be limiting and that in alternative embodiments, the image quality classifier component 50 may be located separately from the device that captures the images, such as part of a remote (e.g., cloud-based) server in communication with otoscopic video capture device 10. Thus, regardless of the embodiment that is implemented, the user of otoscopic video capture device 10 will be provided with assistance with the capture of videos that helps to ensure that videos of suitable quality are readily and easily captured.

[0037] FIG.3 is a block diagram of computing device 15 according to one exemplary embodiment. As noted above, computing device 15 may be a local device, such as PC or similar device, or a remote device, such as a remote server computer. As seen in FIG.3, the exemplary computing device 15 includes an input apparatus 55 (which in the illustrated embodiment is a keyboard or a touchscreen), a display 60 (which in the illustratedembodiment is an LCD or other display), and a processor apparatus 65. A user is able to provide input into processor apparatus 65 using input apparatus 55, and processor apparatus 65 provides output signals to display 60 to enable display 60 to display information to the user, such as determined TM features and / or AOM diagnostic information as described herein. Processor apparatus 65 comprises a processor portion and a memory portion. The processor portion may be, for example and without limitation, a microprocessor (μP), a microcontroller, or some other suitable processing device, that interfaces with the memory portion. The memory portion can be any one or more of a variety of types of internal and / or external storage media such as, without limitation, RAM, ROM, EPROM(s), EEPROM(s), FLASH, and the like that provide a storage register, i.e., a machine readable medium, for data storage such as in the fashion of an internal storage area of a computer, and can be volatile memory or nonvolatile memory. The memory portion has stored therein a number of routines that are executable by the processor potion. One or more of the routines implement (by way of computer / processor executable instructions) a diagnostic classifier component 70 that comprises a trained neural network (e.g., a DR-RNN) as described herein. Diagnostic classifier component is thus configured to receive the tympanic membrane image data as described herein and to determine and output certain predetermined TM features and certain AOM diagnostic information as described herein.

[0038] While the embodiments shown in FIGS.1-3 contemplate diagnostic classifier component 70 being part of computing device 15 (local or remote) that is separate from otoscopic video capture device 10, that is meant to be exemplary only. It will be understood that in alternative embodiments, diagnostic classifier component 70 may be implemented in otoscopic video capture device 10 in processor apparatus 40. In that embodiment, otoscopic video capture device 10 can be used to both capture the video and process the video to determine features and diagnose AOM.

[0039] FIG.4 is a schematic diagram of a neural network 75 that is implemented as part of diagnostic classifier component 70 according to one particular non-limiting exemplary embodiment of the disclosed concept. In the illustrated exemplary embodiment, neural network 75 comprises a DR-RNN, which stands for deep residual recurrent neural network, and comprises a neural network having a deep residual neural network followed by a recurrent neural network. FIG.4 illustrates the learning and decision-making paradigm of neural network 75. As seen in FIG.4, neural network 75 includes a deep residual neural network portion 80, a long short-term memory recurrent neural network portion 85, a Bahdanau attention layer portion 90, a fully connected neural network portion 95, a featuredetermination and output layer portion 100, a perceptron portion 105, and a diagnosis output portion 110. The DR-RNN is similar to the visual cognition of a human brain.

[0040] As noted elsewhere herein, diagnostic classifier component 70 including neural network 75 is structured and configured to receive tympanic membrane image data and based thereon: (i) predict a plurality of predetermined features (e.g., color, position, translucency, distinct erythema, and air-fluid interface) of the tympanic membrane of the patient (including a predicted probability for each feature and an impact for each feature as described herein), (ii) determine a diagnosis (i.e., AOM v. no AOM) associated with the patient based on the feature determinations / predictions, (iii) determine a confidence in the determined diagnosis, and (iv) determine the predominant contributing features (e.g., the two main features) that led to that diagnosis, including the percentage impact that such features had on reaching the diagnosis. As noted elsewhere herein, the subset of features that are to be included and output as the predominant contributing features may be selectable by the user, such as, without limitation, the top two or three features that had the largest impact on the diagnosis

[0041] In the exemplary embodiment, the received tympanic membrane image data is processed through all layers of neural network 75, each applying transformations using learned weights and activation functions. Fully connected neural network portion 95 predicts the probability that each of the five features (color, position, translucency, distinct erythema, and air-fluid interface) is associated with AOM. The predicted probability for each of the features is provided to perceptron portion 105. Perceptron portion 105 applies learned weights and biases to the predicted probability for each of the features to compute a final AOM probability score. More specifically, to compute the final AOM probability score, Perceptron portion 105 outputs raw scores (logits) for each of the two possible classes (i.e., AOM and Not AOM). These logits are converted into probabilities using the softmax function, which ensures that all values are between 0 and 1 and sum to 1, creating a probability distribution over the two possible outcomes (AOM vs. Not AOM). The final overall diagnosis (i.e., AOM or Not AOM) and a confidence in the final overall diagnosis is a function of the predicted probability. For example, in the exemplary embodiment, a 50% probability indicates maximum uncertainty, meaning no confidence in the prediction. Thus, for predicted probability of 50% or less, zero confidence is assigned. For predicted probabilities of greater than 50%, the overall diagnosis is assigned a confidence value that is greater than zero and less than or equal to 100%. Also in the exemplary embodiment, confidence is defined as the inverse of entropy of probability distribution. For binaryclassification tasks like AOM vs non-AOM classification according to the exemplaryembodiment of the disclosed concept, confidence = 1+P x log(P) + (1 P) x log(1 P), where Pis the probability of a diagnosed class.

[0042] Moreover, each feature is assigned a learned weight and bias during training. Perceptron portion 105 determines a raw impact score for each feature by applying a linear transformation to the predicted AOM probability for the feature based on the learned weight and bias for the feature. Thus, in the exemplary embodiment, the impact of a feature is computed as: Impact=(Predicted AOM Probability of Feature x Weight) + Bias. For example, if the color feature has an 80% probability of AOM (computed by fully connected neural network portion 95) and perceptron portion 105 assigns it a 30% learned weight with a 5% learned bias, its pre-normalization raw impact score will be: 0.8 x 0.3 + 0.05 = 0.29. The raw impact scores for each feature are then normalized so that their sum equals 1.

[0043] FIG.5 is diagram of an exemplary API raw output (comprising (i)-(iii) above) generated by processor apparatus 65 based on the outputs of neural network 75. The API raw output may then be provided to an application, such as a mobile / web application, that is configured to display the outputs (i.e., the AOM diagnostic information as described herein) according to end user needs.

[0044] FIG.6 is diagram of an alternative exemplary API raw output (comprising (i)- (iii) above) generated by processor apparatus 65 based on the outputs of neural network 75. As seen in FIG.6, this API raw output includes an “explanations” portion that identifies the determined predominant contributing features that led to that diagnosis so that those explanations / features can be displayed to the user. In the illustrated example, the “explanations” include “Fluid Level” (corresponding to the air-fluid interface feature), “No Erythema” (corresponding to the distinct erythema feature), “Color” (corresponding to the color feature), and “Translucent” (corresponding to the translucency feature). In this exemplary embodiment, the determined predominant contributing features are those features that have an impact that is greater than a predetermined level, such as, without limitation 20%. The API raw output may then be provided to an application, such as a mobile / web application, that is configured to display the outputs (i.e., the AOM diagnostic information as described herein) according to end user needs.

[0045] In the exemplary embodiments described above, the exemplary predetermined features are color, position, translucency, distinct erythema, and air-fluid interface. It will be appreciated, however, that that is meant to be exemplary only, and that two or more, three or more, or four or more of those features, in any combination, may also be used within thescope of the disclosed concept. Furthermore, one or more other features, either alone or in combination with one or more of color, position, translucency, distinct erythema, and air- fluid interface may also be utilized without departing from the scope of the invention.

[0046] While specific embodiments of the invention have been described in detail, it will be appreciated by those skilled in the art that various modifications and alternatives to those details could be developed in light of the overall teachings of the disclosure. Accordingly, the particular arrangements disclosed are meant to be illustrative only and not limiting as to the scope of disclosed concept which is to be given the full breadth of the claims appended and any and all equivalents thereof.

Claims

What is claimed is:

1. A system for assessing acute otitis media (AOM) in a patient, comprising: a computing device having a processor apparatus, wherein the processor apparatus implements a diagnostic classifier component that comprises a trained neural network, the processor apparatus being structured and configured to: receive tympanic membrane image data representing one or more images of a tympanic membrane of the patient; provide the tympanic membrane image data to the diagnostic classifier component; and process the tympanic membrane image data in the diagnostic classifier component to determine: (i) a plurality of tympanic membrane features from the tympanic membrane image data, (ii) a diagnosis of whether the patient has AOM based on the plurality of tympanic membrane features, (iii) a confidence in the diagnosis, and (iv) an identification one or more of the tympanic membrane features determined to be predominant contributing features that led to the diagnosis.

2. The system according to claim 1, wherein the processor apparatus is structured and configured to determine a percentage impact on reaching the diagnosis for each of the one or more of the tympanic membrane features determined to be predominant contributing features. The system according to claim 1, wherein the one or more of the tympanic membrane features determined to be predominant contributing features is a subset of all of the tympanic membrane features.

4. The system according to claim 1, wherein the processor apparatus is structured and configured to determine a percentage impact on reaching the diagnosis for each of the tympanic membrane features, and wherein the one or more of the tympanic membrane features determined to be predominant contributing features are the tympanic membrane features that have a determine percentage impact that is greater than a predetermined level.

5. The system according to claim 4, wherein the predetermined level is 20%.

6. The system according to claim 1, wherein the processor apparatus is structured and configured to output the diagnosis, the confidence, and information indicative of the one or more of the tympanic membrane features determined to be predominant contributing features for display by a display apparatus.

7. The system according to claim 1, wherein the tympanic membrane features include a position feature based on a bulginess of the tympanic membrane, a color feature based on a color of the tympanic membrane, a translucency feature based on a translucency of the tympanic membrane , a distinct erythema feature based on whether distinct erythema is present in the tympanic membrane , and an air-fluid interface feature based on whether an air fluid interface is present in association with the tympanic membrane.

8. The system according to claim 7, wherein each of the tympanic membrane features is a binary feature.

9. The system according to claim 1, wherein the trained neural network is a trained deep residual-recurrent neural network (DR-RNN).

10. The system according to claim 9, wherein the trained DR-RNN includes a plurality of portions each including a number of layers, wherein the plurality of portions include a deep residual neural network portion, a long short-term memory recurrent neural network portion, a Bahdanau attention layer portion, a fully connected neural network portion, a feature determination and output layer portion, a perceptron layer portion, and a diagnosis output portion.

11. The system according to claim 1, wherein the trained DR-RNN implements combined learning through weighted loss functions for both the features and the diagnosis.

12. The system according to claim 1, further comprising an otoscopic image capture device for capturing the one or more images of the tympanic membrane of the patient.

13. The system according to claim 10, wherein the otoscopic image capture device comprises a smartphone having an otoscopic adapter coupled thereto.

14. A method for assessing acute otitis media (AOM) in a patient, comprising: receiving tympanic membrane image data representing one or more images of a tympanic membrane of the patient; providing the tympanic membrane image data to a diagnostic classifier component that comprises a trained neural network; and processing the tympanic membrane image data in the diagnostic classifier component to determine: (i) a plurality of tympanic membrane features from the tympanicmembrane image data, (ii) a diagnosis of whether the patient has AOM based on the plurality of tympanic membrane features, (iii) a confidence in the diagnosis, and (iv) an identification one or more of the tympanic membrane features determined to be predominant contributing features that led to the diagnosis.

15. The method according to claim 14, the processing further comprising determining a percentage impact on reaching the diagnosis for each of the one or more of the tympanic membrane features determined to be predominant contributing features.

16. The method according to claim 14, wherein the one or more of the tympanic membrane features determined to be predominant contributing features is a subset of all of the tympanic membrane features.

17. The method according to claim 14, the processing further comprising determining a percentage impact on reaching the diagnosis for each of the tympanic membrane features, wherein the one or more of the tympanic membrane features determined to be predominant contributing features are the tympanic membrane features that have a determine percentage impact that is greater than a predetermined level.

18. The method according to claim 17, wherein the predetermined level is 20%.

19. The method according to claim 14, further comprising outputting the diagnosis, the confidence, and the one or more of the tympanic membrane features determined to be predominant contributing features for display by a display apparatus.

20. The method according to claim 14, wherein the tympanic membrane features include a position feature based on a bulginess of the tympanic membrane, a color feature based on a color of the tympanic membrane, a translucency feature based on a translucency of the tympanic membrane, a distinct erythema feature based on a determination of whether distinct erythema is present in the tympanic membrane, and an air-fluid interface feature based on a determination of whether an air fluid interface is present in association with the tympanic membrane.

21. The method according to claim 20, wherein each of the tympanic membrane features is a binary feature.

22. The method according to claim 14, wherein the trained neural network is a trained deep residual-recurrent neural network (DR-RNN).

23. The method according to claim 22, wherein the trained DR-RNN includes a plurality of portions each including a number of layers, wherein the plurality of portions include a deep residual neural network portion, a long short-term memory recurrent neural network portion, a Bahdanau attention layer portion, a fully connected neural network portion, a feature determination and output layer portion, a perceptron layer portion, and a diagnosis output portion.

24. The method according to claim 14, wherein the trained DR-RNN implements combined learning through weighted loss functions for both the features and the diagnosis.

25. An apparatus for assessing whether a number of images of a tympanic membrane of a patient are suitable for use in a system configured for assessing acute otitis media (AOM) in the patient using a trained neural network, comprising: a processor apparatus implementing a trained image quality classifier component comprising a trained frame-level, binary quality filter, wherein the trained image quality classifier component has been previously trained using a plurality of frames from a plurality of training images and expert annotation data indicating suitability of the frames for use in the system, the processor apparatus being structured and configured to: receive the tympanic membrane image data representing the number of images of the tympanic membrane of the patient; provide the tympanic membrane image data to the trained image quality classifier component; and process the tympanic membrane image data in the trained image quality classifier component to determine whether the tympanic membrane image data are suitable for use in the system.

26. The apparatus according to claim 25, wherein the processor apparatus is structured and configured to provide a video and / or audio output to a user if the trained image quality classifier component determines that the tympanic membrane image data is not suitable for use in the system.

27. The apparatus according to claim 25, wherein the tympanic membrane image data represents a plurality of video frames, wherein the trained image quality classifier component is structured and configured to process the video frames and determine that thetympanic membrane image data is not suitable for use in the system if at least a certain amount of the video frames are rejected by the trained image quality classifier component.

28. An otoscopic image capture device, comprising: a speculum; a camera coupled to the speculum for capturing the number of images of the tympanic membrane of the patient and generating tympanic membrane image data; and a processor apparatus according to claim 21.

29. The otoscopic image capture device according to claim 28, wherein the processor apparatus implements a video capture component comprising voice recognition software to enable capture of the video of the tympanic membrane of the patient using voice commands.

30. The otoscopic image capture device according to claim 28, wherein the processor apparatus implements a video capture component structured and configured to enable voice input of impressions regarding an appearance of the tympanic membrane and a presumptive diagnosis.

31. The otoscopic image capture device according to claim 28, wherein the processor apparatus implements a video capture component structured and configured to adjust a focus and brightness of the camera and to crop the captured video.

32. The otoscopic image capture device according to claim 28, wherein the otoscopic image capture device comprises a smartphone.

33. A method for assessing whether a number of images of a tympanic membrane of a patient are suitable for use in a system configured for assessing acute otitis media (AOM) in the patient using a trained neural network, the method comprising: receiving the tympanic membrane image data representing the number of images of the tympanic membrane of the patient; providing the tympanic membrane image data to a trained image quality classifier component comprising a trained frame-level, binary quality filter, wherein the trained image quality classifier component has been previously trained using a plurality of frames from a plurality of training images and expert annotation data indicating suitability of the frames for use in the system; and processing the tympanic membrane image data in the trained image quality classifier component to determine whether the tympanic membrane image data is suitable for use in the system.

Citation Information

Patent Citations

  • Automatic extraction of echocardiographic measurement results from medical images

    CN111511287B

  • Surgical task data derivation from surgical video data

    CN116806349A

  • Apparatuses and methods for mobile imaging and analysis

    US20150065803A1

  • Method and apparatus for aiding in the diagnosis of otitis media by classifying tympanic membrane images

    US20170209078A1

  • Efficient surgical center workflow procedures

    US20210019672A1