System and method for automated capture of imaging data
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-02-09
- Publication Date
- 2026-08-13
Smart Images

Figure FI2026050060_13082026_PF_FP_ABST
Abstract
Description
[0001] SYSTEM AND METHOD FOR AUTOMATED CAPTURE OF IMAGING DATA
[0002] TECHNICAL FIELD
[0003] The present disclosure relates to the field of medical diagnostics. More particularly, the present disclosure relates to otoscopic imaging technologies for examination of the ear.
[0004] BACKGROUND
[0005] Otoscopy is a diagnostic procedure involving visual inspection of the external auditory canal and tympanic membrane in order to identify anatomical landmarks and pathological conditions such as inflammation, perforation, effusion, cerumen impaction, or other abnormalities of the ear. It is routinely performed using traditional or digital otoscopes across clinical environments such as primary care, pediatrics, otolaryngology, emergency medicine, and telemedicine. The diagnostic value of otoscopy depends on image quality, illumination, focus, orientation, and consistent visualization of relevant anatomical landmarks.
[0006] Known otoscopic systems include traditional optical otoscopes and digital otoscopes capable of capturing still images or video. While digital otoscopes improve visualisation compared to optical devices, existing systems largely rely on manual positioning, manual activation of image capture, and manual selection, storage, and documentation of images. As a result, image quality and diagnostic reliability remain highly operatordependent and subject to variability caused by hand stability, experience level, patient movement, illumination conditions, and anatomical variation.Various existing techniques seek to improve individual aspects of otoscopic imaging, such as digital capture, post-processing, or image review. However, such techniques typically address isolated components of the workflow and do not comprehensively address the range of limitations encountered during image acquisition, evaluation, documentation, and data handling in clinical practice.
[0007] For example, one known solution discloses mobile imaging and analysis of tympanic membrane images but does not provide real-time feedback to assist in positioning or control of the otoscope, thereby increasing the risk of missed detections or false positives. The existing solution also lacks structured reporting and comprehensive image storage or sharing capabilities. Another known solution describes diagnosing middle ear diseases using otoscopic images and a trained network model. However, image acquisition remains sequential and manually controlled manner, which may limit efficiency during clinical examinations.
[0008] Existing otoscopic systems further lack mechanisms to consistently support image selection, operator guidance, and uniform documentation quality across different examinations and operators.
[0009] Image enhancement techniques such as noise reduction, illumination correction, and feature extraction are known in medical imaging generally. However, their application within otoscopic workflows is limited to offline processing or requires manual intervention, which reduces their practical effectiveness during regular clinical use. Existing otoscopic systems further lack automated mechanisms for selecting optimal frames from live image streams, for providing real-time feedback to improvepositioning, or for ensuring consistent documentation quality across examinations and operators.
[0010] During otoscopic examination, image quality may be adversely affected by patient movement, anatomical variations, suboptimal illumination, cerumen obstruction, noise, motion artifacts, and operator-dependent factors such as hand stability and experience. In existing otoscopic systems, these limitations may lead to repeated image capture, which increases examination time and patient discomfort. Existing systems typically lack automated mechanisms for real-time frame selection, image quality assessment, or seamless integration of captured data into electronic medical records.
[0011] Therefore, in light of the foregoing discussion, there is a need to overcome the aforementioned limitations to enhance the consistency, reliability, and efficiency of otoscopic examinations in clinical practice.
[0012] SUMMARY
[0013] The aim of the present disclosure is to provide a method, a system, and a computer-readable storage medium for automated capture of imaging data of an ear using an otoscope for ear examination. The aim of the disclosure is achieved by a method, a system, and a computer-readable storage medium for automated capture of imaging data of an ear using an otoscope for ear examination, as defined in the appended independent claims to which reference is made. Advantageous features are set out in the appended dependent claims.The embodiments of the present disclosure substantially enable improved consistency in the capture and handling of otoscopic imaging data during ear examinations.
[0014] Additional aspects, advantages, features, and objects of the present disclosure are made apparent from the drawings and the detailed description of the illustrative embodiments construed in conjunction with the appended claims that follow.
[0015] Throughout the description and claims of this specification, the words "comprise", "include", "have", and "contain" and variations of these words, for example "comprising" and "comprises", mean "including but not limited to", and do not exclude other components, items, integers or steps not explicitly disclosed also to be present. Moreover, the singular encompasses the plural unless the context otherwise requires. In particular, where the indefinite article is used, the specification is to be understood as contemplating plurality as well as singularity, unless the context requires otherwise.
[0016] BRIEF DESCRIPTION OF THE DRAWINGS
[0017] FIG. 1 is a schematic illustration of a system for automated capture of imaging data of an ear using an otoscope for ear examination in accordance with an embodiment of the present disclosure;
[0018] FIG. 2 is an illustration of a functional block diagram representing a system for automated capture of imaging data of an ear using an otoscope in accordance with an embodiment of the present disclosure;FIGS. 3A-3C is an illustration of a process flow for automated capture of imaging data of an ear using an otoscope in accordance with an embodiment of the present disclosure;
[0019] FIG. 4A is an illustration of an example User Interface (UI) system in accordance with an embodiment of the present disclosure;
[0020] FIG. 4B is an illustration of an example exploded view of input methods of FIG. 4A in accordance with an embodiment of the present disclosure;
[0021] FIGS. 5A and 5B are illustrations of detection of an activation event in an otoscope and automated ear-side detection using a system in accordance with an embodiment of the present disclosure;
[0022] FIG. 6 is an illustration of a user interface while capturing imaging data in accordance with an embodiment of the present disclosure;
[0023] FIG. 7 is an illustration of a user interface of an examination report generated by a reporting module of a system in accordance with an embodiment of the present disclosure;
[0024] FIG. 8 is an illustration of a flowchart illustrating steps of a method for automated capture of imaging data of an ear using an otoscope for ear examination in accordance with an embodiment of the present disclosure; and
[0025] FIG. 9 is an illustration of an exploded view of a computing architecture / system according to an embodiment of the present disclosure.In the accompanying drawings, an underlined number is employed to represent an item over which the underlined number is positioned or an item to which the underlined number is adjacent. A non-underlined number relates to an item identified by a line linking the non-underlined number to the item. When a number is non-underlined and accompanied by an associated arrow, the non-underlined number is used to identify a general item at which the arrow is pointing.
[0026] DETAILED DESCRIPTION OF EMBODIMENTS
[0027] The following detailed description illustrates embodiments of the present disclosure and ways in which they can be implemented. Although some modes of carrying out the present disclosure have been disclosed, those skilled in the art would recognize that other embodiments for carrying out or practicing the present disclosure are also possible.
[0028] In a first aspect, the present disclosure provides a method for automated capture of imaging data of an ear using an otoscope for ear examination, comprising:
[0029] detecting an activation event of an otoscope using one or more sensors configured to detect at least one of: movement, orientation or proximity of the otoscope;
[0030] upon detecting the activation event of the otoscope, automatically capturing, in a first position of the otoscope, imaging data associated with the ear;
[0031] processing the imaging data associated with the ear to determine whether the imaging data complies with at least one pre-defined condition associated with the captured imaging data;providing a feedback of the processed imaging data and automatically re-capturing imaging data associated with the ear with a second position of the otoscope, if the imaging data associated with the ear does not comply with the at least one pre-defined condition associated with the captured imaging data; and
[0032] replacing the imaging data that did not comply with the at least one predefined condition with the re-captured imaging data.
[0033] The method is of advantage in that using the automated capture of the imaging data reduces examination time, and improves the quality of the imaging data (e.g., an image) through the feedback in real-time, which enhances clinical documentation using automatic report generation with improved accuracy. The method provides the technical effect of enabling automated convergence toward diagnostically usable ear imaging data by integrating activation-event detection, real-time quality evaluation, feedback, and re-capture within a single examination workflow, thereby reducing operator-dependent variability and manual intervention during otoscopic image acquisition.
[0034] The method provides an automated image capture mechanism in which the imaging data is captured in response to the detection of an activation event of the otoscope, thereby reducing the dependency on manual capture operations and improving ease of use. The automated assessment of the imaging data quality combined with real-time feedback enables accurate capturing / re-capturing of imaging data and reduces the inaccuracy in clinical documentation. The imaging data that is ultimately used for examination and documentation is the imaging data that satisfies the pre-defined condition following the automated evaluation and recapture process.In an embodiment, the method identifies anatomical regions or characteristics of the ear from the captured imaging data, including cases in which the imaging data is not explicitly associated with a particular ear at the time of capture. Image processing techniques, including noise reduction and feature extraction, may be applied to improve the accuracy and diagnostic utility of the imaging data. By automating capture, quality evaluation, and feedback-driven re-capture, the method reduces operator dependency, minimizes repeated capture attempts, and shortens examination time. By generating an examination report based on the captured imaging data and associated metadata, the method improves the consistency and reliability of clinical documentation.
[0035] The method may detect the activation event of the otoscope using the one or more sensors configured to detect movement or proximity of the otoscope relative to the ear. The activation event corresponds to a transition of the otoscope from an idle state to an active examination state, as detected by sensor data indicative of movement, orientation change, or proximity to the ear. The activation event may correspond, for example, to insertion of the otoscope tip into an ear canal, and positioning of the otoscope at an examination position, or a pre-defined motion trajectory indicative of examination intent. Upon detecting the activation event, the otoscope automatically captures imaging data when the otoscope is in a first position without requiring a dedicated manual capture command. The imaging data may comprise one or more images, video frames, or a video stream.
[0036] The captured imaging data is then processed using machine vision techniques to determine whether it complies with at least one pre-defined condition associated with diagnostic quality. If the at least one pre-definedcondition is not satisfied, the feedback is provided to an operator visually, acoustically, or haptically, and the method automatically triggers recapture when the otoscope is moved to a second position. The re-capture is performed without restarting the examination workflow, thereby allowing successive capture attempts to be carried out within a single examination session. This eliminates manual capture errors, reduces operator dependency, ensures only diagnostically relevant images are retained, and minimizes repeated examinations and patient discomfort. The re-captured imaging data is used as the imaging data for examination and documentation.
[0037] In one embodiment, the activation event of the otoscope is detected using proximity sensing. One or more proximity sensors, such as infrared sensors, capacitive sensors, ultrasonic sensors, or optical distance sensors, are configured to detect when a distal tip of the otoscope approaches or enters the ear canal within a pre-defined distance threshold. When the proximity value falls below the threshold, the condition is interpreted as an activation event and image capture or transitions the camera into an active preview or capture state is automatically initiated.
[0038] In another embodiment, the activation event is detected based on recognition of a pre-defined motion pattern of the otoscope. The sensor data from the sensors, such as accelerometers and gyroscopes, is analyzed to identify movement trajectories characteristic of an otoscopic examination. The method may apply temporal filtering, pattern matching, or machine-learning-based classifiers to distinguish examination-related motion from non-intentional movement, such as handling, transport, or repositioning between subjects.In some embodiments, the feedback is provided to guide the operator during capturing and re-capturing of imaging data when the pre-defined conditions related to image quality are not met. The feedback may be delivered through a visual indicator, an LED of the otoscope, an audible cue or a haptic. The visual indicators may be overlaid on a display of a computing device, which shows the live otoscopic image, such as focus guides, alignment markers, quality bars, or color-coded indicators representing tympanic membrane visibility. This provides real-time guidance without requiring additional hardware. The LED on the otoscope may indicate capture readiness or quality status using color or blinking patterns.
[0039] Optionally, the otoscope generates vibration signals when the imaging data quality is insufficient or when optimal positioning is achieved. The audible cue, such as tones or a spoken prompt, informs the operator of capture status or required adjustments for capturing the imaging data. In some embodiments, multiple feedback modalities are combined, for example, the visual indicator overlays may be supplemented by haptic confirmation upon achieving the required imaging data quality.
[0040] Optionally, when using the method, the at least one pre-defined condition comprises computing a tympanic membrane visibility score, determining a threshold, comparing the computed tympanic membrane visibility score with the threshold, and determining if the tympanic membrane visibility score meets the threshold for a pre-determined duration.
[0041] Optionally, when using the method, the processing the imaging data associated with the ear to determine whether the imaging data complies with at least one pre-defined image quality condition for a predeterminedduration, providing feedback based on the processing and automatically re-capturing imaging data associated with the ear without terminating the examination workflow, and using imaging data that satisfies the predefined condition as the imaging data.
[0042] The tympanic membrane visibility score may be computed using image features such as edge density, circular membrane detection, patterns, color distribution, or trained neural network outputs. The method further includes determining a threshold and verifying that the tympanic membrane visibility score meets or exceeds the threshold for a predetermined duration, thereby avoiding false positives due to motion or partial visibility.
[0043] In an embodiment, the pre-defined threshold is adjusted based on patient-specific parameters, such as age or estimated ear canal geometry. For example, lower or modified thresholds may be applied for pediatric patients or anatomically narrow ear canals. This improves imaging data quality assessment, and image capture across different patient populations without operator intervention. The pre-defined condition may be determined based on a combination of multiple image quality scores, including focus quality, illumination uniformity, and degree of occlusion. The scores may be weighted or evaluated jointly to determine the overall acceptability of the imaging data. This optional embodiment provides the technical effect of preventing acceptance of imaging data based on momentary or unstable tympanic membrane visibility by requiring sustained satisfaction of the pre-defined condition before use.
[0044] Optionally, the imaging data associated with the ear that complies with the at least one pre-defined condition is stored in a provisional validationbuffer, wherein, the provisional validation buffer is configured to hold the imaging data together with associated quality metrics comprising at least the tympanic membrane visibility score and one or more acquisition parameters.
[0045] In an embodiment, the imaging data that meets the pre-defined condition is stored in a provisional validation buffer. The provisional validation buffer temporarily holds the imaging data along with associated quality metrics, including at least the tympanic membrane visibility score and acquisition parameters such as focus setting, illumination level, and sensor orientation. Only the imaging data that complies with the predefined condition for a predetermined duration is sent from the provisional validation buffer for permanent storage or reporting. The provisional validation buffer refers to a memory area that is configured to temporarily store the captured imaging data together with associated image quality metrics while the captured imaging data is evaluated against at least one pre-defined acceptance condition, wherein the imaging data is retained in the memory area for a validation period and is used for further processing, storage, or reporting only if the pre-defined acceptance condition is satisfied. This optional embodiment provides the technical effect of enabling controlled acceptance of imaging data by separating temporary validation storage from subsequent use, thereby reducing the risk of prematurely using imaging data that does not satisfy the pre-defined condition and that way to save memory space.
[0046] Optionally, the method further comprises
[0047] detecting using one or more sensors, sensor data indicative of movement, orientation, or proximity of the otoscope;fusing the sensor data to determine context data comprising at least one of: the activation event, a trajectory of the movement, determination of an ear side; and
[0048] based on the context data, automatically transitioning a camera state of the otoscope and associating the imaging data with the determined ear side.
[0049] In this optional embodiment, the sensor data fusion enables the method to distinguish between different examination contexts and to automatically associate captured imaging data with a corresponding ear without requiring manual input. This optional embodiment provides the technical effect of reducing manual ear-side selection and state management during otoscopic examination by automatically deriving examination context and ear association from fused sensor data, thereby eliminating a source of operator labeling error that would otherwise propagate through downstream clinical documentation and reporting.
[0050] Optionally, the method further comprises determining an operational state of the otoscope and adapting the automatic transition of the camera state. This optional embodiment provides the technical effect of enabling state-aware control of camera behaviour during otoscopic examination by aligning camera state transitions with the operational state of the otoscope, thereby reducing power consumption and storage utilization by maintaining the camera in an inactive state during non-examination periods and activating image acquisition only when the otoscope is in an operational position.
[0051] In an embodiment, the sensor data indicative of movement, orientation, or proximity is continuously detected and fused to determine the contextdata. The context data may include detection of the activation event, determination of the movement trajectory, and identification of the examined ear side (left or right). Based on the context data, the camera state of the otoscope is automatically transitioned, for example, between idle, preview, capture, and validation states. The imaging data is automatically associated with the determined ear side. The operational state of the otoscope (e.g., standby, examination, storage, transmission) may also be determined, and camera behavior may be adapted accordingly.
[0052] The ear side being examined may be determined from the orientation of the operator's hand holding the otoscope. Orientation data obtained from the sensors, such as accelerometers and gyroscopes integrated in the otoscope, is analyzed to determine whether the otoscope is being held in a left-hand or a right-hand orientation relative to a reference frame. Based on this orientation, the method automatically determines whether a left ear or a right ear is under examination and associates the captured imaging data accordingly. This enables automatic ear-side identification without requiring additional dedicated sensors or manual input.
[0053] In another embodiment, a user interface may be provided to allow the operator to manually confirm or override the automatically determined ear side. The override may be provided via a touchscreen input, button, gesture, or voice command prior to finalizing storage or reporting of the imaging data. This reduces the error in associating the captured imaging data with the wrong ear.
[0054] Optionally, the imaging data is captured using at least one capture trigger. The capture trigger defines a condition under which image or videocapture is initiated during the examination workflow. This optional embodiment provides the technical effect of enabling controlled initiation of imaging data capture by defining trigger-based capture conditions within the otoscopic examination workflow, thereby preventing acquisition of imaging data during non-examination positioning of the otoscope and reducing the proportion of non-diagnostic frames that require subsequent processing and evaluation.
[0055] Optionally, when using the method, the at least one capture trigger comprises an automated trigger based on machine vision confirmation of tympanic membrane visibility, preset timer intervals, force sensing, capacitive touch detection, or temperature sensing.
[0056] In an embodiment, the imaging data is captured using one or more capture triggers. The capture triggers may include automated triggers based on machine vision confirmation of tympanic membrane visibility, preset timer intervals, force sensing, capacitive touch detection, or temperature sensing. Multiple triggers may be combined to form a trigger condition, such as tympanic membrane visibility confirmation or capacitive touch detection. The capturing of the imaging data may be triggered when a force sensor integrated in the otoscope detects that a contact force between the otoscope and the ear canal exceeds a predefined threshold. That is, the capturing of the imaging data is initiated only when the detected force is within an acceptable range for a predetermined duration. The capturing of the imaging data may be triggered based on the detection of a temperature change indicative of insertion of the otoscope into the ear canal. A temperature sensor, integrated on the otoscope, may detect a transition from ambient temperature to body temperature, which is used as a capture trigger oractivation confirmation. This optional embodiment provides the technical effect of enabling flexible and context-responsive initiation of imaging data capture by providing redundant capture initiation pathways, such that the system maintains operability when an individual sensor modality is unavailable or obstructed, for example when cerumen obscures optical confirmation or when ambient temperature conditions impair temperature-based detection.
[0057] Optionally, the method further comprises generating an examination report by compiling the imaging data and associated metadata. The examination report may be generated by compiling the imaging data and associated metadata, including ear side, quality metrics, timestamps, and acquisition parameters. This optional embodiment provides the technical effect of enabling automated consolidation of imaging data and examination-related information into a single examination report without requiring manual compilation, thereby preserving structured associations between imaging data, ear-side labels, and quality metrics during transfer to an external electronic medical record system and reducing transcription errors that arise from manual data entry.
[0058] Optionally, the method further comprises integrating the examination report with an electronic medical record system using a secure data transfer protocol. The examination report may be exported in one or more formats and integrated with an electronic medical record (EMR) system using a secure data transfer protocol, such as encrypted network communication or standardized medical data interfaces in a manner that preserves the confidentiality and integrity of the examination data. The examination report may be generated on a remote computing system, such as a cloud-based server, to which the captured imaging data andassociated metadata are transmitted securely. The report generation, storage, and access are performed remotely, thereby enabling the operator to review reports from different locations and devices.
[0059] In an embodiment, the captured imaging data and the generated examination reports are stored locally on the otoscope or an associated local computing device without depending on network connectivity. The captured imaging data may be synchronized with an external system at a later time when the network connectivity becomes available.
[0060] The examination report may be generated using a structured reporting template selected based on a clinical specialty or a patient age group. The reporting template defines standardized fields, terminology, and layout appropriate for the examination context, such as pediatric otoscopy or otolaryngology practice.
[0061] In a second aspect, the present disclosure provides a system for automated capture of imaging data of an ear using an otoscope for ear examination, comprising:
[0062] a sensor module that is configured to detect an activation event of an otoscope using one or more sensors configured to detect at least one of: movement, orientation or proximity of the otoscope;
[0063] a capture module that is configured to automatically capture, in a first position of the otoscope, imaging data associated with the ear, upon detecting the activation event of the otoscope;
[0064] an automated ear detection module configured to (i) determine an ear side based on sensor fusion of data obtained from the sensor module and (ii) perform machine vision processing on the imaging data associatedwith the ear to determine whether the imaging data complies with at least one pre-defined condition associated with the captured imaging data; and a machine vision module configured to provide a feedback of the processed imaging data and enable the capture module to automatically re-capture imaging data associated with the ear with a second position of the otoscope, if the imaging data associated with the ear does not comply with the at least one pre-defined condition associated with the captured imaging data, wherein the system replaces the imaging data that did not comply with the at least one pre-defined condition with the re-captured imaging data.
[0065] The system enables the capture module to automatically acquire imaging data associated with an ear when the otoscope is positioned in a first operational position. The capture module is actuated by at least one capture trigger, such that imaging data are acquired without requiring manual user input. The system provides the technical effect of enabling coordinated sensor-based activation, automated image capture, and machine-vision-driven validation within a single integrated architecture for otoscopic examination, thereby reducing the latency between otoscope positioning and acquisition of diagnostically usable imaging data compared to systems requiring separate manual steps for activation, capture, and quality assessment.
[0066] In an embodiment, the system comprises an image capturing device operatively associated with the otoscope for acquiring the imaging data of the ear. The capture module is configured to control the image capturing device in response to the at least one capture trigger. The image capturing device and the at least one capture trigger are communicatively coupled via a network. The network may be implemented as a wirednetwork, a wireless network, or a combination thereof, including an Internet-based network.
[0067] The at least one capture trigger may comprise an automated capture trigger. The automated capture trigger may be generated based on confirmation from a machine vision module indicating that a position and orientation of the otoscope satisfy the at least one pre-defined condition for capturing the imaging data. The automated capture trigger may be generated based on one or more of: a preset timer, a force-sensing element configured to detect contact or pressure between the otoscope and the ear, a capacitive touch sensor configured to detect handling of the otoscope, or a probe temperature sensor configured to detect ear canal temperature indicative of insertion of the otoscope to a captureready position.
[0068] Optionally, the at least one capture trigger comprises a manual capture trigger. The manual capture trigger may be initiated via one or more user controls presented on a user interface of a computing device operatively connected with the otoscope. The computing device may comprise a tablet computer, desktop computer, personal computer, electronic notebook, or smartphone. In an embodiments, the manual capture trigger may be initiated via voice commands processed by a speech recognition module, or via one or more foot pedals or buttons associated with the otoscope or the system, thereby enabling hands-free operation. The system allows switching between automated and manual capture triggers.
[0069] In an embodiment, the capture module is configured to automatically capture first imaging data associated with a first ear of a subject and second imaging data associated with a second ear of the subject. The firstear may be a right ear and the second ear a left ear, or vice versa. Upon completion of image acquisition for one ear, the system is configured to initiate image acquisition for the other ear. The first operational position of the otoscope may correspond to a right-side handling position or a leftside handling position.
[0070] During operation, the machine vision module applies an image processing algorithm, including a convolutional neural network, to determine whether the imaging data satisfy the at least one pre-defined condition. The predefined condition may comprise recognition of anatomical characteristics of a tympanic membrane. Such anatomical characteristics may include, for example, shape, colour, and pre-defined anatomical landmarks. The system is configured to generate feedback based on the detected anatomical characteristics. Optionally, the pre-defined condition comprises computing a tympanic membrane visibility score, comparing the computed score with a pre-defined threshold, and determining whether the score satisfies the threshold for a predetermined duration.
[0071] The system may apply an image enhancement technique to the captured imaging data to improve image quality. Such techniques may include contrast adjustment, sharpening, noise reduction, or correction of non-uniform illumination. By way of example, the sharpening may be performed using a 3x3 median filter. Optionally, the imaging data that satisfy the pre-defined condition are stored in a provisional validation buffer. The provisional validation buffer is configured to temporarily store the imaging data together with associated quality metrics, including at least the tympanic membrane visibility score and one or more acquisition parameters.In an embodiment, the system receives a live video stream from the image capturing device and continuously computes the tympanic membrane visibility score. The capture module is enabled to acquire the imaging data only when the computed score satisfies the pre-defined threshold. If the system determines that the imaging data do not satisfy the pre-defined condition, the system outputs feedback based on the processed imaging data and automatically initiates re-capture with the otoscope in a second position, thereby allowing immediate correction without restarting the workflow. The re-captured imaging data are then used as the imaging data for the examination. The system may further output guidance indicating a next permissible workflow action.
[0072] In an embodiment, the system supports multiple workflows, including a screening workflow, a retake workflow, and an otoscopy workflow. The workflow may be modified via the user interface based on user actions, including capturing the imaging data, viewing a live video stream, or selecting an ear for re-capture. The system dynamically updates its operational behaviour based on the selected workflow.
[0073] Optionally, the system automatically determines the ear side of the subject using sensor fusion in combination with machine learning algorithms, thereby eliminating manual ear-side selection. Sensor inputs may be received from one or more of a probe temperature sensor, an ambient light sensor, an electromagnetic positioning sensor, or one or more proximity sensors. The machine vision module may fuse the sensor inputs to validate ear-side determination, including under non-standard operating conditions.In an embodiment, the system recognizes characteristic ergonomic motion patterns corresponding to the left ear and right ear examinations by analysing tilt and rotation patterns of the otoscope. In another embodiment, the system visually identifies the ear and displays a corresponding label indicating the detected ear side. The labels may comprise, for example, "L" for the left ear and "R." for the right ear, thereby providing visual confirmation. The user interface may allow the operator to manually confirm or override the automatically determined ear side.
[0074] In an embodiment, the screening workflow comprises automated image capture, automatic ear labelling, and guided workflow steps to complete an ear examination. In an embodiment, the retake workflow comprises viewing a live video stream, manual selection of an ear, and re-capture of imaging data. In an embodiment, the otoscopy workflow comprises viewing and / or capturing imaging data using manual capture triggers.
[0075] Optionally, in the system, the one or more sensors is configured to detect at least movement, an orientation or proximity of the otoscope as sensor data. Optionally, the sensor module fuses the sensor data to determine context data comprising at least one of: the activation event, a trajectory of the movement or determination of the ear side, and based on the context data, configured to automatically transition a camera state of the otoscope and associate the imaging data with the determined ear side. In an embodiment, the system may be in a resting state and initiate a process of the ear examination upon detecting the activation event of the otoscope. In an embodiment, the system initiates the process of ear examination through at least one of: motion detection using the one or more sensors or automatic workflow which is triggered using a scheduledstep in a sequence in the process of the ear examination. This optional system configuration provides the technical effect of enabling automatic context-aware camera control and ear-side association by fusing sensor data within the system architecture, thereby reducing power consumption and processing load by transitioning the camera between active and standby states in response to detected changes in otoscope orientation and proximity to the ear canal.
[0076] In an example embodiment, the otoscope is configured to transition from a resting state to an active state (i.e. an activation event) when the otoscope is taken out of a cradle, which is detected by the one or more sensors. Upon detection of the activation event, the system initiates automated image capture by enabling an automated trigger. The automated trigger is actuated when the machine vision module analyzes a live video stream acquired by the otoscope and determines that a tympanic membrane visible in the live video stream satisfies at least one pre-defined condition.
[0077] The system may receive a user input as a voice command to initiate the capturing of the imaging data. If the imaging data does not comply with the at least one pre-defined condition, the system enables re-capturing of imaging data, without terminating the process of the ear examination. The system may capture the imaging data for both ears sequentially. Further, the system automatically compiles the captured imaging data and generates an examination report. The generated examination report is presented to a user for correction, annotation, or confirmation before finalization, and the system enables exporting of the examination report.Optionally, in the system, the automated ear detection module uses an image recognition algorithm comprising a convolutional neural network configured to automatically identify anatomical characteristics indicative of the determined ear, and utilize a motion sequence analysis to determine the anatomical characteristics. In this optional configuration, image-based feature recognition is combined with motion-based analysis to support reliable identification of anatomical context during otoscopic examination. The image recognition algorithm comprises a machine learning model configured to receive image frames from the capture module and to identify anatomical characteristics of the ear, including anatomical landmarks such as the tympanic membrane, umbo, cone of light, and malleus handle. The machine learning model outputs a tympanic membrane visibility score indicating the proportion of diagnostically relevant anatomical structures visible in the frame. In a preferred embodiment, the machine learning model comprises a convolutional neural network trained on a dataset of annotated otoscopic images. In alternative embodiments, the machine learning model comprises a vision transformer, a pre-trained foundation model finetuned on otoscopic image data, or another image classification architecture suitable for real-time anatomical landmark detection. The motion sequence analysis component processes time-series data from the MEMS accelerometer and gyroscope to determine the trajectory and orientation of the otoscope during insertion, which is correlated with the output of the machine learning model to confirm ear-side association. The image recognition algorithm operates in real-time during the examination, processing each captured frame to determine whether the pre-defined quality condition is satisfied.The system reduces noise introduced during the capturing of the imaging data by applying the image recognition algorithm, thereby improving the accuracy of the captured imaging data. Further, the method reduces the examination time by automating manual capture steps and applying realtime image processing. This optional system configuration provides the technical effect of improving anatomical context identification during otoscopic examination by combining convolutional neural network-based image recognition with motion sequence analysis, thereby enabling reliable automated ear-side determination without requiring operator input even when individual anatomical landmarks are partially occluded by cerumen or other obstructions.
[0078] Optionally, the system further comprises a user interface module configured to:
[0079] visualize a tympanic membrane visibility score and a first operational state of the otoscope determined by a state controller;
[0080] provide a prompt for a second operation state of the otoscope determined by the state controller; and
[0081] display the feedback comprising one or more image quality metrics and a determined ear side indicator.
[0082] The user interface module may be implemented as a state interface that is configured to display a tympanic membrane visibility score, an operational state of the otoscope, image quality metrics, and ear side indicators, and to prompt a next operational state of the otoscope. The user interface module provides a user interface that may display user information, ear media content / captured imaging data, and an action bar. The user information includes identifying data associated with a user, such as a name, an identification number, and a date of examination. In oneembodiment, the user is a patient. The user interface enables review and annotation of the captured imaging data via control operations accessible through the action bar.
[0083] The user interface provides navigation controls for moving forwards and backwards through the captured imaging data associated with the ear. The user interface may further include playback controls such as a play button for playing the captured imaging data, a pause button for pausing the captured imaging data, a rewind button for rewinding the captured imaging data, and a forward button for forwarding the captured imaging data. This optional system configuration provides the technical effect of enabling guided interaction between the operator and the automated capture system, wherein the user interface dynamically updates quality indicators and alignment markers in response to real-time sensor data and otoscope movement, thereby guiding the operator toward an optimal sensor positioning for image acquisition and reducing the number of recapture iterations required to satisfy the pre-defined condition.
[0084] Optionally, the system further comprises a reporting module that is configured to compile the imaging data into an examination report.
[0085] The reporting module is configured to compile captured images and videos together with relevant metadata associated with the ear in order to generate an examination report. The reporting module may automatically aggregate imaging data associated with the right ear and the left ear into a single examination report. The examination report may be generated upon completion or termination of at least one workflow of the ear examination during the examination for subsequent use.The imaging data may comprise one or more images and / or video segments of the right ear and the left ear. In one embodiment, the system enables review and annotation of the examination report using a computing device. The computing device may comprise a display unit configured to present the examination report. The display unit may be implemented using, for example, a high-resolution OLED display or an e-ink display.
[0086] The association of the imaging data corresponding to the right ear and the left ear within the examination report improves documentation accuracy and supports follow-up examinations and clinical consultations. The system may be integrated with a telemedicine platform to enable remote review and consultation. Optionally, the system may be applied in veterinary platform for examination of animal ears.
[0087] The system may be configured to export the examination report in one or more formats, including electronic health record (EHR) formats, picture archiving and communication systems (PACS), NOAH, PDF, and DICOM. In one embodiment, the system integrates with existing electronic medical record systems using secure data transfer protocols, such as SFTP, HTTPS, or FTPS. The system enables re-capture of imaging data associated with the ear at any stage of the examination workflow, including after capturing the imaging data of the right ear, after capturing the imaging data of the left ear, or after generation of the examination report. In another embodiment, the system allows restarting of the ear examination at any point, thereby providing operational flexibility. This optional system configuration provides the technical effect of enabling centralized compilation of imaging data within the system architecture to support generation of examination reports in formats compatible withheterogeneous medical record systems, and of permitting re-capture of imaging data at any stage of the examination workflow without loss of previously captured data, thereby reducing the need to restart the examination from the beginning when additional imaging data is required.
[0088] Optionally, the machine vision module fuses one or more sensor inputs from any of: a gyroscopic sensor, a micro-electromechanical systems (MEMS) sensor, proximity sensors, a probe temperature sensor, an ambient light sensor, or an electromagnetic positioning sensor. The system may detect an activation event of the otoscope based on sensed movement of the otoscope and / or inputs received via one or more user input methods, including a physical button, a voice command, or a graphical user interface. In another embodiment, the activation event is detected using machine vision processing. The machine vision processing optionally comprises a machine learning or artificial intelligence algorithm that is configured to recognize examination-related conditions from the imaging data.
[0089] In an embodiment, the MEMS sensor and the gyroscopic sensor are configured to detect a vertical-to-horizontal transition of the otoscope for detecting the activation event of the otoscope. Such a vertical-to-horizontal transition may be used as an indicator of an activation event corresponding to the initiation of an ear examination.
[0090] The activation event may be detected based on movement and / or proximity of the otoscope relative to the ear. The one or more sensors may be configured to detect at least one pre-defined motion pattern indicative of an examination action. The pre-defined motion pattern may comprise one or more parameters corresponding to pre-definedmovement and / or pre-defined proximity and may be stored in a memory. By way of example, the pre-defined movement or proximity parameters may include tilt patterns, rotational patterns, or combinations thereof.
[0091] Optionally, the otoscope communicatively connected with the system may comprise a handheld otoscope or a mountable otoscope, allowing the system to be adapted to different clinical environments and examination purposes. In an embodiment, the system is configured to detect middle ear effusion by integrating additional sensing modalities, such as temperature measurement or gas analysis, thereby providing additional diagnostic information in conjunction with the captured imaging data. This optional system configuration provides the technical effect of improving robustness of automated otoscopic examination by enabling the machine vision module to operate based on fused inputs from multiple sensor modalities, thereby cross-validating activation signals and reducing falsepositive activation events caused by ambient interference on a single sensor channel.
[0092] In a third aspect, the present disclosure provides a computer-readable storage medium having instructions stored thereon which, when executed by a processor, cause the processor to perform a method for automated capture of imaging data of an ear using an otoscope for ear examination, the method comprising:
[0093] detecting an activation event of an otoscope using one or more sensors configured to detect at least one of: movement, orientation or proximity of the otoscope;
[0094] upon detecting the activation event of the otoscope, automatically capturing, in a first position of the otoscope, imaging data associated with the ear;processing the imaging data associated with the ear to determine whether the imaging data complies with at least one pre-defined condition associated with the captured imaging data;
[0095] providing a feedback of the processed imaging data and automatically re-capturing imaging data associated with the ear with a second position of the otoscope, if the imaging data associated with the ear does not comply with the at least one pre-defined condition associated with the captured imaging data; and
[0096] replacing the imaging data that did not comply with the at least one pre-defined condition with the re-captured imaging data.
[0097] The technical effects of the method as described above apply correspondingly to the computer-readable storage medium.
[0098] The advantages of the present system, the present method and the associated computer-readable storage medium, are thus disclosed above, wherein automated capture and processing of the imaging data reduce examination time and improve the quality and accuracy of the captured imaging data by providing real-time feedback during image capturing.
[0099] Embodiments of the present disclosure detect the anatomical characteristics of the ear from the captured imaging data to determine the ear side, including cases in which the imaging data is not explicitly associated with a particular ear at the time of capture. Embodiments of the present disclosure reduce noise and enhance the imaging data by applying image processing and filtering techniques, thereby improving the accuracy of the captured imaging data.
[0100] Embodiments of the present disclosure reduce manual intervention and overall examination time by automating capture steps and applying real-time image processing. Embodiments of the present disclosure improve operational accuracy and flexibility during the ear examination by providing immediate feedback and by automatically re-capturing or refining the imaging data when the pre-defined conditions are not satisfied. Embodiments of the present disclosure enhance clinical documentation by automatically generating the examination report with increased accuracy.
[0101] Some aspects of the invention:
[0102] A method for automated otoscopic imaging, comprises:
[0103] detecting an activation event of an otoscope using one or more sensors configured to detect movement or proximity of the otoscope; initiating a right ear examination by capturing image and / or video data via at least one trigger selected from manual triggers including a touch screen button, foot pedal activation, and voice commands, and automatic triggers including machine vision confirmation of tympanic membrane visibility, preset timer intervals, force sensing, capacitive touch detection, and temperature sensing;
[0104] initiating a left ear examination subsequent to the right ear examination by capturing image and / or video data using a similar trigger mechanism;
[0105] compiling a report comprising the captured image and / or video data and associated metadata;
[0106] exporting the compiled report in one or more formats selected from EHR, PACS, NOAH, PDF, and DICOM; and
[0107] providing internal undo / redo functionality and restart capability at one or more stages of the examination process.
[0108] A system for automated otoscopic imaging, comprising:a sensor module configured to detect an activation event using one or more sensors configured to detect movement or proximity of the otoscope;
[0109] a capture module configured to record image and / or video data during an ear examination;
[0110] an automated ear detection module configured to determine ear orientation (right or left) based on sensor fusion of data from the sensor module and machine vision analysis of captured images;
[0111] a user interface module configured to receive multi-modal inputs including touch, voice commands, foot pedal inputs, and motion gestures, and to display a main viewport, navigation controls, and status indicators reflecting system state and active input modes;
[0112] a reporting module configured to compile the captured data and metadata into a report and export the report in one or more formats selected from EHR, PACS, NOAH, PDF, and DICOM; and
[0113] internal undo / redo mechanisms enabling reversal of capture steps and workflow transitions at various stages of the examination.
[0114] A computer-readable storage medium having instructions stored thereon which, when executed by a processor, cause the processor to perform a method for automated otoscopic imaging, the method comprising:
[0115] detecting an activation event using one or more sensors configured to detect movement or proximity of the otoscope;
[0116] capturing right ear image and / or video data via one or more trigger mechanisms selected from manual triggers including a touch screen button, foot pedal activation, and voice commands, and automatic triggers including machine vision confirmation, preset timer intervals, force sensing, capacitive touch detection, and temperature sensing;capturing left ear image and / or video data following the right ear capture using similar trigger mechanisms;
[0117] compiling the captured data and associated metadata into a report; and
[0118] exporting the report in one or more selectable formats including EHR, PACS, NOAH, PDF, and DICOM,
[0119] wherein the method further comprises performing internal undo / redo operations and restart functions at predetermined stages of the examination.
[0120] The method of claim 1, further comprising providing real-time image enhancement and feedback during the image capture process.
[0121] Optionally, detecting the activation event further comprises detecting a vertical-to-horizontal transition of the otoscope from a resting state. Optionally, the sensor module includes a MEMS accelerometer and a gyroscope configured to detect the vertical-to-horizontal transition for device activation.
[0122] Optionally, the automated ear detection module employs image recognition algorithms, including but not limited to convolutional neural networks, to identify anatomical features indicative of a right ear or a left ear and utilizes motion sequence analysis to confirm ear orientation. Optionally, wherein the detecting step utilizes sensor fusion from accelerometers, gyroscopes, and machine vision analysis to automatically determine ear orientation.
[0123] Optionally, the capture trigger comprises manual triggers including a touch screen button, foot pedal activation, and voice commands.
[0124] Optionally, the capture trigger comprises automatic triggers including machine vision confirmation, preset timer intervals, force sensing, capacitive touch detection, and temperature sensing.Optionally, the user interface module displays a main viewport with realtime feedback that includes image quality metrics, ear orientation indicators, and active input method status.
[0125] The system of claim 2, further comprising a reporting module configured to automatically generate a comprehensive report that includes video and still images captured for both ears and to export the report in multiple formats including EHR, PACS, NOAH, PDF, and DICOM.
[0126] The method of claim 1, further comprising automatically generating a comprehensive report that includes captured video and still images for both ears and integrating the report with electronic medical record systems via secure data transfer protocols.
[0127] The system of claim 2, further comprising additional sensor inputs including a probe temperature sensor, an ambient light sensor, and an electromagnetic positioning sensor to enhance ear-side detection accuracy.
[0128] The aforementioned steps are only illustrative, and other alternatives can also be provided where one or more steps are added, one or more steps are removed, or one or more steps are provided in a different sequence without departing from the scope of the claims herein.
[0129] DETAILED DESCRIPTION OF THE DRAWINGS
[0130] FIG. 1 is a schematic illustration of a system 100 for automated capture of imaging data of an ear using an otoscope 112 for ear examination in accordance with an embodiment of the present disclosure. The system 100 comprises a sensor module 102, a capture module 106, an automated ear detection module 108, and a machine vision module 110.
[0131] The system 100 is communicatively connected with an otoscope 112.The sensor module 102 comprises one or more sensors 104A-N configured to detect an activation event of the otoscope 112 using the one or more sensors 104A-N. The one or more sensors 104A-N may detect at least one of movement, orientation or proximity of the otoscope 112, and generate corresponding sensor data. In one embodiment, the sensor module 102 fuses the sensor data to determine a context comprising at least the activation event, a movement trajectory of the otoscope 112 or an ear side.
[0132] Upon detection of the activation event, the capture module 106 is configured to automatically capture imaging data associated with an ear when the otoscope 112 is in a first position. The ear may be a right ear or a left ear. Based on the determined context, the system 100 automatically transitions a camera state of the otoscope 112 and associates the captured imaging data with the corresponding ear side.
[0133] The automated ear detection module 108 is configured to determine the ear side based on sensor fusion data obtained from the sensor module 102, and to perform machine vision processing on the captured imaging data to determine whether the imaging data satisfies at least one predefined image quality condition. The machine vision module 110 provides feedback of the processed imaging data and enables the capture module 106 to automatically re-capture imaging data with the otoscope 112 in a second position, if the imaging data associated with the ear does not comply with the at least one pre-defined condition associated with the captured imaging data. Then, the system 100 uses the re-captured imaging data as the imaging data for further processing.Referring next to FIG. 2, there is shown an illustration of a functional block diagram representing a system 200 for automated capture of imaging data of an ear using an otoscope 202 in accordance with an embodiment of the present disclosure. The system 200 comprises a sensor module 204, a capture module 208, an automated ear detection module 210, a user interface module 214, a reporting module 216, and a triggering module 218, which are implemented in hardware, software, or a combination thereof. The system 200 is communicatively coupled to the otoscope 202.
[0134] The sensor module 204 comprises one or more sensors 206A-N configured to detect an activation event of the otoscope 202. The capture module 208 captures imaging data associated with an ear upon detecting the activation event of the otoscope 202. The triggering module 218 enables the capture module 208 to capture the imaging data upon detecting the activation event of the otoscope 202. The triggering module 218 may perform an automated trigger and a manual trigger. In an embodiment, the automated trigger enables the capture module 208 to capture the imaging data automatically based on one or more detected parameters, while the manual trigger enables the capture module 208 to capture the imaging data upon receipt of at least one trigger input (e.g., a user input).
[0135] The automated ear detection module 210 comprises an image recognition algorithm 212 such as a convolutional neural network, configured to automatically identify anatomical features indicative of the ear, and to determine an ear side using motion sequence analysis. The user interface module 214 operates as a state interface that is configured to display a tympanic membrane visibility score, an operational state of the otoscope202, image quality metrics, and ear side indicators, and to prompt a next operational state. The reporting module 216 is configured to compile the captured imaging data into an examination report.
[0136] Referring next to FIGS. 3A-3C, there is shown an illustration of a process flow for automated capture of imaging data of an ear using an otoscope in accordance with an embodiment of the present disclosure. At step 302, the system is in an idle state. At step 304, the system detects an activation event of the otoscope using one or more sensors. The activation event may be determined using micro-electromechanical systems (MEMS) sensors, accelerometers, camera movement, proximity sensing, or machine vision, or user inputs such as buttons, voice commands, motion controls, or user interface interactions. At step 306, the system automatically captures imaging data of an ear using one or more triggers 308. The one or more triggers 308 may be an automatic trigger.
[0137] When the system is triggered using the one or more triggers 308, the system initiates either a right ear capture method 310 or a left ear capture method 312. In the right-ear capture method 310, imaging data of a right ear is captured in a first position of the otoscope at step 314 and evaluated against at least one pre-defined condition at step 316. If the at least one pre-defined condition is not satisfied, imaging data is automatically re-captured in a second position of the otoscope at step 318, and the evaluation is repeated at the step 316. Once the at least one pre-defined condition is satisfied, the re-captured imaging data is used as the imaging data and the imaging data is accepted at step 320.
[0138] In the left ear capture method 312, imaging data of a left ear is captured in a first position of the otoscope at step 322 and evaluated against atleast one pre-defined condition at step 324. If the at least one pre-defined condition is not satisfied, imaging data is automatically re-captured in a second position of the otoscope at step 326, and the evaluation is repeated at the step 324. Once the at least one pre-defined condition is satisfied, the re-captured imaging data is used as the imaging data and the imaging data is accepted at step 328. In an embodiment, the system moves to the idle position at the step 302, when there is no event detected. Following completion of one ear, the system may initiate capture of the other ear.
[0139] At step 330, the system compiles a report. In an embodiment, the report is an examination report comprising the imaging data of the right ear and the imaging data of the left ear. At step 332, the system exports the report in a format such as electronic health record (EHR) formats, picture archiving and communication systems (PACS), NOAH, PDF, or DICOM. Then, the system moves to the idle position.
[0140] In FIG. 4A, there is shown an illustration of an example User Interface (UI) system 400 in accordance with an embodiment of the present disclosure. The UI system 400 includes display elements 402, input methods 410, UI states 420, and available actions 428. The display elements 402 comprise a main viewport 404, navigation controls 406, and status indicators 408. The main viewport 404 displays a live otoscope feed, captured images, and an examination report.
[0141] The navigation controls 406 are aligned with a physical layout to support muscle memory. The status indicators 408 indicates a system state, active input methods, and a progress of an ear examination. In an embodiment, the status indicators 408 can be a status bar. The inputmethods 410 comprise at least one of motion controls 412, voice commands 414, foot pedals 416, and touch controls 418, to capture or re-capture imaging data, navigate in the imaging data, switch from a first ear to a second ear or vice versa, and review the imaging data. In an embodiment, the motion controls 412, the voice commands 414, the foot pedals 416, and the touch controls 418 are received from a user.
[0142] The III states 420 comprise active controls 422 that highlights currently available actions, inactive controls 424 that hides or mutes irrelevant options, and feedback (e.g. visual, auditory, or haptic feedback) 426 that provides immediate confirmation of inputs from the input methods 410. The available actions 428 comprise navigation actions 430 that manage progression of a workflow, capture actions 432 that control capturing of the imaging data, and save or delete functions, and review actions 434 that enable review, annotation, saving, deletion, and export of the imaging data. In an embodiment, the capturing of the imaging data comprises capturing of images or videos.
[0143] In FIG. 4B, there is shown an illustration of an example exploded view of the input methods 410 of FIG. 4A in accordance with an embodiment of the present disclosure. The input methods 410 comprise the motion controls 412, the voice commands 414, the foot pedals 416, and the touch controls 418, each configured to control one or more operations of an otoscope. The motion controls 412 are configured to receive inputs from one or more motion sensors, including a gyroscope 436, an accelerometer 438, and motion gestures 440. In one embodiment, the gyroscope 436 and the accelerometer 438 detect orientation, tilt, and rotational movement of the otoscope, while the motion gestures 440 correspond to pre-defined gestures / movement patterns. In anembodiment, the detected motion inputs are mapped to control operations of the otoscope, including automated capture of imaging data, switching between ear sides, view adjustments, and real-time feedback relating to the imaging data.
[0144] The voice commands 414 comprise a voice trigger 442, and a command set 444 for controlling the operations of the otoscope. In an embodiment, the voice trigger 442 comprises a wake word for activating voice control functionality. In another embodiment, the command set 444 comprises a standardized command set of pre-defined voice commands associated with corresponding otoscope operations. The foot pedals 416 comprise a left pedal 446, a center pedal 448, and a right pedal 450, configured to enable navigation, capture and review of the imaging data.
[0145] In an embodiment, the left pedal 446 enables leftward movement or selection, the right pedal 450 enables rightward movement or selection, and the center pedal 448 initiates capturing of the imaging data. In another embodiment, the left and right pedals assist in positioning the otoscope for accurate image capture. The touch controls 418 comprise screen buttons 452, and touch gestures 454, displayed on a user interface, for controlling operations of the otoscope.
[0146] In FIGS. 5A and 5B, there are shown illustrations of detection of an activation event in an otoscope 502 and automated ear-side detection using a system in accordance with an embodiment of the present disclosure. Referring to FIG. 5A, a subject 504 with a right ear 506 and a left ear 508 is illustrated. In an embodiment, the subject 504 is a human subject. The system is configured to detect an activation event when the otoscope 502 is lifted from a rest position and transitions froma vertical orientation to a horizontal orientation, as indicated at step 510.
[0147] In an embodiment, the vertical orientation corresponds to an idle state, and the horizontal orientation corresponds to an active state. In another embodiment, the activation event is detected using one or more motion sensors, including accelerometers and gyroscopic sensors.
[0148] The system is further configured to detect characteristic ergonomic movement arcs of the otoscope 502 to determine whether imaging is being performed on the right ear 506 or the left ear 508. In an embodiment, ear-side determination is performed by analyzing tilt and rotational patterns of the otoscope 502. The system determines that the otoscope is oriented toward the right ear 506 when the otoscope is tilted and rotated toward a right-side direction, as indicated at step 512.
[0149] Similarly, the system determines that the otoscope is oriented toward the left ear 508 when the otoscope is tilted and rotated toward a left-side direction, as indicated at step 514.
[0150] Referring to FIG. 5B, the system provides a visualization on the otoscope 502 comprising ear illustrations with labels and recommended grip and approach angles for both the left and right ears, as indicated at step 516.
[0151] The labels may include "R" for the right ear 506 and "L" for the left ear 508, as shown in a view 518. In an embodiment, the labels provide visual confirmation of the automatically detected ear side.
[0152] The system may further provide visual cues, including steady illumination, pulsing patterns, flashing sequences, and colour transitions, to indicate system state and confirm the image capture process. In an embodiment, the colour transitions comprise red for the right ear and blue for the left ear, or vice versa. In another embodiment, the system receives an inputfrom one or more additional sensors to assist in ear-side detection under non-standard conditions. Preferably, the one or more sensors may be a proximity sensor, an ambient light sensor, and a probe temperature sensor.
[0153] Referring next to FIG. 6, there is shown an illustration of a user interface while capturing imaging data in accordance with an embodiment of the present disclosure. The user interface comprises a main viewport 602 configured to display a live otoscope feed, a status bar 604 indicating an active ear side and a status of one or more input methods, and navigation controls 606. The navigation controls 606 enable initiation of at least one of: capturing of imaging data using a capture option 608, review of the imaging data using a review option 610, or execution of context-sensitive actions through an action bar 612.
[0154] In an embodiment, the status bar 604 comprise visual status indicators including an active ear indicator, where "R" indicates the right ear and "L" indicates the left ear. The status bar 604 further comprises input method status indicators, including a voice command status indicator 4231, a motion input status indicator 4241, and a foot pedal input status indicator 4220. The input method status indicators are updated by the system in response to changes in the operational state of the otoscope, such that input methods appropriate to the current examination context are indicated as active and input methods that are unavailable or contextually inappropriate are indicated as inactive. For example, the foot pedal input status indicator 4220 may transition from an inactive state to an active state when the system detects that the otoscope has been lifted from a resting position, indicating that the operator's hands are occupied and that hands-free input via the foot pedal is available. In an embodiment,the navigation controls 606 further comprise an export option 614 for exporting the imaging data. In another embodiment, the export of the imaging data is enabled only after the imaging data has been captured for both the right ear and the left ear.
[0155] In FIG. 7, there is shown an illustration of a user interface of an examination report generated by a reporting module of a system in accordance with an embodiment of the present disclosure. The user interface displays user information 702, right ear media contents (e.g., captured imaging data of the right ear) 704, and left ear media contents e.g., captured imaging data of the left ear) 706, and an action bar 708. The user information 702 includes identifying data associated with a user, such as a name, an identification number, and a date of examination. In one embodiment, the user is a patient.
[0156] The right ear media contents 704 include an image view 710 configured to display a still image of a right ear and a video playback region 712 configured to display a video of the right ear. The user interface enables review and annotation of the right ear media contents 704 via control operations accessible through the action bar 708. The left ear media contents 706 comprises an image view 714 displaying a still image of a left ear and a video playback region 716 displaying a video of the left ear. The user interface further enables review and annotation of the left ear media contents 706 through the action bar 708.
[0157] In one embodiment, the image view 710 of the right ear and the image view 714 of the left ear each display corresponding imaging data associated with the respective ear. In another embodiment, the user interface provides navigation controls (718, 720) for moving forwardsand backwards through imaging data associated with the right ear. Likewise, navigation controls (722, 724) enable selection of different imaging data associated with the left ear.
[0158] In an embodiment, the video playback region 712 and the video playback region 716, each present corresponding video data.
[0159] The user interface may further include playback controls such as a play button 726 for playing the video, a pause button 728 for pausing the video, a rewind button 730 for rewinding the video, and a forward button 732 for forwarding the video along a timeline 734 of the video playback region 712 associated with the right ear. Similarly, the user interface may further include playback controls such as a play button 736 for playing the video, a pause button 738 for pausing the video, a rewind button 740 for rewinding the video, and a forward button 742 for forwarding the video along a timeline 744 of the video playback region 716 associated with the left ear.
[0160] In an embodiment, the right ear media contents 704 and the left ear media contents 706 are associated with metadata, including timestamps, ear side, and examination-related information. The action bar 708 may further include controls for exporting the examination report.
[0161] Referring next to FIG. 8, there is shown an illustration of a flowchart illustrating steps of a method for (namely, a method of) automated capture of imaging data of an ear using an otoscope for ear examination, in accordance with an embodiment of the present disclosure. At step 802, an activation event of an otoscope is detected using one or more sensors that are configured to detect movement or proximity of the otoscope. At step 804, imaging data associated with the ear is automatically capturedin a first position of the otoscope, upon detecting the activation event of the otoscope. At step 806, the imaging data associated with the ear is processed to determine whether the imaging data complies with at least one pre-defined condition associated with the captured imaging data. At step 808, a feedback of the processed imaging data is provided and imaging data associated with the ear is automatically re-captured with a second position of the otoscope, if the imaging data associated with the ear does not comply with the at least one pre-defined condition associated with the captured imaging data. At step 810, the re-captured imaging data is used as the imaging data.
[0162] In FIG. 9, there is shown an illustration of an exploded view of a system, including a computing architecture, in accordance with an embodiment of the present disclosure. The exploded view depicts a system that comprises at least one input interface 902, a control module that comprises a data processing arrangement 904, a memory 906 and a nonvolatile storage 908, processing instructions 910, a shared / distributed storage 912, and an image capturing device that comprises a processor 914, a memory 916 and a non-volatile storage 918 and an output interface 920. The functions of the data processing arrangement 904, the least one input interface 902 are as described above and cooperate together to implement methods of the disclosure as described in the foregoing.
[0163] Modifications to embodiments of the present disclosure described in the foregoing are possible without departing from the scope of the present disclosure as defined by the accompanying claims. Expressions such as "including", "comprising", "incorporating", "have", "is" used to describe and claim the present disclosure are intended to be construed in a non-exclusive manner, namely allowing for items, components or elements not explicitly described also to be present. Reference to the singular is also to be construed to relate to the plural.
Claims
CLAIMS1. A method for automated capture of imaging data of an ear using an otoscope for ear examination, comprising:detecting an activation event of an otoscope using one or more sensors (104A-N, 206A-N) configured to detect at least one of: movement, orientation or proximity of the otoscope (112, 202, 502); upon detecting the activation event of the otoscope, automatically capturing, in a first position of the otoscope, imaging data associated with the ear;processing the imaging data associated with the ear to determine whether the imaging data complies with at least one pre-defined condition associated with the captured imaging data;providing a feedback of the processed imaging data and automatically re-capturing imaging data associated with the ear with a second position of the otoscope, if the imaging data associated with the ear does not comply with the at least one pre-defined condition associated with the captured imaging data; andreplacing the imaging data that did not comply with the at least one pre-defined condition with the re-captured imaging data.
2. The method of claim 1, wherein the at least one pre-defined condition comprises computing a tympanic membrane visibility score, determining a threshold, comparing the computed tympanic membrane visibility score with the threshold, and determining if the tympanic membrane visibility score meets the threshold for a predetermined duration.
3. The method of claim 1 or 2, wherein the imaging data associated with the ear that complies with the at least one pre-defined condition is storedin a provisional validation buffer, wherein the provisional validation buffer is configured to hold the imaging data together with associated quality metrics comprising at least the tympanic membrane visibility score and one or more acquisition parameters.
4. The method of any of the preceding claims, wherein the method further comprisesdetecting, using one or more sensors (104A-N, 206A-N), sensor data indicative of movement, orientation, or proximity of the otoscope (112, 202, 502);fusing the sensor data to determine context data comprising at least one of: the activation event, a trajectory of the movement, determination of an ear side; andbased on the context data, automatically transitioning a camera state of the otoscope (112, 202, 502) and associating the imaging data with the determined ear side.
5. The method of claim 4, wherein the method further comprises determining an operational state of the otoscope (112, 202, 502) and adapting automatic transition of the camera state.
6. The method according to any of the preceding claims, wherein the imaging data is captured using at least one capture trigger.
7. The method of claim 6, wherein the at least one capture trigger comprises an automated trigger based on machine vision confirmation of tympanic membrane visibility, preset timer intervals, force sensing, capacitive touch detection, or temperature sensing.
8. The method according to any of the preceding claims, wherein the method comprises generating an examination report by compiling the imaging data and associated metadata.
9. The method of claim 8, wherein the method further comprises integrating the examination report with an electronic medical record system using a secure data transfer protocol.
10. A system (100, 200) for automated capture of imaging data of an ear using an otoscope (112, 202, 502) for ear examination, comprising: a sensor module (102, 204) that is configured to detect an activation event of an otoscope (112, 202, 502) using one or more sensors (104A-N, 206A-N) configured to detect at least one of: movement, orientation or proximity of the otoscope (112, 202, 502); a capture module (106, 208) that is configured to automatically capture, in a first position of the otoscope (112, 202, 502), imaging data associated with the ear, upon detecting the activation event of the otoscope (112, 202, 502);an automated ear detection module (108, 210) configured to (i) determine an ear side based on sensor fusion of data obtained from the sensor module (102, 204) and (ii) perform machine vision processing on the imaging data associated with the ear to determine whether the imaging data complies with at least one pre-defined condition associated with the captured imaging data; anda machine vision module (110) configured to provide a feedback of the processed imaging data and enable the capture module (106, 208) to automatically re-capture imaging data associated with the ear with a second position of the otoscope (112, 202, 502), if the imaging data associated with the ear does not comply with the at least one pre-definedcondition associated with the captured imaging data, wherein the system (100, 200) replaces the imaging data that did not comply with the at least one pre-defined condition with the re-captured imaging data.
11. The system (100, 200) of claim 10, wherein the one or more sensors (104A-N, 206A-N) is configured to detect at least movement, an orientation or proximity of the otoscope (112, 202, 502) as sensor data, wherein the sensor module (102, 204) fuses the sensor data to determine context data comprising at least one of: the activation event, a trajectory of the movement or determination of the ear side, and based on the context data, configured to automatically transition a camera state of the otoscope (112, 202, 502) and associate the imaging data with the determined ear side.
12. The system (100, 200) of claim 10 or 11, wherein the automated ear detection module (108, 210) uses an image recognition algorithm (212) comprising a convolutional neural network configured to automatically identify anatomical characteristics indicative of the determined ear, and utilize a motion sequence analysis to determine the anatomical characteristics.
13. The system (100, 200) according to any of the claims 10 to 12, wherein the system (100, 200) further comprises a user interface module (214) configured to:visualize a tympanic membrane visibility score and a first operational state of the otoscope (112, 202, 502) determined by a state controller;provide a prompt for a second operation state of the otoscope (112, 202, 502) determined by the state controller; anddisplay the feedback comprising one or more image quality metrics and a determined ear side indicator.
14. The system (100, 200) according to any of the claims 10 to 13, wherein the system (100, 200) further comprises a reporting module (216) that is configured to compile the imaging data into an examination report.
15. The system (100, 200) of claim 10, wherein the machine vision module (110) fuses one or more sensor inputs from any of: a gyroscopic sensor, a micro-electromechanical systems (MEMS) sensor, a probe temperature sensor, an ambient light sensor, or an electromagnetic positioning sensor.
16. A computer-readable storage medium having instructions stored thereon which, when executed by a processor, cause the processor to perform a method for automated capture of imaging data of an ear using an otoscope for ear examination, the method comprising:detecting an activation event of an otoscope using one or more sensors (104A-N, 206A-N) configured to detect at least one of: movement, orientation or proximity of the otoscope (112, 202, 502); upon detecting the activation event of the otoscope, automatically capturing, in a first position of the otoscope, imaging data associated with the ear;processing the imaging data associated with the ear to determine whether the imaging data complies with at least one pre-defined condition associated with the captured imaging data;providing a feedback of the processed imaging data and automatically re-capturing imaging data associated with the ear with a second position of the otoscope, if the imaging data associated with theear does not comply with the at least one pre-defined condition associated with the captured imaging data; andreplacing the imaging data that did not comply with the at least one pre-defined condition with the re-captured imaging data.