Annotation Data Collection Using Eye-Based Tracking
Gaze-based tracking methods create annotated training datasets for AI systems by mapping user interactions and gaze to sample images, addressing the lack of explicit annotations in current protocols and enhancing AI training with detailed spatial data.
Patent Information
- Application Number
- JP2023506110
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-07-27
- Filing Date
- 2021-07-20
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2041-07-20
AI Technical Summary
Current slide scoring protocols lack explicit annotation requirements and fail to capture detailed spatial information during pathologist's decision-making, leading to loss of valuable data for training AI systems.
A computer-implemented method for creating a training dataset using gaze-based tracking, which includes images of samples with monitored user interactions and gaze mappings to pixels, enabling the generation of heat maps and overlays that indicate gaze duration and operations performed.
Enables robust and scalable annotation data collection for AI training, capturing detailed user interactions and gaze patterns without disrupting the user's workflow, thereby improving the accuracy of AI systems in tasks like pathology reporting and quality assurance.
Smart Images

Figure 0007787150000001 
Figure 0007787150000002 
Figure 0007787150000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates generally to methods, systems, and apparatus for performing annotation data collection, and more particularly to methods, systems, and apparatus for performing annotation data collection using gaze-based tracking, in some cases for training artificial intelligence ("AI") systems (which may include, but are not limited to, at least one of neural networks, convolutional neural networks ("CNNs"), learning algorithm-based systems, or machine learning systems, etc.).
[0002] [Related projects] This application claims priority to U.S. Provisional Patent Application No. 63 / 057,105, filed July 27, 2020, the entire disclosure of which is incorporated herein by reference.
[0003] [Copyright Notice] A portion of the disclosure of this patent document contains material that is subject to copyright protection. The copyright owner has no objection to the copying by anyone of this patent document or the patent disclosure as it appears in the U.S. Patent and Trademark Office patent files or records, but otherwise reserves all copyright rights whatsoever. [Background technology]
[0004] Thousands of stained histopathology slides are viewed and scored daily in clinical and research laboratories. Traditionally, such slides have been scored under a microscope; more recently, slides have been scanned and scored on a display screen. Developing digital analysis methods for scanned slides to assist pathologists requires access to a large number of pathologists' annotations to train algorithms (e.g., algorithms that identify regions of interest, diagnoses, treatments, etc.), including deep learning, machine learning, or other algorithms. However, current slide scoring protocols (either under a microscope or on a display screen) lack explicit annotation requirements or non-obtrusive annotation capabilities. As a result, in some cases, a significant amount of expert annotations (i.e., which precise regions on the slide guided the pathologist's decision) are not recorded and are therefore lost. Some conventional techniques address this problem by collecting a pathologist's region of interest ("ROI") on a glass slide using a microscope equipped with a video camera that tracks and records the pathologist's field of view ("FOV") while examining the slide. This information can later be aligned with a whole slide image ("WSI") digital slide and used to train a convolutional neural network ("CNN") on the diagnostic or treatment regions of interest in the WSI. While this method collects annotations in a non-intrusive manner during the pathologist's daily work, it still lacks valuable information about the specific cells or structures the pathologist was focusing on within the FOV.
[0005] Furthermore, current slides are routinely diagnosed and scored without providing any local information to support the pathologist's decision, while collecting the detailed spatial annotations required to develop AI-based algorithms is costly and time-consuming. Summary of the Invention [Problem to be solved by the invention]
[0006] Therefore, there is a need for a more robust and scalable solution for performing annotation data collection, and more particularly, there is a need for methods, systems, and apparatus for performing annotation data collection using gaze-based tracking, in some cases for training AI systems. [Means for solving the problem]
[0007] According to a first aspect, a computer-implemented method includes automatically creating a training dataset including a plurality of recordings, wherein one recording includes an image of a sample of an object, an indication of a monitored user's manipulation of an indication of the sample, and a ground truth indication of a monitored gaze of the user observing the sample on a display or through an optical device mapped to pixels of the image of the sample, the monitored gaze including at least one location of the sample being observed by the user and a time spent observing the at least one location.
[0008] In a further embodiment of the first aspect, the sample of the object is selected from the group consisting of a biological sample, a live cell culture in a microwell plate, a slide of a pathology tissue sample for generating a pathology report, a 3D radiology image, and a manufactured microarray for identification of manufacturing defects.
[0009] In a further embodiment of the first aspect, the method further comprises training a machine learning model on the training dataset to generate predicted line of sight results for a target in response to input of a target image of a target sample of a target object.
[0010] In a further embodiment of the first aspect, the ground truth representation of the monitored line of sight comprises a total time that the monitored line of sight is mapped to each particular pixel of the image over an observation time interval.
[0011] In a further embodiment of the first aspect, the ground truth representation of the monitored line of sight comprises: (i) a heat map corresponding to the image of the sample, wherein the intensity of each pixel of the heat map correlates with the total time that the monitored line of sight is mapped to each pixel, the pixels of the heat map being indicative of different actual sizes of the sample at multiple zoom levels defined by the monitored operation, and / or a panning operation of the monitored operation. and (ii) a heat map normalized to pixels located in different portions of the sample that are non-simultaneously visible on a display obtained by a panning operation of the monitored operation (i.e., pixels indicating different actual sizes of the sample at multiple zoom levels defined by the monitored operation, or pixels located in different portions of the sample that are non-simultaneously visible on a display obtained by a panning operation of the monitored operation, or both); and (ii) an overlay on the image of the sample, wherein features of the overlay correspond to the extent of the gaze and / or are indicative of the total time (i.e., correspond to the extent of the gaze, or are indicative of the total time, or both).
[0012] In a further embodiment of the first aspect, the ground truth representation of the monitored line of sight comprises an ordered time sequence that dynamically maps adaptations of the monitored line of sight of different fields being observed to different specific pixels over an observation time interval.
[0013] In a further embodiment of the first aspect, the ground truth representation of the monitored gaze is shown as at least one of (i) a directed line overlaid on pixels of the image of the sample showing dynamic adaptation of the monitored gaze, and (ii) presenting the ordered time sequence together with an indication of the time spent in each field of view.
[0014] In a further embodiment of the first aspect, the records of the training dataset further include a ground truth representation of operations performed by the user to adjust the field of view of the sample, which are mapped to a ground truth representation of the monitored line of sight and the pixels of the image.
[0015] In a further embodiment of the first aspect, the sample is observed as a magnified image thereof, and the user operations associated with the mapping of the monitored line of sight to particular pixels of the image are selected from the group comprising zooming in, zooming out, panning left, panning right, panning up, panning down, adjusting the light, adjusting the focus, and adjusting the zoom of the image.
[0016] In a further embodiment of the first aspect, the sample is observed through a microscope, and monitoring a gaze includes acquiring gaze data from at least one first camera that follows a pupil of the user while the user is observing the sample under the microscope, and the image of the sample being manipulated is captured by a second camera while the user is observing the sample under the microscope, and the computer-implemented method further includes acquiring a scanned image of the sample and aligning the scanned image of the sample with the image of the sample captured by the second camera, and mapping includes mapping the monitored gaze to pixels of the scanned image using the alignment to the image captured by the second camera.
[0017] In a further embodiment of the first aspect, the monitored line of sight is represented as a weak annotation, and the records of the training dataset further include at least one of the following additional ground truth labels of the image of the sample: when the sample comprises a sample of tissue of a subject, a pathology report created by the user observing the sample, a pathological diagnosis created by the user observing the sample, a sample score indicating a pathological evaluation of the sample created by the user observing the sample, at least one clinical parameter of the subject represented in the sample, historical parameters of the subject, and results of a treatment administered to the subject; when the sample comprises a manufactured microarray, a user-provided indication of at least one manufacturing defect, a pass / fail indication of a quality assurance test; when the sample comprises a live cell culture, cell growth rate, cell density, cell homogeneity, and cell heterogeneity; and one or more other user-provided data items.
[0018] In a further embodiment of the first aspect, the method further comprises training a machine learning model on the training dataset to generate a target predicted pathology report and / or pathology diagnosis and / or sample score result in response to input of a target image of a target biological sample of pathological tissue of the target individual and a target gaze of a target user when the sample comprises a tissue sample of the subject; to generate a target manufacturing defect and / or quality check pass / fail indication result in response to input of a target image of a target manufactured microarray when the sample comprises the manufactured microarray; and to generate a target cell growth rate, target cell density, target cell homogeneity, and target cell heterogeneity result when the sample comprises a live cell culture.
[0019] According to a second aspect, a computer-implemented method for assisting in visual analysis of a sample of an object includes providing a target image of the sample of the object to a machine learning model trained on a training dataset including a plurality of records, the records including an image of the sample of the object, an indication of a monitored user's manipulation of a presentation of the sample, and a ground truth representation of a monitored line of sight of the user observing the sample on a display or through an optical device mapped to pixels of the image of the sample, the monitored line of sight including at least one location of the sample being observed by the user and time spent observing the at least one location; and obtaining, as a result of the machine learning model, an indication of a predicted monitored line of sight of a pixel of the target image.
[0020] In a further embodiment of the second aspect, the results include a heat map of multiple pixels mapped to pixels of the target image, the intensity of the pixels of the heat map being correlated to the predicted time of gaze, and the pixels of the heat map being normalized to pixels indicating different actual sizes of the sample at multiple zoom levels determined by the monitored operation and / or pixels located in different parts of the sample that are non-simultaneously visible on a display obtained by a panning operation of the monitored operation.
[0021] In a further embodiment of the second aspect, the results include a time series showing dynamic gaze mapped to pixels of the target image over a time interval, and the computer-implemented method further includes monitoring in real time a gaze of a user observing the target image, comparing a difference between the real-time monitoring and the time series, and generating an alert when the difference exceeds a threshold.
[0022] In a further embodiment of the second aspect, the records of the training dataset further include a ground truth representation of an operation by the user mapped to a ground truth representation of the monitored gaze and the pixels of the image, and the results include a prediction of an operation for a presenter of the target image.
[0023] In a further embodiment of the second aspect, the method further comprises monitoring a user's manipulation of the sample presentation in real time, comparing a difference between the real-time monitoring of manipulation and a prediction of the manipulation, and generating an alert when the difference exceeds a threshold.
[0024] According to a third aspect, a computer-implemented method for assisting in visual analysis of a sample of an object includes providing a target image of the sample to a machine learning model and obtaining, as a result of the machine learning model, a sample score indicative of a visual evaluation of the sample, wherein the machine learning model is trained on a training dataset including a plurality of records, the records including: an image of the sample of the object; a representation of a monitored user's interaction with a presentation of the sample; a ground truth representation of a monitored gaze of the user observing the sample on a display or through an optical device mapped to pixels of the image of the sample, the monitored gaze including at least one location of the sample observed by the user and a time spent observing the at least one location; and a ground truth representation of a sample visual evaluation score assigned to the sample.
[0025] According to a fourth aspect, an eye-tracking component integrated with a microscope between an objective lens and an eyepiece comprises an optical device that directs a first set of electromagnetic frequencies reflected back from each eye of a user observing a sample under the microscope to a respective first camera that generates a representation of the user's tracked gaze, and simultaneously directs a second set of electromagnetic frequencies from the sample under the microscope to a second camera that captures an image indicative of the field of view observed by the user.
[0026] In a further embodiment of the fourth aspect, the first set of electromagnetic frequencies are infrared (IR) frequencies generated by an IR source, and the first camera is a near IR camera. the second set of electromagnetic frequencies comprises a visible light spectrum, the second camera comprises a red-green-blue (RGB) camera; the optical device comprises a beam splitter that directs the first set of electromagnetic frequencies from the IR source to an eyepiece where the user's eye is located, directs the back-reflected first set from the user's eye via the eyepiece to the NIR camera, and directs the second set of electromagnetic frequencies from the sample to the second camera and the eyepiece; the optical device that separates the electromagnetic light waves from a single optical path after reflection from two eyes into two optical paths to two of the first cameras is selected from the group consisting of: a polarizer and / or waveplate (i.e., a polarizer or waveplate or both) that directs different polarizations into different optical paths, and / or an infrared spectrum light source in conjunction with a dichroic mirror and a spectral filter, and / or applying amplitude modulation at different frequencies in each optical path for heterodyne detection.
[0027] A further understanding of the nature and advantages of particular embodiments may be realized by reference to the remaining portions of the specification and the drawings. In the drawings, like reference numerals are used to refer to like components. In some cases, a sublabel is associated with a reference numeral to represent one of multiple similar components. When a reference numeral is recited without specifying the sublabel present, it is intended to refer to all such multiple similar components. It should be noted that, as used herein, "and / or" means covering one element, any combination thereof, or the sum of two or more elements connected by the phrase. [Brief explanation of the drawings]
[0028] [Figure 1] FIG. 1 is a schematic diagram illustrating a system for implementing annotation data collection using gaze-based tracking, according to various embodiments. [Figure 2A] FIG. 1 is a schematic diagram illustrating a non-limiting example of annotation data collection using gaze-based tracking, according to various embodiments. [Figure 2B] FIG. 1 is a schematic diagram illustrating a non-limiting example of annotation data collection using gaze-based tracking, according to various embodiments. [Figure 3A] 10A-10C are schematic diagrams illustrating various other non-limiting examples of annotation data collection using gaze-based tracking, according to various embodiments. [Figure 3B] 10A-10C are schematic diagrams illustrating various other non-limiting examples of annotation data collection using gaze-based tracking, according to various embodiments. [Figure 3C] 10A-10C are schematic diagrams illustrating various other non-limiting examples of annotation data collection using gaze-based tracking, according to various embodiments. [Figure 3D] 10A-10C are schematic diagrams illustrating various other non-limiting examples of annotation data collection using gaze-based tracking, according to various embodiments. [Figure 4A] FIG. 1 is a flow diagram illustrating a method for performing annotation data collection using gaze-based tracking, according to various embodiments. [Figure 4B] FIG. 1 is a flow diagram illustrating a method for performing annotation data collection using gaze-based tracking, according to various embodiments. [Figure 4C] FIG. 1 is a flow diagram illustrating a method for performing annotation data collection using gaze-based tracking, according to various embodiments. [Figure 4D] FIG. 1 is a flow diagram illustrating a method for performing annotation data collection using gaze-based tracking, according to various embodiments. [Figure 5A] FIG. 1 is a flow diagram illustrating a method for training an AI system based on annotation data collected using gaze-based tracking, according to various embodiments. [Figure 5B]FIG. 1 is a flow diagram illustrating a method for training an AI system based on annotation data collected using gaze-based tracking, according to various embodiments. [Figure 5C] FIG. 1 is a flow diagram illustrating a method for training an AI system based on annotation data collected using gaze-based tracking, according to various embodiments. [Figure 5D] FIG. 1 is a flow diagram illustrating a method for training an AI system based on annotation data collected using gaze-based tracking, according to various embodiments. [Figure 6] FIG. 1 is a block diagram illustrating an exemplary computer or system hardware architecture, according to various embodiments. [Figure 7] FIG. 1 is a block diagram illustrating a network system of computers, computing systems, or system hardware architectures that can be used in accordance with various embodiments. [Figure 8] FIG. 1 is a block diagram of components of a system for creating a training dataset of images annotated with representations of monitored gaze and / or monitored actions (i.e., monitored gaze or monitored actions, or both), and / or training machine learning model(s) on this training dataset, according to various embodiments. [Figure 9] 1 is a flowchart of a method for automatically creating an annotated training dataset containing sample images of objects annotated with supervised gaze for training an ML model, according to various embodiments. [Figure 10] 1 is a flowchart of a method of inference by a machine learning model trained on a training dataset of images annotated with indications of monitored gaze and / or monitored manipulation, according to various embodiments. [Figure 11] 1 is a schematic diagram illustrating a heat map overlaid on an image of an observed field of view of a sample of an object, according to various embodiments. [Figure 12] 1 is a schematic diagram of components installed on a microscope for monitoring the line of sight of a user viewing a sample under the microscope, according to various embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0029] An aspect of some embodiments of the present invention relates to a system, method, computing device, and / or code instructions (stored in memory and executable by one or more hardware processors) for automatically creating an annotated training dataset for training a machine learning model (i.e., a system, method, computing device, and / or code instructions (stored in memory and executable by one or more hardware processors)). The annotated training dataset includes a plurality of records. Each record includes an image of a sample of the object, also referred to herein as a first sample (e.g., an image of a histology slide, optionally a whole slide image (WSI), or an image of a manufactured product such as a microarray), and an indication of each user's monitored gaze (sometimes referred to herein as attention data) acquired during an observation session in which the respective user is observing the sample (e.g., when the user is looking at the current field of view (FOV) seen in the eyepiece of the microscope and / or the current field of view presented on the display (i.e., the current field of view (FOV) seen in the eyepiece of the microscope, or the current field of view presented on the display, or both)), and optionally an indication of monitored actions taken by each user to adjust the presentation of the sample during the observation session. The monitored gaze represents ground truth. The monitored gaze may be represented as weak annotations of the image. The ground truth monitored gaze is mapped to pixels of the image of the respective sample. The monitored gaze includes one or more locations (e.g., regions) of the sample the user is observing and / or the time spent observing each location. Because the magnification of a sample may be so large that if the FOV represents only a portion of the entire sample, it may not be possible to adequately inspect the entire sample on the display, the user may select a different FOV and / or adjust the presentation of the FOV to visualize the sample (i.e., select a different FOV, or adjust the presentation of the FOV, or both), for example, zoom in, zoom out, pan, adjust focus, adjust light, and adjust the image zoom.The monitored gaze can be represented, for example, as a heat map, where pixels of the heat map can indicate the total viewing time during the viewing session that the user gazed on the portion of the sample corresponding to each pixel of the heat map. The pixels of the heat map can be normalized to pixels that indicate different actual sizes of the sample at different zoom levels determined by the monitored operation and / or pixels located in different portions of the sample that are non-simultaneously visible on the display resulting from a panning operation of the monitored operation. The record can include additional data. This additional data can be additional labels that represent ground truth along with the monitored gaze. Examples of additional data include a visual assessment score for the sample, which can be a result provided by a user reviewing the sample. When the sample is a tissue sample obtained from a subject, the visual assessment score can be, for example, a clinical score and / or a pathological (i.e., clinical score and / or pathological) diagnosis, e.g., a pathology report. When the sample is a manufactured product, such as a manufactured microarray, the visual assessment score can be an indication of one or more defects found in the product.
[0030] The samples may be of objects that cannot be observed in their entirety by a user, for example, samples of objects that cannot be presented at a size suitable for visual inspection on a display and / or under a microscope (i.e., on a display or under a microscope, or both). When a sample is presented at a zoom-in level suitable for visual inspection, a portion of the sample is presented on the display and / or shown under the microscope, while other portions of the sample are not presented. The user performs operations to visually examine the remainder of the sample, for example, zooming out, panning, and / or zooming in on other areas.
[0031] Examples of sample objects include: *A tissue sample, such as a sample of pathological tissue obtained as a biopsy. The tissue sample may be viewed as a prepared slide, such as a whole image slide. When such a slide is viewed under a microscope and / or on a screen at a zoom level sufficient to examine detail (e.g., a single cell, the interior of a cell, a group of cells), some portions of the image are visible, while much of the remaining image is invisible. A pathologist (or other user) visually examining the sample views the WSI or slide by panning to view different fields at different magnification levels. The pathologist examines the tissue sample to, for example, generate a pathology report, provide a clinical diagnosis, and / or calculate a clinical score that may be used to determine whether to administer chemotherapy (i.e., generate a pathology report, provide a clinical diagnosis, and / or administer chemotherapy) or other therapeutic agent. *For example, live cell cultures in microwell plates. *Other biological samples. *Radiological images, such as three-dimensional CT and / or MRI images (i.e., CT and / or MRI images). A radiologist viewing such a 3D image can view a single 2D slice at a time and can scroll back and forth along the z-axis to view upper and lower 2D slices. The radiologist can zoom in on a particular portion of the image. The radiologist can examine individual organs one at a time by repeatedly scrolling up and down to different organs. Multiple organs can be evaluated; for example, when looking for metastatic disease, the radiologist can examine each organ for the presence of tumor. The radiologist examines the 3D image to, for example, prepare a radiology report, provide a clinical diagnosis, and / or calculate a clinical score. *The object can be a manufactured product, such as a microarray (e.g., a glass slide with approximately one million DNA molecules attached in a regular pattern), a cell culture, a silicon chip, a micro-electromechanical system (MEMS), etc. A user may view the manufactured product or an image of it as part of a quality assurance process to identify manufacturing defects and / or indicate whether the product passed or failed quality assurance testing.
[0032] Optionally, an image of the FOV of the sample showing what the user is gazing at is captured using the monitored line of sight. The image of the FOV can be registered with an image of the sample, such as a WSI obtained by scanning a tissue sample slide, and / or an image of a manufactured product (e.g., a hybridized DNA microarray) captured by a camera. When the monitored line of sight is mapped to the image of the FOV, the registration between the image of the FOV and the image of the sample enables the monitored line of sight to be mapped to the image of the sample. Optionally, to enable the monitored line of sight to be mapped to the image of the sample, a display (e.g., a heat map) of the monitored line of sight corresponding to different FOVs (e.g., at different zoom levels) is normalized using data from operations (e.g., zoom level operation, pan operation, image zoom). In other words, because a magnified sample is usually very large, a user usually observes different fields of view of the sample. Each field of view can represent a portion of the sample currently shown in the microscope eyepiece and / or presented on a display. The FOV can be associated with a certain magnification. The same region of a sample can be viewed as different FOVs under different magnifications. Each FOV is mapped to an image of the sample, such as a whole-slide image of a pathology sample on a slide, and / or a larger image of a manufactured product, such as a hybridized DNA microarray. The mapping can be at the pixel level and / or at the pixel group level, allowing the user's observation location (e.g., by tracking pupil movement) to be mapped to a single pixel and / or pixel group of the FOV and / or image of the sample (e.g., WSI).
[0033] Various machine learning models can be trained on the training dataset according to the data structure of the records of the training dataset. In one example, the ML model generates a predicted target gaze result in response to input of a target image of the target sample. In another example, the ML model generates a predicted target manipulation result in response to input of a target image and / or monitored gaze of the target sample. The predicted target gaze and / or manipulation can be used, for example, to train a novice user (e.g., a pathologist) in learning how to examine and / or manipulate a new sample, and / or to guide a user on what to look at on a new sample, and / or as a form of quality assurance for a user observing a new sample to verify that the user viewed and / or manipulated the sample according to standard practice. In yet another example, the ML model generates a visual inspection result, e.g., a clinical score, a clinical diagnosis (e.g., a clinical diagnosis of a medical image such as a pathology slide and / or a 3D radiology image), and / or an indication of a defect in a manufactured product (e.g., a pass / fail quality check, where the defect is located), in response to input of the target image and / or target gaze and / or target manipulation. In yet another example, the ML model combines the target gaze and the visual inspection by generating an indication of where the features that caused the visual inspection were found in the sample, such as which region(s) of a microarray had a defect that caused a quality assurance test to fail, or which region(s) of a pathology slide were used to calculate a clinical score that indicated the patient should be treated with chemotherapy or another therapeutic agent.
[0034] The monitored gazes and / or monitored actions taken by the user may be collected in the background and do not necessarily require active user input. The monitored gazes and / or monitored actions are collected while the user is observing samples based on the user's standard practice workflow and do not interfere with and / or alter the standard practice workflow.
[0035] At least some implementations described herein address the technical problem of creating annotations of sample images of objects for training machine learning models. Annotating sample images of objects is technically challenging for several reasons.
[0036] First, each sample of an object may contain numerous details for examination. For example, a tissue sample has a large number of biological components, such as cells, blood vessels, and inter-cellular objects (e.g., nuclei), present within the sample. In another example, a manufactured microarray may have a large number of DNA molecule clusters, such as approximately one million (or other values). Training a machine learning model requires a large number of annotations. Traditionally, labeling is performed manually. The challenge is that people qualified to perform this manual task are typically trained subject matter experts (e.g., pathologists, quality assurance engineers), and these experts are in short supply and difficult to find to create a large number of labeled images. Even if such trained subject matter experts are identified, manual labeling is time-consuming because each sample image may contain thousands of features of various types and / or conditions. Some types of object features are difficult to distinguish using images, thereby requiring even more time to correctly annotate them. Moreover, manual labeling is prone to errors, for example in distinguishing between different types of cellular objects.
[0037] Second, each sample requires a full visual inspection, with additional time spent inspecting critical features. To be time-efficient, a subject matter expert needs to know how much time to spend performing a full visual inspection and when to spend additional time looking at specific features. Therefore, as described herein, a capture of the time spent at each location being observed is collected and used in recording to create a training dataset. For some objects, such as tissue samples, each sample is unique, with different structures and / or cells located in different locations and / or having different configurations. A subject matter expert has knowledge of how to inspect such samples to obtain the required visual data without missing critical features, for example, for generating a pathology report. For objects such as microarrays where DNA is arranged in a regular pattern, a subject matter expert has knowledge of how to inspect a large field of view with a regular pattern to identify abnormalities and, for example, pass / fail a quality assurance inspection. In yet another example, most people have very similar anatomical structures, so in anatomical images (e.g., 3D CT scans, MRIs), the heart, lungs, stomach, liver, and other organs are almost always located in the same relative locations. However, in most cases, a visual inspection of all of the organs is required to identify clinical features that may be unique to each organ. In some systemic diseases, different organs are affected differently as part of different pathological manifestations of the same underlying disease. This diagnosis is made by examining various visual findings.
[0038] Third, to gain a global and / or local understanding of individual features and / or the interactions between features, sample annotation requires observing the sample in various fields of view obtained using various observation parameters, such as various zoom levels, various light intensities, focus, various image scans, and / or panning across the sample. For example, in the case of microarrays, a visual inspector opens a quality control image using feature extraction and views the image at several magnifications. In addition, the inspector observes the image in both a standard scale and a logarithmic scale. The standard scale is generally used to observe bright features at the top of the image. The logarithmic scale is generally used to observe dimmer features at the bottom of the image. The inspector subjectively passes or fails slides based on the type and severity of defects he or she identifies. Examples of anomalies that result in failure include draggers, scratches, empty pockets, merging, nozzle issues, and honeycomb. At least some embodiments described herein improve machine learning techniques by automatically generating annotations of sample images (e.g., histopathology slides, 3D radiology images, microarrays, and other manufactured products) for training machine learning models.
[0039] Using standard techniques, individual samples of objects (e.g., cells) are manually annotated by a user (e.g., a pathologist). The user generates a result for the sample, e.g., a report (e.g., a tissue sample and / or radiology image report), a pass / fail for product quality assurance, etc. This result is based on the sample's features, which can serve as ground truth annotations.
[0040] In at least some embodiments, the improvement resides in monitoring the gaze of a user (e.g., a pathologist, radiologist, quality assurance technician) while performing standard practice tasks of reading a sample of an object (e.g., a pathology tissue sample, a radiology image, or the observation of a manufactured product such as a DNA microarray under a microscope and / or presented as an image on a display), and optionally monitoring the user's manipulation of the sample (e.g., panning, zoom level, focus, scaling, light). Monitoring the gaze and / or manipulation of the sample can occur without necessarily requiring active input from the user and / or can occur in the background while the user performs their task based on a standard practice workflow, without necessarily requiring an interruption and / or change in the workflow. The user's gaze is monitored and mapped to a pattern of how the user is looking, for example, first a quick scan of the entire sample, then zooming in on a specific area, zooming out to get a view of the larger tissue structure, then zooming in again, etc., by considering sample locations (e.g., pixels) that indicate where the user is looking, and / or the time spent at each observation location. The image of the sample is annotated using an indication of the monitored gaze and / or monitored operation, such as creating a heat map, where pixel intensity indicates the total observation time at the location of the sample corresponding to the pixel in the heat map. The pixels in the heat map can be normalized to pixels indicating different actual sizes of the sample at different zoom levels determined by the monitored operation and / or pixels located in different portions of the sample that are non-simultaneously visible on the display, obtained by panning the monitored operation. A weak label in the form of a visual indication (e.g., clinical score, pathological score, pathological report, clinical diagnosis, quality assurance pass / fail indication for the object, indication of defects found in the object) can be assigned to the sample based on the results manually created by the user.Other data can be included in this weak label, for example, an audio label created from an audio message recorded by, for example, an audio sensor that records a short verbal note made by a user while observing a sample, as described herein.
[0041] Before describing at least one embodiment of the invention in detail, it is to be understood that the invention is not necessarily limited in its application to the details of construction and the arrangement of components and / or methods set forth in the following description and / or illustrated in the drawings and / or examples. The invention is capable of other embodiments and of being practiced or carried out in various ways.
[0042] The present invention may be a system, method, and / or computer program product that may include computer-readable storage medium(s) having computer-readable program instructions that cause a processor to perform aspects of the present invention.
[0043] A computer-readable storage medium may be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific, non-exhaustive examples of computer-readable storage media include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, and any suitable combination thereof. As used herein, computer-readable storage media is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses passing through a fiber optic cable), or electrical signals transmitted over a wire.
[0044] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium into each computing / processing device, or may be downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards these computer-readable program instructions for storage in a computer-readable storage medium within the respective computing / processing device.
[0045] The computer-readable program instructions for carrying out the operations of the present invention may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or source or object code written in any combination of one or more programming languages. Programming languages include object-oriented programming languages such as Smalltalk, C++, and traditional procedural programming languages such as the "C" programming language or similar. The computer-readable program instructions may execute entirely on the user's computer as a standalone software package, partially on the user's computer, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be to an external computer (e.g., over the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, a field-programmable gate array (FPGA), or a programmable logic array (PLA) can execute computer-readable program instructions by utilizing state information of the computer-readable program instructions to personalize the electronic circuitry to perform aspects of the present invention.
[0046] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0047] These computer-readable program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, executed by the processor of the computer or other programmable data processing apparatus, produce means for performing the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams. These computer-readable program instructions can also be stored on a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other device to function in a particular manner, such that a computer-readable storage medium having instructions stored thereon includes an article of manufacture containing instructions that perform aspects of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.
[0048] Computer-readable program instructions may be loaded onto a computer, other programmable data processing apparatus, or other device and cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, such that the instructions executing on the computer, other programmable apparatus, or other device perform the functions / operations specified in one or more blocks of the flowcharts and / or block diagrams.
[0049] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent an instruction module, instruction segment, or instruction portion, including one or more executable instructions that implement the specified logical function(s). In some alternative implementations, the functions shown in the blocks may occur out of the order shown in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, may be implemented by a special-purpose hardware-based system that performs the specified functions or operations or executes a combination of special-purpose hardware and computer instructions.
[0050] Reference is now made to Figure 8, which is a block diagram of components of a system 800 for creating a training dataset of images annotated with indications of monitored gaze and / or monitored actions and / or training machine learning model(s) on this training dataset, according to some embodiments of the present invention. System 800 may be an alternative to and / or may be combined (e.g., using one or more components) with the systems described with reference to Figures 1, 2A-2B, 3A-3D, 6, 7, and 12.
[0051] System 800 may optionally perform the operations of the methods described with reference to Figures 4A-4D, 5A-5D, 9, 10, and 11 by hardware processor(s) 802 of computing device 804 executing code instructions 806A and / or 806B stored in memory 806.
[0052] Computing device 804 can be embodied as, for example, a client terminal, a server, a virtual server, a laboratory workstation (e.g., a pathology workstation), a quality assurance workstation, a manufacturing workstation, a procedure room (e.g., an operating room) computer and / or server, a virtual machine, a computing cloud, a mobile device, a desktop computer, a thin client, a smartphone, a tablet computer, a laptop computer, a wearable computer, a glasses computer, and a watch computer. Computing device 804 can include advanced visualization workstations that may be implemented as add-ons to laboratory workstations and / or quality assurance workstations and / or other devices that present images of samples of objects to a user (e.g., a subject matter expert).
[0053] Different architectures of the system 800 based on the computing device 804 can be implemented, for example, a central server-based implementation and / or a localized-based implementation.
[0054] In an example of a central server-based implementation, the computing device 804 may include locally stored software that performs one or more of the operations described with reference to FIGS. 4A-4D, 5A-5D, 9, 10, and 11 and / or may operate as one or more servers (e.g., network servers, web servers, computing clouds, virtual servers) that provide services (e.g., one or more of the operations described with reference to FIGS. 4A-4D, 5A-5D, 9, 10, and 11) to one or more client terminals 808 (e.g., user client terminals such as remotely located laboratory workstations, remotely located quality assurance workstations, remotely located manufacturing workstations, remote picture archiving and communication system (PACS) servers, remote electronic medical record (EMR) servers, remote sample image storage servers, remotely located pathology computing devices, desktop computers, etc.) over a network 810. The computing device 804 may, for example, provide software as a service (SaaS) to the client terminal(s) 808, provide applications for local download to the client terminal(s) 808 as add-ons to a web browser, a tissue sample image viewer application, a quality assurance image review application, and / or provide the client terminal(s) 808 with the ability to use remote access sessions via a web browser or the like.
[0055] In one embodiment, multiple gaze monitoring devices 826 monitor the gaze of each user observing a sample on the imaging device 812 (e.g., a microscope and / or a display), and optionally multiple manipulation monitoring devices 850 monitor manipulation of each sample by each user (e.g., panning, zooming in / out, adjusting light, adjusting focus, adjusting scale). Exemplary gaze monitoring devices 826 are described, for example, with reference to FIGS. 2A and 2B, 3A-3D, and 12. Images of each sample are captured (e.g., by the imaging device 812 and / or another device). The monitored gaze data and / or monitored manipulation data and / or sample images can be provided to a respective client terminal 808. Each of the multiple client terminals 808 provides the monitored gaze data and / or monitored manipulation data and / or images to the computing device 804, optionally via the network 810. The computing device can create each annotated dataset 822A by annotating sample images with corresponding ground truth monitored gaze data and / or monitored interaction data and / or other data (e.g., clinical scores) as described herein. One or more training datasets 822C can be created from the annotated dataset(s) 822A as described herein. One or more machine learning models 822B can be trained on the training dataset(s) 822C as described herein. Training of the ML model(s) 822B can be performed locally by the computing device 804 and / or remotely by another device (e.g., a server) that can provide the trained ML model(s) 822B to the computing device 804 and / or can be remotely accessed by the computing device 804.In another embodiment, the computing device 804 obtains a respective image of each sample from each of a plurality of client terminals 808, feeds each image into the trained ML model 822B, and obtains a respective result, such as a heat map indicating areas for the user to view, which is provided to the corresponding client terminal 808 for local presentation and / or user (e.g., for training and monitoring users as described herein).
[0056] In locally based embodiments, each respective computing device 804 is used by a particular user, e.g., a particular pathologist and / or a particular quality assurance technician, and / or a user group, at a facility, such as a hospital and / or pathology laboratory and / or manufacturing facility. The computing device 804 receives monitored gaze data and / or monitored operation data and / or visual assessments and / or other data (e.g., audio tags) and / or images of samples, e.g., directly and / or via an image repository, such as a server 818 (e.g., a PACS server, cloud storage, hard disk). The computing device 804 may locally generate annotated dataset(s) 822A, create training dataset(s) 822C, and / or train ML model(s) 822B, as described herein. The computing device 804 can locally feed sample images to the trained ML model(s) 822B as described herein to obtain results that can be used locally (e.g., for presentation on a display, for use in training a user, for use in guiding a user).
[0057] Sample images can be locally fed into one or more machine learning models 822B to obtain results. The results can be presented, for example, on a display 826, stored locally in a data storage device 822 of the computing device 804, and / or fed into another application that can be locally stored in the data storage device 822. The results can be used, for example, for user training, monitoring a user's work for quality assurance, and / or assisting a user, as described herein. Training of the machine learning model(s) 822B can be performed locally by each respective computing device 804 based on sample images and / or gaze data; for example, various pathology laboratories can each train their own set of machine learning models using their own sample and gaze data from their own pathologists. In another example, various manufacturing facilities can each train their own set of machine learning models using their own sample and gaze data from their own quality assurance technicians. In another example, the trained machine learning model(s) 822B are obtained from another device, such as a central server.
[0058] The computing device 804 receives images of samples of objects captured by one or more imaging devices 812. Exemplary imaging device(s) 812 include scanners and cameras. The images of the samples can be presented on a display implementation of the imaging device(s) 812. In another example, the imaging device 812 is implemented as a microscope, and the images of the samples are viewed by a user through the microscope.
[0059] The imaging device(s) 812 may generate and / or present two-dimensional (2D) images of a sample of the object, e.g., an image of the entire sample, such as an image of the entire slide in the case of a tissue sample, and / or an image of the entire microarray in the case of a fabricated microarray being evaluated for manufacturing defects. Note that the sample may represent 3D data in which features of the object at different depths are revealed by adjusting the focus.
[0060] The sample images captured by the imaging machine 812 can be stored in an image repository such as a server(s) 818, for example, a storage server (e.g., a PACS, EHR server, manufacturing server and / or quality assurance server), a computing cloud, virtual memory, and a hard disk.
[0061] The annotated dataset(s) 822A are created by annotating image(s) of the sample(s) with ground truth representations of gaze, and / or manipulation data, and / or other data, as described herein.
[0062] The training dataset(s) 822C may be created based on the annotated dataset(s) 822A as described herein.
[0063] The machine learning model(s) 822B may be trained on the training dataset(s) 822C as described herein.
[0064] The computing device 804 can receive sample images and / or monitored gaze data and / or monitored operations and / or other data from the imaging device 812 and / or gaze monitoring device 826 and / or operation monitoring device(s) 814 using one or more data interfaces 820, such as a wired connection (e.g., a physical port), a wireless connection (e.g., an antenna), a local bus, a port for connecting a data storage device, a network interface card, other physical interface implementations, and / or a virtual interface (e.g., a software interface, a virtual private network (VPN) connection, an application programming interface (API), a software development kit (SDK)). Alternatively or additionally, the computing device 804 can receive sample images and / or monitored gaze data and / or monitored operations from the client terminal(s) 808 and / or the server(s) 818.
[0065] The hardware processor(s) 802 may be implemented, for example, as central processing unit(s) (CPU), graphics processing unit(s) (GPU), field programmable gate array(s) (FPGA), digital signal processor(s) (DSP), and application specific integrated circuit(s) (ASIC). The processor(s) 802 may include one or more processors (homogeneous or heterogeneous), which may be arranged for parallel processing as a cluster and / or as one or more multi-core processing units.
[0066] Memory 806 (also referred to herein as program store and / or data storage device), such as random access memory (RAM), read-only memory (ROM), and / or storage devices, such as non-volatile memory, magnetic media, semiconductor memory devices, hard drives, removable storage devices, and optical media (e.g., DVD, CD-ROM), stores code instructions that are executed by hardware processor(s) 802. Memory 806 stores code 806A and / or training code 806B that implement one or more operations and / or features of the methods described with reference to Figures 4A-4D, 5A-5D, 9, 10, and 11.
[0067] The computing device 804 may include a data storage device 822 that stores data, e.g., annotated dataset(s) 822A of sample images annotated with monitored gaze data and / or monitored action data, machine learning model(s) 822B as described herein, and / or training dataset 822C for training the machine learning model(s) 822B as described herein. The data storage device 822 may be embodied, for example, as a memory, a local hard drive, a removable storage device, an optical disk, a storage device, and / or a remote server and / or computing cloud (e.g., accessed via the network 810). It should be noted that executable code portions of the data stored in the data storage device 822 may be loaded into the memory 806 for execution by the processor(s) 802.
[0068] Computing device 804 may include a data interface 824, optionally a network interface for connecting to network 810, such as one or more of a network interface card, a wireless interface for connecting to a wireless network, a physical interface for connecting to a cable for network connection, a virtual interface implemented in software, network communications software providing an upper layer of network connectivity, and / or other implementations. Computing device 804 may use network 810 to access one or more remote servers 818 to, for example, download updated versions of machine learning model(s) 822B, code 806A, training code 806B, and / or training dataset(s) 822C.
[0069] Computing device 804 can communicate using network 810 (or another communication channel, such as through a direct link (e.g., cable, wireless) and / or an indirect link (e.g., via an intermediate computing device, such as a server, and / or a storage device)) with one or more of the following: *As described herein, for example, a client terminal(s) 808 when the computing device 804 acts as a server providing image analysis services (e.g., SaaS) to remote terminals. *For example, a server 818 implemented in conjunction with a PACS and / or electronic medical record server and / or manufacturing server / quality assurance server that may store sample images captured by imaging device 812 and / or gaze monitoring data captured by gaze monitoring device 826 and / or operational data captured by operational monitoring device 814 of various users.
[0070] It should be noted that the imaging interface 820 and the data interface 824 can exist as two independent interfaces (e.g., two network ports), as two virtual interfaces on a common physical interface (e.g., a virtual network on a common network port), and / or can be combined into a single interface (e.g., a network interface).
[0071] The computing device 804 includes or is in communication with a user interface 826. The user interface 826 includes mechanisms designed for a user to input data (e.g., generate a report) and / or view data (e.g., view a sample). Exemplary user interfaces 826 include, for example, one or more of a touch screen, a microscope, a display, a keyboard, a mouse, and voice-activated software using a speaker and microphone.
[0072] Reference is now also made to Figure 9, which is a flowchart of a method for automatically creating an annotated training dataset containing sample images of objects annotated with supervised gaze for training an ML model, according to some embodiments of the present invention.
[0073] 9, a sample of an object is provided at 902. The sample may be, for example, a biological sample, a chemical sample, and / or a manufactured sample (e.g., an electrical and / or mechanical component).
[0074] Examples of samples include microscope slides of tissue, which may be pathological tissue and live cell cultures (e.g., prepared by slicing frozen sections and / or formalin-fixed paraffin-embedded (FFPE) slides). The sample may be contained in other ways, such as in at least one of a transparent sample cartridge, a vial, a tube, a capsule, a flask, a vessel, a receptacle, a microarray, or a microfluidic chip. Tissue samples may be obtained during surgery, such as during a biopsy procedure, an FNA procedure, a core biopsy procedure, a colonoscopy for colon polyp removal, surgery for the removal of an unknown mass, surgery for the removal of a benign cancer, and / or surgery for the removal of a malignant cancer, or surgery for the treatment of a medical condition. Tissue may be obtained from bodily fluids, such as urine, synovial fluid, blood, and cerebrospinal fluid. The tissue may be in the form of a combined cell group, such as a histological slide. The tissue may be in the form of individual cells or cell clusters suspended in a bodily fluid, for example in the form of a cytological sample.
[0075] In another example, the sample may be a statistically selected sample of manufactured products such as microarrays (e.g., DNA microarrays), silicon chips, and / or electrical circuits that may be selected for quality assurance evaluation to identify manufacturing defects and / or make pass / fail decisions, etc.
[0076] At 904, the line of sight of a user viewing the sample is monitored. The sample may be viewed under a microscope and / or other optical device, and / or an image of the sample may be presented on a display for viewing by the user. The image and / or view may be at any magnification.
[0077] The monitored line of sight can be collected, for example, while a user is observing and / or analyzing a sample using a microscope and / or by viewing a display and providing result data, without interrupting, slowing down, or obstructing the user.
[0078] A user's gaze can be monitored by tracking the user's pupil movements as the user views a sample under a microscope and / or presented on a display, for example, using the devices described with reference to Figures 2A and 2B, 3A-3D, and 12. The user's pupil movements can be tracked, for example, by a camera, as described herein.
[0079] Pupil movement is mapped to an area within the field of view of the sample that the user is viewing. Pupil movement can be tracked at various resolution levels that indicate various levels of precision in what the user is actually viewing, such as a single cell or group of cells within the area in the case of tissue, and / or microscopic features such as DNA strands and / or microscopic electrical and / or mechanical components in the case of manufactured products. Pupil movement can be mapped to areas of various sizes, for example, to a single pixel of the FOV and / or image of the sample, and / or to a group of pixels, and / or to the FOV as a whole. As described herein, wider and / or lower resolution tracking can be used for weak annotation of the FOV and / or image of the sample in the training dataset. Weak annotation of the FOV and / or image of the sample in the training dataset can be from any gaze coordinate at any resolution.
[0080] Optionally, the user's gaze is tracked as a function of time. An indication of the time the monitored gaze maps to each particular region of the sample over a time interval can be obtained. The time spent can be defined, for example, per FOV (e.g., as described herein) and for each pixel and / or group of pixels of the image that maps to the FOV of the sample the user is observing. For example, during a 10-minute observation session, the user spends 1 minute looking at one FOV and 5 minutes looking at another FOV. Alternatively or additionally, the monitored gaze can be represented in and / or include an ordered time sequence that dynamically maps the adaptation of the monitored gaze of the different fields of view being observed to different particular pixels over the observation time interval. For example, the user spent the first minute of the observation session looking at a first FOV located near the center of the sample, then spent 5 minutes looking at a second FOV located to the right of the first FOV, and then spent another 2 minutes looking back at the first FOV.
[0081] The monitored gaze can be visualized and / or implemented as a data structure, optionally a heat map, corresponding to the image of the sample. A pixel of the heat map corresponds to a pixel and / or group of pixels in the image of the sample. The respective intensities of each pixel of the heat map that correlate with the representation of the monitored gaze are mapped to each respective pixel, e.g., the pixel intensity values of the heat map represent the total time the user spent observing those pixels. The heat map can be presented as an overlay on the image of the sample. The mapping of time to pixel intensity values can be based, for example, on set thresholds (e.g., less than 1 minute, between 1 and 3 minutes, and more than 3 minutes) and / or relative time spent (e.g., more than 50% of the total time, between 20% and 50% of the total time, and less than 20% of the total time), or other techniques.
[0082] The heat map can represent the total time spent in each pixel and / or each region of the FOV. A time sequence representation showing the dynamic adaptation of the monitored gaze (i.e., what the user looks at as a function of time) can be calculated in addition to and / or instead of the heat map. The monitored gaze representing gaze as a function of time can be represented, for example, as a directed line overlaid on the pixels of the sample image and / or overlaid on the heat map. In another example, the monitored gaze can be represented as an ordered time sequence of each FOV labeled with an indication of the time spent in each FOV. Each FOV can be mapped to the sample image (WSI), for example, shown as a boundary representing the FOV overlaid on the sample image.
[0083] A monitored gaze can be represented using other data structures, such as a vector of coordinates in the coordinate system of the FOV indicating where the user is looking. The vector can be a time sequence indicating locations in the FOV that the user looked at over a period of time. In yet another example, a monitored gaze can be represented using one or more successive overlays, each including markings (e.g., color, shape, highlighting, outline, shading, pattern, jet color map, etc.) over the region of the FOV the user is gazing at and representing the monitored gaze for a small time interval (e.g., 1 second, 10 seconds, 30 seconds, etc.). In yet another example, a monitored gaze can be represented by indicating the time spent gazing at each FOV region, and images of the FOV can be sequentially arranged according to the user's observation of the FOV. For example, a contour (or other marking, such as shading, highlighting, etc.) over the FOV indicating the time spent gazing at the region shown in the FOV can be represented. Time can be represented, for example, by metadata, the thickness of the contour, and / or the color and / or intensity of the marking. Multiple contours can be presented, each representing a different gaze point. For example, three circles may be presented on the FOV, with one red circle representing a 3-minute fixation and two blue circles representing fixations of less than 30 seconds.
[0084] Reference is now made to FIG. 11, which is a schematic diagram showing a heat map (represented as a jet color map) overlaid on an image 1102 of an observed field of view of a sample of an object, in accordance with some embodiments of the present invention. In this illustrated case, the sample is a tissue slide obtained from a subject. High intensity pixel values 1104 represent areas where the user spent a significant amount of time observing the area, medium intensity pixel values 1106 represent areas where the user spent a moderate amount of time observing the area, and low intensity pixel values 1108 represent areas where the user spent little time observing the area. Referring again to 906 of FIG. 9, an FOV of the sample being observed by the user is captured. The FOV can optionally be captured dynamically as a function of time. A time sequence of the FOV observed by the user can be generated.
[0085] When being observed using a microscope, the FOV can be captured by a camera (optionally a different camera than the camera used to track the eye movements of the user observing the sample under the microscope) that captures images of the sample as seen under the microscope while the user is observing the sample under the microscope.
[0086] When the FOV is presented on a display, the FOV presented on the display can be captured, for example, by performing a screen capture operation.
[0087] At 908, manipulation(s) of the FOV presentation of the sample by each user can be monitored. The sample, when magnified, may be very large and may not be viewable by the users simultaneously to allow for proper analysis. Thus, the user may manipulate the image of the sample being viewed on the display and / or manipulate the slide and / or microscope to generate a different FOV.
[0088] Examples of manipulations include zooming in, zooming out, panning left, panning right, panning up, panning down, adjusting the light, adjusting the magnification, and adjusting the focus (e.g., in focus, out of focus) using axial axis (z-axis) scanning. If the sample is a tissue slide, the slide can be adjusted along the z-axis using the z-axis knob to observe different depths of the tissue under the microscope. If the sample is a 3D (three-dimensional) image, the 3D image can be sliced into 2D (two-dimensional) planes using back and forth scrolling to obtain 2D (two-dimensional) planes.
[0089] When an image of a sample is presented on a display, operation can be monitored by monitoring user interactions with a user interface associated with the display, e.g., icons, keyboard, mouse, and touch screen. When a sample is viewed under a microscope, operation can be monitored by sensors associated with various components of the microscope that detect, for example, which zoom lens is being used and / or sensors associated with components that adjust the position and / or light intensity of the sample.
[0090] The operations can be monitored as a function of time, for example as a time sequence showing which operations were performed over an observation time interval.
[0091] The monitored actions can be correlated with the monitored gaze, e.g., to correspond to the same timeline. For example, one minute into the observation session, the user switches the zoom from 50x to 100x, and the FOV changes from FOV_1 to FOV_2.
[0092] The manipulations (which can be used as ground truth labels) can be represented as an overlay on the image of the sample (e.g., acquired in 910) indicating the total time and / or sequence. For example, when a heat map corresponds to the image of the sample, the respective intensities of the respective pixels of the heat map correlate with the total time, the monitored line of sight is mapped to each respective pixel. In another example, one or more boundaries (e.g., circles, squares, irregular shapes) are overlaid on the image of the sample, the dimensions of each boundary correspond to the extent of the line of sight, and the markings on each boundary indicate the total time (e.g., the thickness and / or color of the boundary).
[0093] At 910, an image of the sample can be acquired. The image can be an image of the entire sample, such as a whole sample image, such as a whole slide image, and / or a high-resolution image of the entire product. The image of the sample can be acquired, for example, by scanning the slide with a scanner and / or using a high-resolution camera. Alternatively or additionally, the image of the sample can be created as a combination of images of the FOV of the sample.
[0094] At 912, the image of the sample can be aligned with an image of the FOV of the sample. The image of the FOV is captured while the user is viewing the sample under the microscope and depicts what the user is viewing using the microscope. Note that when the user is viewing the sample on the display, alignment to the image of the sample is not necessarily required because the user is directly viewing the FOV of the image of the sample.
[0095] Registration can be achieved by a registration process, for example, by matching features of the image in the FOV with the image of the sample, and can be strict and / or flexible, such as when the tissue sample may be physically moved during processing.
[0096] At 914, the monitored line of sight is mapped to pixels of the image of the sample, optionally the scanned image and / or the WSI.
[0097] The mapping can be performed using an FOV that is aligned with an image of the sample (e.g., a WSI and / or scanned image), i.e., the monitored line of sight is mapped to an FOV that is aligned with the image of the sample, thereby allowing the monitored line of sight to be directly mapped to pixels of the image of the sample.
[0098] The mapping of the monitored lines of sight can be done for each pixel of the image of the sample, and / or for each group of pixels of the image of the sample, and / or for each region of the image of the sample.
[0099] The representation of the monitored line of sight, optionally a heat map, may be normalized to pixels of the image of the sample. The monitored line of sight may first be correlated to FOVs acquired at different zoom levels and / or may first be correlated to FOVs acquired at different locations of the sample (not simultaneously visible on the display), so the monitored line of sight may require normalization to map to pixels of the image of the sample.
[0100] At 916, additional data, optionally metadata, associated with the sample can be obtained. The additional data can be assigned as weak labels to the entire image of the sample, for example, the scanned image and / or the WSI.
[0101] Examples of additional data include: *When the sample is tissue of a subject and / or a radiology image of the subject, a pathology / radiology report created by the user viewing the sample, a pathology / radiology diagnosis (e.g., type of cancer) created by the user viewing the sample, a sample score indicating a pathology / radiology evaluation of the sample (e.g., percentage of tumor cells, Gleason score) created by the user viewing the sample, at least one clinical parameter of the subject whose tissue is represented in the sample (e.g., stage of cancer), a historical parameter of the subject (e.g., smoker), and the results of a treatment administered to the subject (e.g., chemotherapy response results). *When the sample is a manufactured microarray, a user-provided indication of at least one manufacturing defect present in the manufactured microarray and / or an indication of whether the manufactured microarray passed or failed quality assurance testing. *When samples involve live cell cultures, cell growth rate, cell density, cell homogeneity, and cell heterogeneity. *Other User-Provided Data.
[0102] The additional data may be obtained, for example, from manual input provided by a user, automatically extracted from a pathology report, and / or automatically extracted from the subject's electronic health record (e.g., medical history, diagnostic codes).
[0103] At 918, one or more of the features described with respect to 902-916 are repeated for multiple different users, where each user may be observing a different tissue sample.
[0104] At 920, a training dataset of multiple records can be created, where each record can include an image of a sample (e.g., a scanned image, WSI) and one or more of a monitored line of sight mapped to pixels of the image of the sample, respective user actions taken to adjust the field of view of the sample, and additional data that can serve as target inputs and / or ground truth.
[0105] The target inputs and ground truth can be specified according to the desired outputs of the ML model being trained, as described herein.
[0106] At 922, one or more machine learning models are trained on the training dataset.
[0107] In one example, an ML model is trained to generate predicted gaze results for a target in response to an input of a target image of a target sample. Such a model can be used, for example, to train a new pathologist / radiologist on how to gaze at a sample and / or as a quality control measure to verify that the pathologist / radiologist is properly gazing at the sample. In another example, such a model can be used to train a new quality assurance technician on how to gaze at manufactured products for manufacturing defect evaluation and / or quality assurance, and / or to monitor the quality assurance technicians. In another example, such a model can be used to train a new subject matter expert on how to gaze at live cell cultures. Such an ML model can be trained on a training dataset including multiple records, each record including an image of a sample and a ground truth representation of the monitored gaze of a user observing the sample. When the records of the training dataset also include actions taken by a user observing the sample in the record, actions taken by the target user observing the target image can be provided as input to the ML model.
[0108] In another example, the ML model is trained to generate results for additional data, such as a target visual assessment, for example, a quality assurance result (e.g., pass / fail, identified manufacturing defects), a sample clinical score, a pathology / radiology diagnosis, and / or a target predicted pathology / radiology report. In another example, when the sample is a live cell culture, the ML model is trained to generate results for a target cell growth rate, a target cell density, a target cell homogeneity, and / or a target cell heterogeneity. Such a model can be used, for example, to determine additional data for the target sample. The ML model is provided with inputs of one or more of a target image of the target sample, monitored gazes of a target user observing the target sample, and presentation actions of the target sample performed by the target user. Such an ML model can be trained on a training dataset including multiple recordings, each recording including an image of the sample, a ground truth representation of the additional data for the sample, and monitored gazes and optionally monitored actions of a respective user observing the sample in the recording.
[0109] In yet another example, an ML model is trained to generate a predicted outcome of a target manipulation of a presentation of a sample in response to an input of a target image of the target sample. Such a model can be used, for example, to train a new domain expert on how to manipulate the sample to obtain an FOV that allows proper observation of the sample, and / or can be used as a quality control measure to verify that an existing domain expert is properly manipulating the sample to obtain an FOV that allows proper observation of the sample. Such an ML model can be trained on a training dataset that includes multiple recordings, each recording including an image of the sample and a ground truth representation of the manipulation performed by a user observing the sample. When the training dataset records also include the gaze of the user observing the sample in the recording, the gaze of the target user observing the target image can be provided as input to the ML model.
[0110] Exemplary architectures of the machine learning models described herein include, for example, statistical classifiers and / or other statistical models, neural networks of various architectures (e.g., convolutional, fully connected, deep, encoder-decoder, recurrent, graph), support vector machines (SVMs), logistic regression, k-nearest neighbors, decision trees, boosting, random forests, regressors, and / or any other commercial or open source package that enables regression, classification, dimensionality reduction, supervised, unsupervised, semi-supervised, or reinforcement learning. Machine learning models can be trained using supervised and / or unsupervised techniques.
[0111] Reference is now made to Figure 10, which is a flowchart of a method of inference by a machine learning model trained on a training dataset of images annotated with indications of monitored gaze and / or monitored action, according to some embodiments of the present invention. At 1002, one or more machine learning models are provided. The machine learning models are trained using the techniques described with respect to Figure 9, for example, as described with respect to 922 of Figure 9.
[0112] The gaze of a user viewing the sample target image is monitored in real time at 1004. An example of a gaze monitoring technique is described, for example, with reference to 904 in FIG.
[0113] At 1006, manipulation of a sample presentation by a user can be monitored in real time. The manipulation can be a user adjusting a microscope setting and / or adjusting the presentation of an image of the sample on a display. Examples of samples are described, for example, with reference to 902 in FIG. 9. The sample can be placed under the microscope for viewing by the user. Examples of techniques for monitoring manipulation and / or exemplary manipulations are described, for example, with reference to 908 in FIG. 9.
[0114] At 1008, a target image of the sample is provided to the machine learning model(s). Optionally, a monitored gaze of a user observing the sample is provided to the machine learning model in addition to the target image. Alternatively or additionally, monitored actions taken by the user are provided to the machine learning model in addition to the target image. Alternatively or additionally, one or more other data items, such as tissue type and / or medical history, as described with reference to 916 of FIG. 9, are provided to the machine learning model in addition to the target image.
[0115] The image of the sample can be acquired, for example, by scanning a histopathology slide to create a WSI and / or by capturing a high-resolution image of the product. When the sample is a live cell culture, the image can be acquired, for example, by a high-resolution camera and / or a camera connected to a microscope. The user can observe the image of the sample on a display. Additional exemplary details of acquiring an image of the sample are described, for example, with reference to 910 of FIG. 9.
[0116] For example, different results may be generated depending on the training dataset used to train the machine learning model(s) and / or depending on the inputs provided to the machine learning model, as described with reference to 922 in Figure 9. Examples of processes based on the results of the machine learning models are described with reference to 1010 and 1012 relating to diagnosing and / or treating subjects, and 1014-1018 relating to training new users and / or quality control of users.
[0117] At 1010, for example, a sample score (e.g., a visual assessment score) indicating a pathological / radiological assessment of a sample (e.g., a biological sample, a tissue sample, a radiological image, and / or a live cell culture), and / or a pass / fail result of a quality assurance test (e.g., of a manufactured product), and / or other examples of data described with reference to 916 of FIG. 9 can be obtained as a result of the machine learning model.
[0118] At 1012, in the case of a medical sample (e.g., a biological sample, a tissue sample, a radiological image, and / or a live cell culture), the subject may be treated and / or evaluated according to the sample score. For example, if the sample score exceeds a threshold, the subject may be administered chemotherapy; if the pathology diagnosis indicates a certain type of cancer, the subject may undergo surgery, etc. In the case of a manufactured product (e.g., a microarray), the object may be further processed if the sample score indicates passing quality assurance tests and / or no major manufacturing defects, and / or the object may be rejected if the sample score indicates failing quality assurance tests and / or there are major manufacturing defects.
[0119] Alternatively or additionally to 1010-1012, at 1014, an indication of a predicted line of sight and / or a predicted operation is obtained as a result of the machine learning model. The predicted monitored line of sight may be, for example, per pixel and / or per group of pixels and / or per region of the target image. The predicted operation may be a manipulation of the entire image and / or a manipulation of the current FOV, for example, zooming in and / or panning the field of view.
[0120] The monitored gaze can be represented as a heat map. The heat map can include a plurality of pixels that map to pixels of the target image. The intensity of the pixels in the heat map correlates with the predicted time of gaze. Additional exemplary details of the heat map are described, for example, with reference to 908 in FIG. 9.
[0121] The predicted line of sight can be presented on a display, for example as an overlay on an image of the sample.
[0122] The predicted monitored gaze and / or predicted operations can be represented as a time series showing dynamic gazes mapped to pixels of the target image over a time interval and / or operations performed during different times of the time interval.
[0123] At 1016, the real-time monitoring of the manipulation is compared to a prediction of the manipulation and / or the real-time monitoring of the gaze is compared to a prediction of the gaze.
[0124] For example, this comparison can be performed by calculating a difference indicating the amount of similarity and / or dissimilarity, e.g., the number of pixels between the predicted line of sight and the actual line of sight. In another example, the difference between the real-time monitoring and the time series is compared, and an alert is generated when the difference exceeds a threshold.
[0125] At 1018, one or more actions can be taken. Actions can be taken when the difference exceeds a threshold and / or when the difference indicates statistical dissimilarity. For example, an alert can be generated and / or an instruction can be generated and, for example, presented on a display, played as a video, presented as an image, presented as text, and / or played as an audio file over a speaker. The instruction can indicate to the user that the user's action and / or gaze differs from expected. Such instruction can be provided, for example, to train new subject matter experts and / or to monitor trained subject matter experts as a form of quality control, such as to help ensure that the subject matter experts are following standard practice. The instruction can indicate what the predicted gaze and / or predicted action is so that the user can follow the instruction.
[0126] In another example, instructions indicating predicted gaze and / or predicted manipulation are provided without necessarily monitoring the user's current gaze and / or manipulation, e.g., to guide the user while evaluating a sample.
[0127] At 1020, one or more of the features described with reference to 1004-1008 and / or 1014-1018 are repeated during the observation session, e.g., to dynamically guide the user's gaze and / or actions to sample evaluation and / or for continuous real-time training and / or quality control.
[0128] Reference is now made to Figure 12, which is a schematic diagram of a component 1202 mounted on a microscope 1204 for monitoring the line of sight of a user observing a sample (e.g., a biological sample, a live cell sample, a tissue sample, or a manufactured product such as a microarray) under the microscope, in accordance with some embodiments of the present invention. Component 1202 can be integrated with microscope 1204 and / or designed to be connected and / or disconnected from microscope 1204.
[0129] Component 1202 is placed between objective lens 1212 and eyepiece 1224 of microscope 1204 .
[0130] The component 1202 is designed so that adding the component 1202 does not affect (or does not significantly affect) the optical path and / or the observation experience and / or workflow of a user who uses the microscope. Some infinity correction methods do not affect the optical path and / or the experience and / or the workflow.
[0131] The component 1202 may include an optical device 1206. The optical device 1206 directs a first set of electromagnetic frequencies reflected back from a user's eye 1208, which is observing a sample 1210 under a microscope objective 1212, to a camera 1214. The camera 1214 generates a representation of the user's tracked line of sight. The first set of electromagnetic frequencies may be infrared (IR) frequencies generated by an IR source 1216. The camera 1214 may be a near-IR (NIR) camera. The optical device 1206 simultaneously directs a second set of electromagnetic frequencies from the sample 1210 to a camera 1220, which captures an image indicative of the field of view observed by the user.
[0132] The first set of electromagnetic frequencies and the second set of electromagnetic frequencies can include the visible light spectrum.The camera 1220 can be a red-green-blue (RGB) camera.
[0133] The optical device 1206 may include a beam splitter 1222. The beam splitter 1222 directs a first set of electromagnetic frequencies from the IR source 1216 to an eyepiece 1224 where the user's eye 1208 is located. The beam splitter 1222 simultaneously directs the first set of electromagnetic frequencies reflected back from the user's eye 1208 to the NIR camera 1214 via the eyepiece 1224, and directs a second set of electromagnetic frequencies from the sample 1210 to the camera 1220 and the eyepiece 1224.
[0134] Along the way from the IR source 1216, the IR frequencies pass through a linear polarizer (P) 1226, a polarized beam splitter (PBS) 1228, and a λ / 4 filter 1230 positioned along the plane of incidence. The PBS 1228 allows IR energy to travel from the IR source 1216 to the beam splitter 1222 but prevents IR energy from returning to the IR source 1216, instead reflecting the IR energy to the NIR camera 1214. An NIR filter 1232 is positioned in the optical path from the beam splitter 1222 to the RGB camera 1220 and prevents reflected IR energy from reaching the RGB camera 1220. The optical arrangement can include a linear polarizer (P) positioned along the plane of incidence and a λ / 4 quarter wave plate. These act as optical isolators, converting linearly polarized light into circularly polarized light and preventing the IR back-reflected light set normal to the plane of incidence after returning through λ / 4 from entering the IR light source and directing it to the IR camera.
[0135] It should be noted that while a single eye 1208 and a single eyepiece 1224 are shown, in practice a user would use both eyes and two eyepieces. The optical device 1206 separates the electromagnetic light waves from the single optical path after reflection from the two eyes into two optical paths that go to two of the IR cameras 1216. This separation can be implemented, for example, as one or more of using an infrared spectrum light source shifted at a specific wavelength using polarizers and / or wave plates that direct different polarizations into different paths, and / or dichroic mirrors and spectral filters, and / or adding amplitude modulation at different frequencies for each optical path for heterodyne detection.
[0136] Various embodiments provide tools and techniques for performing annotation data collection, and more particularly, methods, systems, and apparatus for performing annotation data collection using gaze-based tracking, which in some cases may include for training artificial intelligence ("AI") systems (which may include, but are not limited to, at least one of neural networks, convolutional neural networks ("CNNs"), learning algorithm-based systems, or machine learning systems, etc.).
[0137] In various embodiments, a first camera can capture at least one first image of at least one eye of a user while the user is viewing an optical view of a first sample. A computing system can analyze the captured at least one first image of the at least one eye of the user and at least one second image of the optical view of the first sample to determine whether the at least one eye of the user is focused on a particular region of the optical view of the first sample. Based on a determination that the at least one eye of the user is focused on a particular region of the optical view of the first sample, the computing system can identify at least one particular portion of the at least one second image that corresponds to the particular region of the optical view of the first sample. The computing system can collect attention data including the identified at least one particular portion of the at least one second image, and the collected attention data can be stored in database 110a or 110b. According to some embodiments, the collection of attention data may occur without interrupting, slowing down, or obstructing the user while the user is providing result data either while diagnosing the first sample using microscope 115 or while diagnosing the image of the first sample displayed on display screen 120. In some cases, the collected attention data may include, but is not limited to, at least one of: one or more coordinate locations of at least one particular portion of the optical view of the first sample; an attention duration that the user focuses on at least one particular portion of the optical view of the first sample; a zoom level of the optical view of the first sample while the user focuses on at least one particular portion of the optical view of the first sample;In some cases, the identified at least one specific portion of the at least one second image corresponding to a particular region of the optical view of the first sample can include, without limitation, at least one of: one or more specific cells, one or more specific tissues, one or more specific structures, or one or more molecules, etc.
[0138] In some embodiments, the computing system may generate at least one highlight field in the at least one second image that covers at least one identified portion of the at least one second image that corresponds to a particular region of the optical view of the first sample. In some cases, each of the at least one highlight field may include, without limitation, at least one of a color, a shape, or a highlight effect, etc., where the highlight effect may include, but is not limited to, at least one of an outlining effect, a shadowing effect, a patterning effect, a heat map effect, or a jet color map effect, etc.
[0139] According to some embodiments, the at least one second image can be displayed on a display screen. Capturing at least one first image of at least one eye of the user can include capturing the at least one first image of at least one eye of the user with a camera while the user views image(s) or video(s) of the optical view of the first sample displayed as the at least one second image on the display screen of the display device. An eye-tracking device can be used instead of a camera to collect attention data while the user views the image(s) or video(s) of the first sample displayed on the display screen of the display device. Identifying at least one specific portion of the at least one second image corresponding to a specific region of the optical view of the first sample can include identifying, using a computing system, at least one specific portion of the at least one second image displayed on the display screen that corresponds to the specific region of the optical view of the first sample. The computing system can display the at least one second image on the display screen with at least one generated highlight field covering the identified at least one specific portion of the at least one second image that corresponds to the specific region of the optical view of the first sample.
[0140] In some embodiments, the display of the at least one second image on the display screen can shift in response to a command by a user. In some cases, the shifting display of the at least one second image can include at least one of a horizontal shift, a vertical shift, panning, tilting, zooming in, or zooming out of the at least one second image on the display screen, etc. The first camera can track the movement of at least one eye of the user while the user is viewing the shifting display of the at least one second image on the display screen. The computing system can match the tracked movement of the at least one eye of the user with the shifting display of the at least one second image on the display screen based at least in part on one or more of the tracked movement of the at least one eye of the user, an identified at least one specific portion of the at least one second image corresponding to a specific region of the optical view of the first sample, or at least one of a horizontal shift, a vertical shift, panning, tilting, zooming in, or zooming out of the at least one second image on the display screen. Instead of using a camera, an eye-tracking device can be used to collect additional attention data as the user views the shifting display of at least one second image on the display screen of the display device.
[0141] Alternatively, the microscope can project the optical view of the first sample onto an eyepiece through which at least one eye of a user is observing. The second camera can capture at least one second image of the optical view of the first sample. In some cases, capturing at least one first image of at least one eye of the user can include capturing at least one first image of at least one eye of the user using the first camera while the user is viewing the optical view of the first sample through the eyepiece. Identifying at least one specific portion of the at least one second image corresponding to a specific region of the optical view of the first sample can include identifying, using a computing system, at least one specific portion of the at least one second image being observed through the eyepiece that corresponds to the specific region of the optical view of the first sample. In some cases, the computing system can display the at least one second image on a display screen with at least one generated highlight field covering the identified at least one specific portion of the at least one second image corresponding to the specific region of the optical view of the first sample.
[0142] In some cases, the first camera can be one of an infrared ("IR") camera, a back-reflected IR camera, a visible color camera, a light source, a location photodiode, etc. In some cases, the microscope can include two or more of a plurality of mirrors, a plurality of dichroic mirrors, or a plurality of half mirrors that reflect or pass at least one of, but are not limited to, an optical view of the first sample observed through the eyepiece, an optical view of at least one eye of a user observed through the eyepiece and captured as at least one first image by the first camera, or a projection of the generated at least one highlighted field through the eyepiece onto at least one eye of a user, etc.
[0143] According to some embodiments, the projection of the optical view of the first sample onto the eyepiece can be shifted by at least one of adjusting an XY stage carrying a microscope slide including the first sample, changing an objective lens or a zoom lens, adjusting the focus of the eyepiece, etc. The first camera can track movement of at least one eye of the user as the user views the shifted projection of the optical view of the first sample onto the eyepiece. The computing system can match the tracked movement of the at least one eye of the user with the shifted projection of the optical view of the first sample onto the eyepiece based at least in part on one or more of the tracked movement of the at least one eye of the user, identifying at least one specific portion of the at least one second image corresponding to a specific region of the optical view of the first sample, or at least one of adjusting an XY stage carrying a microscope slide including the first sample, changing an objective lens or a zoom lens, adjusting the focus of the eyepiece, etc.
[0144] Alternatively or additionally, the one or more audio sensors may capture one or more voice notes from the user while the user is viewing the optical view of the first sample, and the computing system may map the one or more voice notes captured from the user with the at least one second image of the optical view of the first sample to match the one or more captured voice notes with the at least one second image of the optical view of the first sample.
[0145] According to some embodiments, the computing system can receive outcome data provided by a user. The outcome data includes at least one of a diagnosis of the first sample, a pathology score of the first sample, or a set of identification data corresponding to at least a plurality of portions of the first sample. The computing system can train an AI system (which can generally include, but is not limited to, at least one of a neural network, a convolutional neural network (“CNN”), a learning algorithm-based system, a machine learning system, etc.) based at least in part on at least one of an analysis of at least one captured first image of at least one eye of the user and at least one second image of an optical view of the first sample, or a joint analysis of the collected attention data and the received outcome data, to generate a model used to generate a predicted value. In some embodiments, the predicted value can include, but is not limited to, at least one of a predicted clinical outcome, a predicted attention data, etc.
[0146] According to various embodiments described herein, the annotation data collection system described herein allows for recording of a user's (e.g., a pathologist's) visual attention in addition to tracking the microscope FOV during the scoring process, thus providing highly localized spatial information that supports the overall score of a slide. This information is used to develop algorithms such as tumor localization, classification, and digital scoring in WSIs. Algorithms for localization, classification, and digital scoring of ROIs in WSIs other than tumors can also be developed.
[0147] These and other aspects of the annotation data collection system using gaze-based tracking and / or the training of an AI system based on annotation data collected using gaze-based tracking (i.e., the annotation data collection system using gaze-based tracking, or the training of an AI system based on annotation data collected using gaze-based tracking, or both) are described in more detail with respect to the figures.
[0148] The following detailed description provides further details of a few exemplary embodiments to enable those skilled in the art to practice such embodiments. The described examples are provided for illustrative purposes and are not intended to limit the scope of the invention.
[0149] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the described embodiments. However, it will be apparent to those skilled in the art that other embodiments of the present invention may be practiced without some of these specific details. In other instances, certain structures and devices are shown in block diagram form. Although several embodiments are described herein and various features are attributed to different embodiments, it will be understood that features described with respect to one embodiment may also be combined with other embodiments. Similarly, however, one or more individual features of any described embodiment should not be considered essential to every embodiment of the present invention, since other embodiments of the present invention may omit such features.
[0150] Unless otherwise specified, all numbers used herein to express quantities, dimensions, and the like are to be understood as being modified in all instances by the term "about." In this application, the use of an unspecified number includes the plural unless specifically stated otherwise, and the use of the terms "and / and" and "or / or" means "and / or" unless specifically stated otherwise. Furthermore, the use of the term "comprises" and other forms such as "comprises" should be considered non-exclusive. Also, terms such as "element" or "component" encompass both elements and components comprising one unit and elements and components comprising two or more units, unless specifically stated otherwise.
[0151] The various embodiments described herein embody (in some cases) software products, computer-implemented methods, and / or computer systems (i.e., software products, computer-implemented methods, and / or computer systems), but represent tangible, concrete improvements to existing technology areas including, but not limited to, annotation collection technology, annotation data collection technology, and the like. In other aspects, certain embodiments may include, for example, capturing, with a first camera, at least one first image of at least one eye of a user while the user is viewing an optical view of a first sample; capturing, with a second camera, at least one second image of the optical view of the first sample; analyzing, with a computing system, the captured at least one first image of the at least one eye of the user and the captured at least one second image of the optical view of the first sample to determine whether the at least one eye of the user is focused on a particular region of the optical view of the first sample; and, based on a determination that the at least one eye of the user is focused on a particular region of the optical view of the first sample, capturing, with the computing system, at least one image of the at least one eye of the user corresponding to the particular region of the optical view of the first sample. identifying at least one particular portion of each of the at least one second images; collecting, with the computing system, attention data comprising the identified at least one particular portion of the at least one second image; storing the collected attention data in a database; receiving, with the computing system, result data provided by a user, the result data comprising at least one of a diagnosis of the first sample, a pathology score of the first sample, or a set of identification data corresponding to at least a plurality of portions of the first sample; and based at least in part on at least one of an analysis of the captured at least one first image of at least one eye of the user and the captured at least one second image of the optical view of the first sample, or a joint analysis of the collected attention data and the received result data.The functionality of the user equipment or the system itself (e.g., annotation collection system, annotation data collection system, etc.) can be improved, such as by training at least one of a neural network, a convolutional neural network ("CNN"), an artificial intelligence ("AI") system, or a machine learning system to generate a model used to generate a predictive value (e.g., at least one of a predicted clinical outcome or a predicted attention data, etc.).
[0152] In particular, various embodiments have some degree of abstraction that extends beyond merely conventional computer processing operations. Some examples include capturing, with a first camera, at least one first image of at least one eye of a user as the user views an optical view of a first sample; capturing, with a second camera, at least one second image of the optical view of the first sample; analyzing, with a computing system, the captured at least one first image of the at least one eye of the user and the captured at least one second image of the optical view of the first sample to determine whether the at least one eye of the user is focused on a particular region of the optical view of the first sample; and based on a determination that the at least one eye of the user is focused on the particular region of the optical view of the first sample, capturing, with the computing system, at least one second image of the at least one second image corresponding to the particular region of the optical view of the first sample. and identifying at least one particular portion of the at least one second image; collecting, with a computing system, attention data including the identified at least one particular portion of the at least one second image; storing the collected attention data in a database; receiving, with the computing system, result data provided by a user, the result data including at least one of a diagnosis of the first sample, a pathology score of the first sample, or a set of identification data corresponding to at least a plurality of portions of the first sample; and training at least one of a neural network, a convolutional neural network ("CNN"), an artificial intelligence ("AI") system, or a machine learning system based at least in part on at least one of an analysis of the captured at least one first image of at least one eye of the user and the captured at least one second image of the optical view of the first sample, or a joint analysis of the collected attention data and the received result data, to generate a predictive value (e.g.,The functions described herein can be implemented by devices, software, systems, and methods with certain novel functionality (e.g., steps or operations), such as generating models used to generate at least one of predicted clinical outcome or predictive attention data, etc. These functionality can produce tangible results outside of the implementing computer system, and include, by way of example only, enabling recording of a user's visual attention in addition to tracking the FOV of a sample during the user's visual analysis, thus providing highly localized spatial information supporting global annotation of a sample analyzed by the user; in some cases, this information is used to develop algorithms for sample region of interest ("ROI") localization, classification, and digital scoring of the sample, at least some of which can be observed or measured by the user and / or service provider (i.e., the user, the service provider, or both).
[0153] In one aspect, a method includes using a microscope to project an optical view of a first sample onto an eyepiece through which at least one eye of a user is viewing; using a first camera to capture at least one first image of the at least one eye of the user as the user views the optical view of the first sample through the eyepiece; using a second camera to capture at least one second image of the optical view of the first sample; and using a computing system to analyze the captured at least one first image of the at least one eye of the user and the captured at least one second image of the optical view of the first sample to determine the user's The method may include determining whether at least one eye is focused on a particular region of the optical view of the first sample; based on a determination that at least one eye of the user is focused on a particular region of the optical view of the first sample, using a computing system to identify at least one particular portion of at least one second image being viewed through the eyepiece that corresponds to the particular region of the optical view of the first sample; collecting, using the computing system, attention data including the identified at least one particular portion of the at least one second image; and storing the collected attention data in a database.
[0154] In some embodiments, the first sample can be contained in at least one of a microscope slide, a transparent sample cartridge, a vial, a tube, a capsule, a flask, a vessel, a receptacle, a microarray, or a microfluidic chip, etc. In some cases, the first camera can be one of an infrared ("IR") camera, a back-reflecting IR camera, a visible color camera, a light source, or a location photodiode, etc. In some cases, the microscope can include two or more of a plurality of mirrors, a plurality of dichroic mirrors, or a plurality of half mirrors that reflect or pass at least one of an optical view of the first sample observed through an eyepiece or an optical view of at least one eye of a user observed through the eyepiece and captured as at least one first image by the first camera.
[0155] According to some embodiments, the identified at least one specific portion of the at least one second image corresponding to a particular region of the optical view of the first sample can include at least one of one or more specific cells, one or more specific tissues, one or more specific structures, or one or more molecules, etc. In some cases, identifying the at least one specific portion of the at least one second image can include determining, with a computing system, a coordinate location within the at least one second image of the optical view that corresponds to the identified at least one specific portion of the at least one second image.
[0156] In some embodiments, the method may further include receiving, with a computing system, result data provided by a user, the result data including at least one of a diagnosis of the first sample, a pathology score of the first sample, or a set of identification data corresponding to at least a plurality of portions of the first sample; and training at least one of a neural network, a convolutional neural network (“CNN”), an artificial intelligence (“AI”) system, or a machine learning system based at least in part on at least one of an analysis of at least one captured first image of at least one eye of the user and at least one captured second image of an optical view of the first sample, or a joint analysis of the collected attention data and the received result data, to generate a model used to generate a predicted value. In some cases, the predicted value may include at least one of a predicted clinical outcome, predicted attention data, or the like. In some cases, collecting the attention data may be performed without interrupting, slowing down, or interrupting the user as the user provides the result data while diagnosing the first sample using the microscope.
[0157] According to some embodiments, the method may further include tracking, with the first camera, movement of at least one eye of the user and, with the computing system, simultaneously tracking, with the computing system, at least one of: one or more coordinate locations of the identified at least one particular portion of the at least one second image; an attention duration during which the user focuses on the particular region of the optical view; or a zoom level of the optical view of the first sample while the user focuses on the particular region of the optical view. In some cases, determining whether at least one eye of the user is focusing on the particular region of the optical view of the first sample may include determining whether at least one eye of the user is focusing on the particular region of the optical view of the first sample based at least in part on at least one of: one or more coordinate locations of the identified at least one particular portion of the at least one second image; an attention duration during which the user focuses on the particular region of the optical view; or a zoom level of the optical view of the first sample while the user focuses on the particular region of the optical view.
[0158] In some embodiments, the method may further include capturing, using an audio sensor, one or more voice notes from the user while the user is viewing the optical view of the first sample, and mapping, using a computing system, the one or more voice notes captured from the user with the at least one second image of the optical view of the first sample to match the one or more captured voice notes with the at least one second image of the optical view of the first sample.
[0159] In another aspect, a system can include a microscope, a first camera, a second camera, and a computing system. The microscope can be configured to project an optical view of a first sample into an eyepiece through which at least one eye of a user is observing. The first camera can be configured to capture at least one first image of at least one eye of a user while the user is viewing the optical view of the first sample through the eyepiece. The second camera can be configured to capture at least one second image of the optical view of the first sample. The computing system can include at least one first processor and a first non-transitory computer-readable medium communicatively coupled to the at least one first processor. The first non-transitory computer-readable medium can have stored thereon computer software including a first set of instructions that, when executed by at least one first processor, causes the computing system to: analyze at least one captured first image of at least one eye of the user and at least one captured second image of the optical view of the first sample to determine whether the at least one eye of the user is focused on a particular region of the optical view of the first sample; based on a determination that the at least one eye of the user is focused on a particular region of the optical view of the first sample, identify at least one particular portion of the at least one second image being observed through the eyepiece that corresponds to the particular region of the optical view of the first sample; collect attention data including the identified at least one particular portion of the at least one second image; and store the collected attention data in a database.
[0160] In some embodiments, the first set of instructions, when executed by the at least one first processor, further causes the computing system to: receive outcome data provided by a user, the outcome data including at least one of a diagnosis of the first sample, a pathology score of the first sample, or a set of identification data corresponding to at least a plurality of portions of the first sample; and train at least one of a neural network, a convolutional neural network (“CNN”), an artificial intelligence (“AI”) system, or a machine learning system based at least in part on at least one of an analysis of at least one of the captured at least one first image of at least one eye of the user and the captured at least one second image of the optical view of the first sample, or a joint analysis of the collected attention data and the received outcome data, to generate a model used to generate a predicted value (e.g., at least one of a predicted clinical outcome, a predicted attention data, etc.). In some cases, the predicted value may include at least one of a predicted clinical outcome, a predicted attention data, etc. In some cases, the first camera may be further configured to track movement of at least one eye of the user. In some cases, the computing system may be further configured to simultaneously track at least one of one or more coordinate locations, attention duration, or zoom level of the optical view of the first sample.
[0161] According to some embodiments, determining whether at least one eye of the user is focused on a particular region of the optical view of the first sample may include determining whether at least one eye of the user is focused on a particular region of the optical view of the first sample based at least in part on one or more of tracking one or more coordinate locations of an attention gaze, tracking at least one of the movement and zoom level of the optical view of the first sample, or determining that at least one eye of the user is continuing to look at a portion of the optical view of the first sample.
[0162] In some embodiments, the system can further include an audio sensor configured to capture one or more voice notes from a user while the user is viewing the optical view of the first sample. The first set of instructions, when executed by the at least one first processor, can cause the computing system to map the one or more voice notes captured from the user with the at least one second image of the optical view of the first sample and match the one or more captured voice notes with the at least one second image of the optical view of the first sample.
[0163] In yet another aspect, a method may include receiving at least one first image of at least one eye of a user captured by a first camera while the user is viewing an optical view of a first sample through an eyepiece of a microscope; receiving at least one second image of the optical view of the first sample captured by a second camera; analyzing, with a computing system, the at least one first image and the at least one second image to determine whether the at least one eye of the user is focused on a particular region of the optical view of the first sample; tracking, with the computing system, the user's attention based on the analysis; and collecting, with the computing system, attention data based on the tracking.
[0164] In one aspect, a method may include receiving, with a computing system, collected attention data corresponding to a user viewing an optical view of a first sample; receiving, with the computing system, result data provided by the user, the result data including at least one of a diagnosis of the first sample, a pathology score of the first sample, or a set of identification data corresponding to at least a plurality of portions of the first sample; and training at least one of a neural network, a convolutional neural network (“CNN”), an artificial intelligence (“AI”) system, or a machine learning system based at least in part on a joint analysis of the collected attention data and the received result data to generate a model used to generate a prediction.
[0165] In some embodiments, the first sample can be contained in at least one of a microscope slide, a clear sample cartridge, a vial, a tube, a capsule, a flask, a vessel, a receptacle, a microarray, or a microfluidic chip, etc. In some cases, the predictive value can include at least one of a predicted clinical outcome or predictive attention data, etc.
[0166] According to some embodiments, the collection of attention data may occur without interrupting, slowing, or obstructing the user while the user is providing result data either while diagnosing the first sample using the microscope or while diagnosing the image of the first sample displayed on the display screen. In some cases, the collected attention data may include at least one of one or more coordinate locations of at least one particular portion of the optical view of the first sample, an attention duration that the user focuses on at least one particular portion of the optical view of the first sample, a zoom level of the optical view of the first sample while the user focuses on at least one particular portion of the optical view of the first sample, etc.
[0167] In some embodiments, the attention data can be collected based on at least one first image of at least one eye of a user captured by a first camera while the user is viewing an optical view of a first sample through an eyepiece of the microscope. In some cases, the microscope can include two or more of a plurality of mirrors, a plurality of dichroic mirrors, or a plurality of half mirrors that reflect or pass at least one of the optical view of the first sample observed through the eyepiece or the optical view of at least one eye of the user observed through the eyepiece and captured as at least one first image by the first camera.
[0168] Alternatively, the attention data can be collected using an eye-tracking device while the user views a first image of the optical view of the first sample displayed on the display screen. In some embodiments, the method can further include generating, with a computing system, at least one highlight field overlapping the identified at least one particular portion of the at least one first image displayed on the display screen that corresponds to a particular region of the optical view of the first sample. In some cases, the method may further include: using a computing system to display the generated at least one highlight field on the display screen to overlap the identified at least one particular portion of the at least one first image displayed on the display screen that corresponds to the collected attention data; using an eye-tracking device to track the attention data while the user views the first image of the optical view of the first sample displayed on the display screen; and using the computing system to match the tracked attention data with a representation of the at least one first image of the optical view of the first sample displayed on the display screen based at least in part on at least one of: one or more coordinate locations of the at least one particular portion of the optical view of the first sample; an attention duration during which the user focuses on the at least one particular portion of the optical view of the first sample; or a zoom level of the optical view of the first sample while the user focuses on the at least one particular portion of the optical view of the first sample. In some cases, each of the at least one highlight field may include at least one of a color, a shape, a highlight effect, or the like. The highlighting effect may include at least one of an outlining effect, a shadowing effect, a patterning effect, a heat map effect, or a jet color map effect, or the like.
[0169] According to some embodiments, the method may further include tracking attention data using an eye tracking device and simultaneously tracking, using a computing system, at least one of: one or more coordinate locations of the identified at least one particular portion of the at least one second image of the optical view of the first sample; an attention duration during which the user focuses on a particular area of the optical view; or a zoom level of the optical view of the first sample while the user focuses on a particular area of the optical view.
[0170] In some embodiments, the method may further include capturing, using an audio sensor, one or more voice notes from the user while the user is viewing the optical view of the first sample, and mapping, using a computing system, the one or more voice notes captured from the user with at least one third image of the optical view of the first sample to match the one or more captured voice notes with the at least one third image of the optical view of the first sample.
[0171] In another aspect, an apparatus can include at least one processor and a non-transitory computer-readable medium communicatively coupled to the at least one processor. The non-transitory computer-readable medium can have stored thereon computer software including a set of instructions that, when executed by the at least one first processor, cause the apparatus to: receive collected attention data corresponding to a user viewing an optical view of a first sample; receive result data provided by the user, the result data including at least one of a diagnosis of the first sample, a pathology score of the first sample, or a set of identification data corresponding to at least a plurality of portions of the first sample; and train at least one of a neural network, a convolutional neural network (“CNN”), an artificial intelligence (“AI”) system, or a machine learning system based at least in part on a joint analysis of the collected attention data and the received result data to generate a model used to generate a prediction.
[0172] In yet another aspect, a system can include a first camera, a second camera, and a computing system. The first camera can be configured to capture at least one first image of at least one eye of a user while the user views an optical view of a first sample. The second camera can be configured to capture at least one second image of the optical view of the first sample. The computing system can include at least one first processor and a first non-transitory computer-readable medium communicatively coupled to the at least one first processor. The first non-transitory computer-readable medium can have stored thereon computer software including a first set of instructions that, when executed by at least one first processor, causes the computing system to: receive collected attention data corresponding to a user viewing an optical view of a first sample; receive result data provided by the user, the result data including at least one of a diagnosis of the first sample, a pathology score of the first sample, or a set of identification data corresponding to at least a plurality of portions of the first sample; and train at least one of a neural network, a convolutional neural network (“CNN”), an artificial intelligence (“AI”) system, or a machine learning system based at least in part on a joint analysis of the collected attention data and the received result data to generate a model used to generate a prediction.
[0173] Various modifications and additions can be made to the embodiments discussed without departing from the scope of the invention. For example, while the embodiments described above refer to particular features, the scope of this invention also includes embodiments having different combinations of features and embodiments that do not include all of the features described above.
[0174] Reference is now made to embodiments illustrated by the drawings. Figures 1-12 illustrate some features of methods, systems, and apparatuses for performing annotation data collection, and more particularly, as described above, methods, systems, and apparatuses for performing annotation data collection using gaze-based tracking and / or for training artificial intelligence ("AI") systems based on annotation data collected using gaze-based tracking. The methods, systems, and apparatuses illustrated by Figures 1-7 refer to examples of different embodiments that include various components and steps that may be considered alternatives or may be used in conjunction with one another in various embodiments. The descriptions of the exemplary methods, systems, and apparatuses illustrated in Figures 1-12 are provided for illustrative purposes and should not be considered limiting of the scope of various embodiments.
[0175] Referring now to the figures, FIG. 1 is a schematic diagram illustrating a system 100 for implementing annotation data collection using gaze-based tracking, according to various embodiments.
[0176] 1, system 100 may include computing system 105a and a data store or database 110a local to computing system 105a. In some cases, database 110a may be external to computing system 105a but communicatively coupled to computing system 105a. In other cases, database 110a may be integrated within computing system 105a. System 100, according to some embodiments, may further include a microscope 115 and / or a display device 120 that may enable a user 125 to view a sample (e.g., sample 170) or image(s) or video(s) of the sample. System 100 may further include camera(s) 130, one or more audio sensors 135 (optional), and one or more user devices 140 (optional). The camera 130 can capture an image or video of the user 125 (and in some cases, capture an image or video of at least one eye of the user 125) while the user 125 is within the field of view (“FOV”) 130a of the camera 130. In some cases, the camera 130 can include, without limitation, one or more eye-tracking sensors, one or more motion sensors, one or more tracking sensors, or the like. Instead of the camera 130, an eye-tracking device (not shown in FIG. 1 ) can be used to collect attention data when the user views an optical view of the first sample through the eyepiece of the microscope 115 or when viewing an image or video of the first sample displayed on the display screen of the display device 120. In some cases, the one or more audio sensors 135 can include, but are not limited to, one or more microphones, one or more voice recorders, one or more audio recorders, or the like. In some cases, the one or more user devices 140 can include, without limitation, a smartphone, a mobile phone, a tablet computer, a laptop computer, a desktop computer, a monitor, or the like.The computing system 105a may be communicatively coupled (either wirelessly (depicted by a lightning bolt symbol, etc.) or via a wired connection (depicted by a connecting line)) to one or more of the microscope 115, the display device 120, the camera 130 (or eye-tracking device), the one or more audio sensors 135, and / or the one or more user devices 140. The computing system 105a, the database(s) 110a, the microscope 115, the display device 120, the user 125, the camera 130 (or eye-tracking device), the audio sensors 135, and / or the user devices 140 may be located or installed within a work environment 145. The work environment 145 may include, but is not limited to, one of a laboratory, a clinic, a medical facility, a research facility, a study room, or the like.
[0177] System 100 may further include a remote computing system 105b (optional) and corresponding database(s) 110b (optional) that may be communicatively coupled to computing system 105a via network(s) 150. In some cases, system 100 may further include an artificial intelligence (“AI”) system 105c that may be communicatively coupled to computing system 105a or remote computing system 105b via network(s) 150. In some embodiments, AI system 105c may include at least one of, but is not limited to, machine learning system(s), learning algorithm-based system(s), neural network system(s), or the like.
[0178] By way of example only, network(s) 150 may each include a local area network ("LAN"), including, but not limited to, a fiber network, an Ethernet network, a Token-Ring™ network, etc.; a wide area network ("WAN"); a wireless wide area network ("WWAN"); a virtual network such as a virtual private network ("VPN"); the Internet; an intranet; an extranet; a public switched telephone network ("PSTN"); an infrared network; a wireless network, including, but not limited to, a network operating under any of the IEEE 802.11 suite of protocols, the Bluetooth® protocol, and / or any other wireless protocol known in the art; and / or any combination of these and / or other networks. In particular embodiments, network(s) 150 may each include an access network of an Internet service provider ("ISP"). In another embodiment, network(s) 150 may each include a core network of an ISP and / or the Internet.
[0179] According to some embodiments, the microscope 115 includes, but is not limited to, a processor 155, a data store 160a, user interface device(s) 160b (e.g., touch screen(s), buttons, keys, switch toggles, knobs, dials, etc.), a microscope stage 165a (e.g., an XY stage or an XYZ stage, etc.), a first motor 165b (for autonomously controlling the X-direction movement of the microscope stage), a second motor 165c (for autonomously controlling the Y-direction movement of the microscope stage), a third motor 165d (optionally for autonomously controlling the Z-direction movement of the microscope stage), and a third motor 165e (optionally for autonomously controlling the Z-direction movement of the microscope stage). The processor 155 may include at least one of a data store 160a, a user interface device(s) 160b, a first motor 165b, a second motor 165c, a third motor 165d, an FOV camera 175, an eyepiece 180, an eye camera 185, a projection device 190 (optionally), a wired communication system 195a, and a transceiver 195b. The processor 155 may include at least one of a data store 160a, a user interface device(s) 160b, a first motor 165b, a second motor 165c, a third motor 165d, an FOV camera 175, an eye camera 185, a projection device 190, a wired communication system 195a, a transceiver 195b, or the like.
[0180] During operation, microscope 115 can project an optical view of first sample 170 into eyepiece(s) 180 through which at least one eye of user 125 is viewing. Camera 130 (or an eye-tracking device) or eye-gaze camera 185 can capture at least one first image of at least one eye of user 125 as user 125 views the optical view of the first sample (whether projected through eyepiece(s) 180 of microscope 115 or displayed on a display screen such as display device 120). The computing system 105a, the user device(s) 140, the remote computing system(s) 105b, and / or the processor 155 (if a microscope is used) (collectively, "computing systems") can analyze at least one captured first image of at least one eye of the user 125 and at least one captured second image of the optical view of the first sample to determine whether the at least one eye of the user 125 is focused on a particular region of the optical view of the first sample. Based on a determination that the at least one eye of the user 125 is focused on a particular region of the optical view of the first sample, the computing system can identify at least one particular portion of the at least one second image that corresponds to the particular region of the optical view of the first sample. The computing system can collect attention data including the identified at least one particular portion of the at least one second image and store the collected attention data in the database 110a or 110b. According to some embodiments, the collection of attention data can be performed without interrupting, slowing down, or obstructing the user while the user is providing result data either while diagnosing a first sample using the microscope 115 or while diagnosing an image of the first sample displayed on the display screen 120.In some cases, the collected attention data may include, but is not limited to, at least one of: one or more coordinate locations of at least one particular portion of the optical view of the first sample; an attention duration during which a user focuses on at least one particular portion of the optical view of the first sample; a zoom level of the optical view of the first sample while the user focuses on at least one particular portion of the optical view of the first sample; etc. In some cases, the identified at least one particular portion of the at least one second image corresponding to a particular region of the optical view of the first sample may include, but is not limited to, at least one of: one or more particular cells, one or more particular tissues, one or more particular structures, one or more molecules, etc.
[0181] In some embodiments, the computing system may generate at least one highlight field in the at least one second image that covers at least one identified portion of the at least one second image that corresponds to a particular region of the optical view of the first sample. In some cases, each of the at least one highlight field may include, without limitation, at least one of a color, a shape, or a highlight effect, etc., where the highlight effect may include, but is not limited to, at least one of an outlining effect, a shadowing effect, a patterning effect, a heat map effect, or a jet color map effect, etc.
[0182] According to some embodiments, the at least one second image may be displayed on a display screen (e.g., a display screen of display device 120, etc.). Capturing at least one first image of at least one eye of user 125 may include capturing at least one first image of at least one eye of user 125 with camera 130 while user 125 views image(s) or video(s) of the optical view of the first sample displayed as the at least one second image on the display screen of display device 120. An eye-tracking device may be used instead of camera 130 to collect attention data while the user views the image(s) or video of the first sample displayed on the display screen of display device 120. Identifying at least one specific portion of the at least one second image corresponding to a specific region of the optical view of the first sample may include identifying, using a computing system, at least one specific portion of the at least one second image displayed on the display screen that corresponds to a specific region of the optical view of the first sample. The computing system may display the at least one second image on a display screen (e.g., a display screen of the display device 120, etc.) with the generated at least one highlight field covering the identified at least one particular portion of the at least one second image corresponding to a particular region of the optical view of the first sample.
[0183] In some embodiments, the display of the at least one second image on the display screen can shift in response to a command by a user. In some cases, the shifting display of the at least one second image can include at least one of a horizontal shift, a vertical shift, panning, tilting, zooming in, or zooming out of the at least one second image on the display screen, etc. The camera 130 can track the movement of at least one eye of the user 125 while the user 125 is viewing the shifting display of the at least one second image on the display screen. The computing system can match the tracked movement of the at least one eye of the user 125 with the shifting display of the at least one second image on the display screen based at least in part on one or more of the tracked movement of the at least one eye of the user 125, an identified at least one specific portion of the at least one second image corresponding to a specific region of the optical view of the first sample, or at least one of a horizontal shift, a vertical shift, panning, tilting, zooming in, or zooming out of the at least one second image on the display screen. As the user views the shifting display of at least one second image on the display screen of the display device 120, an eye-tracking device can be used instead of using the camera 130 to collect additional attention data.
[0184] Alternatively, microscope 115 can project an optical view of a first sample (e.g., sample 170) into eyepiece 180 through which at least one eye of user 125 is observing. FOV camera 175 can capture at least one second image of the optical view of the first sample. In some cases, capturing at least one first image of at least one eye of user 125 can include capturing at least one first image of at least one eye of user 125 using gaze camera 185 while user 125 is viewing the optical view of the first sample through eyepiece 180. Identifying at least one specific portion of the at least one second image corresponding to a specific region of the optical view of the first sample can include identifying, using a computing system, at least one specific portion of the at least one second image being observed through eyepiece 180 that corresponds to a specific region of the optical view of the first sample. Generating at least one highlight field in the at least one second image overlying the identified at least one particular portion of the at least one second image corresponding to the particular region of the optical view of the first sample can include generating, with a computing system, the at least one highlight field overlapping the identified at least one particular portion of the at least one second image being observed through the eyepiece 180 that corresponds to the particular region of the optical view of the first sample. The computing system can use the projection device 190 to project the generated at least one highlight field to overlap the identified at least one particular portion of the at least one second image being observed through the eyepiece 180 that corresponds to the particular region of the optical view of the first sample.Alternatively or additionally, the computing system may display the at least one second image on a display screen (e.g., a display screen of display device 120, etc.) with the generated at least one highlight field covering the identified at least one particular portion of the at least one second image corresponding to a particular region of the optical view of the first sample.
[0185] In some cases, FOV camera 175 can be one of an infrared ("IR") camera, a back-reflecting IR camera, a visible color camera, a light source, a location photodiode, etc. In some cases, the microscope can include two or more of a plurality of mirrors, a plurality of dichroic mirrors, or a plurality of half mirrors that reflect or pass at least one of, but are not limited to, an optical view of a first sample observed through eyepiece 180, an optical view of at least one eye of user 125 observed through eyepiece 180 and captured as at least one first image by FOV camera 175, or a projection of the generated at least one highlighted field through eyepiece 180 onto at least one eye of user 125 (if projection device 190 is used or present), etc.
[0186] According to some embodiments, the projection of the optical view of the first sample onto the eyepiece 180 can be shifted by at least one of adjusting the microscope stage 165a on which the microscope slide containing the first sample is mounted, changing the objective lens or zoom lens 165f, or adjusting the focus of the eyepiece 180. The camera 130 or 185 can track the movement of at least one eye of the user 125 as the user 125 views the shifted projection of the optical view of the first sample onto the eyepiece 180. The computing system can match the tracked movement of at least one eye of the user 125 with a shifted projection of the optical view of the first sample onto the eyepiece 180 based at least in part on one or more of the tracked movement of at least one eye of the user 125, an identified at least one specific portion of the at least one second image corresponding to a specific region of the optical view of the first sample, or at least one of adjusting the microscope stage 165a on which the microscope slide containing the first sample is placed, changing the objective lens or zoom lens 165f, or adjusting the focus of the eyepiece 180, etc.
[0187] Alternatively or additionally, the one or more audio sensors 135 may capture one or more voice notes from the user 125 as the user 125 views the optical view of the first sample. The computing system may map the one or more voice notes captured from the user 125 with the at least one second image of the optical view of the first sample to match the one or more captured voice notes with the at least one second image of the optical view of the first sample.
[0188] According to some embodiments, the computing system can receive outcome data provided by a user. The outcome data can include at least one of a diagnosis of the first sample, a pathology score of the first sample, or a set of identification data corresponding to at least a plurality of portions of the first sample. The computing system can train the AI system 105c (which can generally include, but is not limited to, at least one of a neural network, a convolutional neural network (“CNN”), a learning algorithm-based system, a machine learning system, etc.) based at least in part on at least one of an analysis of at least one captured first image of at least one eye of the user and at least one captured second image of an optical view of the first sample, or a joint analysis of the collected attention data and the received outcome data, to generate a model used to generate a predicted value. In some embodiments, the predicted value can include, but is not limited to, at least one of a predicted clinical outcome, a predicted attention data, etc.
[0189] In one aspect, the computing system can receive at least one first image of at least one eye of a user captured by a first camera when the user is viewing an optical view of a first sample through an eyepiece of a microscope, can receive at least one second image of the optical view of the first sample captured by a second camera, can analyze the at least one first image and the at least one second image to determine whether the at least one eye of the user is focusing on a particular region of the optical view of the first sample, can track the user's attention based on this analysis, and can collect attention data based on this tracking.
[0190] In some aspects, a semi-weak annotation data collection system (such as system 100) can collect information about a pathologist's visual attention during the pathologist's routine workflow, without interrupting or altering the workflow. Here, annotations are referred to as "weak" in the sense that they only specify the pathologist's attention while making a decision, rather than specific decisions for each location. Weakly supervised methods (in which one or more scores or classifications are assigned to a microscope slide without spatial information) have been shown to provide accuracy comparable to the state-of-the-art performance of fully supervised methods (in which every pixel is annotated in an image). By tracking a pathologist's visual attention while they examine and grade clinical cases, the system, according to various embodiments, can collect a vast amount of valuable annotation data that can be used, for example, to develop tumor localization and classification algorithms.
[0191] In some embodiments, depending on the grading platform, two modalities for tracing and collecting a pathologist's region of interests ("ROI") during pathology slide scoring can be provided: (1) a display device modality; and / or (2) a microscope modality (i.e., (1) a display device modality, (2) a microscope modality, or both). With respect to the display device modality, i.e., when a pathologist scores a microscope slide while viewing the digital slide, the digital pathology weak annotation collection system can be implemented using an eye-tracking system (or eye-tracking camera, etc.) that tracks the pathologist's gaze while viewing a whole slide image ("WSI") on a screen. Additionally, the coordinates (and in some cases the size and magnification) and duration of the field of view ("FOV") that the user zooms in on are stored. The eye-tracking system can integrate information from both the eye-tracking camera (annotated with a jet color map, etc.) and the WSI FOV (displayed as an RGB image, etc.) (such as shown in Figure 2B).
[0192] For microscopic modalities, i.e., when a pathologist scores slides using a microscope, a digital pathology weak annotation collection system can be implemented using a custom eye-tracking system integrated into the microscope, without interfering with the pathologist's ongoing workflow (e.g., as shown in Figure 3A or 3C). The gaze system can be based on optical tracking of the pathologist's eye movements by detecting a back-reflected infrared ("IR") light source from the pathologist's eye with a digital camera while the pathologist continuously observes the sample through the microscope eyepiece. In addition, a separate digital camera can be used to capture the field of view ("FOV") from which the user is currently observing the pathology slide. Eye-tracking software, integrating information from both the gaze camera and the FOV camera, overlays the ROI observed by the user on the collated FOV during the grading process. Finally, the recorded FOV is aligned with the scanned WSI after the grading process, providing localization of the pathologist grading on the WSI through gaze-based interaction.
[0193] In some embodiments, to make "weak" annotations even stronger, audio recording / recognition capabilities can be included.
[0194] These and other features of system 100 (and its components) are described in more detail below with respect to Figures 2-5. Additionally, while various embodiments are described herein with respect to microscopy-related applications, these various embodiments may be applicable, without limitation, to other fields or technologies in which "weak" annotations can be used. These other fields or technologies include, but are not limited to, locating defects in a manufacturing process by tracking the gaze of an operator while they are solving or performing a predetermined task, etc., locating defects in a malfunctioning machine or system.
[0195] 2A and 2B (collectively "FIG. 2") are schematic diagrams illustrating a non-limiting example 200 of annotation data collection using gaze-based tracking, according to various embodiments. FIG. 2A illustrates a side view of a user observing an image of a sample displayed on a display screen while the user's eye(s) are tracked and image-captured, while FIG. 2B illustrates the image of the sample displayed on the display screen, as shown in direction AA of FIG. 2A.
[0196] Referring to non-limiting example 200 of FIG. 2A , computing system 205 (such as computing system 105 a, remote computing system 105 b, and / or user device(s) 140 of FIG. 1 ) can display an image or video of a first sample on a display screen of display device 210 (such as display device 120 of FIG. 1 ). In some cases, the first sample can include, but is not limited to, at least one of: one or more specific cells, one or more specific tissues, one or more specific structures, or one or more molecules, etc. In some cases, the first sample whose image or video is displayed on the display screen of display device 210 can be contained in at least one of a microscope slide, a transparent sample cartridge, a vial, a tube, a capsule, a flask, a vessel, a receptacle, a microarray, a microfluidic chip, etc. A user 215 (such as user 125 of FIG. 1 ) can observe a first sample image or video displayed on the display screen of display device 210 while a camera or gaze camera 220 (such as camera 130 of FIG. 1 ) captures an image or video of user 215 or at least one eye 230 of user 215. In some cases, camera 220 can have a field of view (“FOV”) 225, while at least one eye 230 can have a field of view 235 that defines an angle 235 a that is rotated approximately 360 degrees around an axis perpendicular to the lens of the user's eye(s) 230. An eye-tracking device can be used in place of camera 220 to collect attention data when the user views the first sample image or video displayed on the display screen of display device 210.
[0197] 2B , viewed in direction AA in FIG. 2A , display screen 210 a of display device 210 can display annotation data collection user interface (“UI”) 240. This user interface can display first sample image(s) or video(s) 245 and can provide user interface inputs or icons (including, but not limited to, display control inputs or icons 240 a, audio annotation control inputs or icons 240 b, etc.). In some cases, display control inputs or icons 240 a can include, but are not limited to, at least one of zoom in, zoom out, zoom scroll bars, focus in, focus out, direction shift control (e.g., shift up, shift down, shift right, shift left, shift right-up, shift left-up, shift right-down, shift left-down, etc.), autofocus, center out or center focus out, color map effect options or highlighting effect options, a single screenshot, or multiple screenshots, etc. In some cases, the audio annotation control inputs or icons 240b may include, but are not limited to, at least one of record, play or pause, stop, mute, audio on, or an audio scroll bar, etc. Also shown in Figure 2B is the camera 220 of Figure 2A.
[0198] During operation, the camera 220 may capture at least one first image of at least one eye 230 of the user 215 while the user 215 views an optical view 245 of a first sample displayed on a display screen 210a, such as the display device 210. The computing system 205 may analyze the captured at least one first image of the at least one eye 230 of the user 215 and the at least one second image of the optical view 245 of the first sample to determine whether the at least one eye 230 of the user 215 is focused on a particular region of the optical view 245 of the first sample displayed on the display screen 210a of the display device 210. An eye-tracking device may be used instead of the camera 220 to collect attention data while the user views an image or video of the first sample displayed on the display screen of the display device 210. Based on a determination that at least one eye 230 of the user 215 is focused on a particular region of the optical view 245 of the first sample displayed on the display screen 210a of the display device 210, or based on the collected attention data, the computing system 205 can identify at least one particular portion of the at least one second image displayed on the display screen 210a of the display device 210 that corresponds to the particular region of the optical view 245 of the first sample. The computing system 205 can generate at least one highlight field 250 in the at least one second image that covers the identified at least one particular portion of the at least one second image that corresponds to the particular region of the optical view 245 of the first sample. The computing system 205 can display the at least one second image on the display screen 210a of the display device 210 with the generated at least one highlight field 250 that covers the identified at least one particular portion of the at least one second image that corresponds to the particular region of the optical view 245 of the first sample.
[0199] In some embodiments, each of the at least one highlight field 250 may include, but is not limited to, at least one of a color, a shape, or a highlight effect, etc., and the highlight effect may include, but is not limited to, at least one of an outlining effect, a shadowing effect, a patterning effect, a heat map effect, a jet color map effect, etc. In some cases, the identified at least one specific portion of the at least one second image corresponding to a specific region of the optical view 245 of the first sample may include, but is not limited to, at least one of one or more specific cells, one or more specific tissues, one or more specific structures, or one or more molecules, etc.
[0200] In some embodiments, the display of the at least one second image on the display screen 210a of the display device 210 can shift in response to a command (whether a verbal command, a keystroke command, a user interface command, etc.) by the user 215. In some cases, the shifting display of the at least one second image can include, without limitation, at least one of a horizontal shift, a vertical shift, a pan, a tilt, a zoom-in, or a zoom-out of the at least one second image on the display screen 210a of the display device 210, etc. The camera 220 can track the movement of at least one eye 230 of the user 215 as the user 215 views the shifting display of the at least one second image on the display screen 210a of the display device 210. The computing system 205 can match the tracked movement of at least one eye 230 of the user 215 with the shifting display of the at least one second image on the display screen 210a of the display device 210 based at least in part on one or more of the tracked movement of the at least one eye 230 of the user 215, the identified at least one particular portion of the at least one second image corresponding to a particular region of the optical view of the first sample, or at least one of a horizontal shift, a vertical shift, a pan, a tilt, a zoom-in, or a zoom-out, etc., of the at least one second image on the display screen. An eye-tracking device can be used instead of using the camera 220 to collect additional attention data when the user is viewing the shifting display of the at least one second image on the display screen 210a of the display device 210.
[0201] 3A-3D (collectively "FIG. 3") are schematic diagrams illustrating various other non-limiting examples 300 and 300' of annotation data collection using gaze-based tracking, according to various embodiments. FIG. 3A illustrates a side view of a microscope whose eyepiece is the eyepiece through which a user views an image of a sample, while FIG. 3B illustrates an image of the sample being projected through the eyepiece, shown in the B-B direction of FIG. 3A. FIG. 3C illustrates example 300', an alternative example to example 300 shown in FIG. 3A, while FIG. 3D illustrates a display screen on which annotated image(s) or annotated video(s) of the sample are displayed.
[0202] Referring to the non-limiting example 300 of FIG. 3A, a computing system 305 (similar to computing system 105a, remote computing system 105b, user device(s) 140, and / or processor 155 of FIG. 1, etc.) can be integrated within microscope 310 (not shown) or can be external but communicatively coupled to microscope 310 (shown in FIG. 3A) and can control various operations of microscope 310. As shown in FIG. 3A, a microscope slide 315 containing a first sample can be positioned on an adjustable microscope stage 320 (e.g., an XY stage or an XYZ stage similar to microscope stage 165a of FIG. 1), and light from a light source 325 (similar to light source 165 of FIG. 1) passes through stage 320, through microscope slide 315, through one of at least one objective or zoom lens 330 (similar to objective or zoom lens(es) 165f of FIG. 1), reflected from or passing through a plurality of mirrors, dichroic mirrors, and / or half mirrors 335, through an eyepiece 340 (similar to eyepiece 180 of FIG. 1), and projected onto at least one eye 345 of a user.
[0203] Microscope 310 may include a field of view ("FOV") camera 350 (such as FOV camera 175 in FIG. 1) that can be used to capture image(s) or video(s) of a first sample contained in microscope slide 315 along a light beam 355 (shown in FIG. 3A as a medium-shaded line 355, for example). Light beam 355 may travel from light source 325 through stage 320, through the first sample contained in microscope slide 315, through at least one objective lens or one of zoom lenses 330, and be reflected off mirrors, dichroic mirrors, and / or half mirrors 335b and 335c to FOV camera 350. In other words, FOV camera 350 may capture image(s) or video(s) (along light beam 355) of the first sample contained in microscope slide 315 that is backlit by light source 325. Eyepiece 340 can collect light for projected image(s) or video(s) of a first sample contained in microscope slide 315 projected by light source 325. Light beam 355 can travel from light source 325 through stage 320, through the first sample contained in microscope slide 315, through one of at least one objective lens or zoom lens 330, reflect off mirror 335c, through half mirror 335b, reflect off mirror 335a, through eyepiece 340, and into at least one of the user's eyes 345. In other words, the user can observe image(s) or video(s) (along light beam 355) of the first sample contained in microscope slide 315 that is back-illuminated by light source 325.
[0204] Microscope 310 may further include an eye camera 360 (such as similar to eye camera 185 of FIG. 1 ) that can be used to capture image(s) or video(s) of at least one user's eye 345 along a light beam 365 (shown in FIG. 3A as a thick, darkly shaded line 365, for example). Light beam 365 may pass from at least one user's eye 345 through eyepiece 340 and be reflected off mirror 335 a, dichroic mirror 335 b, and / or half mirror 335 d to eye camera 360. According to some embodiments, eye camera 360 may include one of, but is not limited to, an infrared (“IR”) camera, a back-reflecting IR camera, a visible color camera, a light source, or a location photodiode, etc.
[0205] During operation, the microscope 310 can project an optical view of a first sample into the eyepiece 340 through which the at least one eye 345 of a user is observing. The line-of-sight camera 360 can capture at least one first image of the at least one eye 345 of the user as the user views the optical view of the first sample observed through the eyepiece 340 of the microscope 310. The FOV camera 350 can capture at least one second image of the optical view of the first sample. The computing system 305 can analyze the captured at least one first image of the at least one eye 345 of the user and the captured at least one second image of the optical view of the first sample to determine whether the at least one eye 345 of the user is focusing on a particular region of the optical view of the first sample. Based on a determination that at least one eye 345 of the user is focused on a particular region of the optical view of the first sample, the computing system 305 may identify at least one particular portion of the at least one second image that corresponds to the particular region of the optical view of the first sample. The computing system 305 may collect attention data including the identified at least one particular portion of the at least one second image and store the collected attention data in a database (e.g., such as database(s) 110a or 110b of FIG. 1 ). According to some embodiments, the collection of attention data may occur without interrupting, slowing down, or disturbing the user as the user provides result data while diagnosing the first sample using the microscope.In some cases, the collected attention data may include, but is not limited to, at least one of: one or more coordinate locations of at least one particular portion of the optical view of the first sample; an attention duration during which a user focuses on at least one particular portion of the optical view of the first sample; a zoom level of the optical view of the first sample while the user focuses on at least one particular portion of the optical view of the first sample; etc. In some cases, the identified at least one particular portion of the at least one second image corresponding to a particular region of the optical view of the first sample may include, but is not limited to, at least one of: one or more particular cells, one or more particular tissues, one or more particular structures, one or more molecules, etc.
[0206] In some embodiments, the computing system 305 may generate at least one highlight field in the at least one second image that overlaps with the identified at least one particular portion of the at least one second image that corresponds to a particular region of the optical view of the first sample. In some cases, each of the at least one highlight field may include, without limitation, at least one of a color, a shape, or a highlight effect, etc., where the highlight effect may include, but is not limited to, at least one of an outlining effect, a shadowing effect, a patterning effect, a heat map effect, or a jet color map effect, etc.
[0207] According to some embodiments, microscope 310 may further include a projection device 370 (such as similar to projection device 190 of FIG. 1 ) that can be used to project the generated at least one highlight field along a light beam 375 (shown in FIG. 3A as a thick, lightly shaded line 375, for example) through eyepiece 340 onto at least one eye 345 of the user. Light beam 375 may travel from projection device 370, reflect off mirror 335 e, pass through half mirror 335 d, reflect off half mirror 335 b, reflect off mirror 335 a, pass through eyepiece 340, and onto at least one eye 345 of the user.
[0208] Figure 3B shows an optical view 380 of the first sample as viewed through the eyepiece 340 of the microscope 310 along the B-B direction of Figure 3A. The optical view 380 includes at least one second image of the first sample 385. As shown in Figure 3B, the optical view 380, in some embodiments, can further include one or more generated highlight fields 390 (in this case, depicted or embodied by a jet color map or the like) that highlight the portion of the first sample 385 on which the user's eye(s) are focused. For example, with respect to a jet color map embodiment, a red region of the color map may represent the highest occurrence or longest duration of eye focus or attention, while a yellow or orange region of the color map may represent the next highest occurrence or next longest duration of eye focus or attention, a green region of the color map may represent a lower occurrence or shorter duration of eye focus or attention, and a blue or purple region of the color map may represent the lowest occurrence or shortest duration of eye focus or attention, but statistically higher than when focus or attention is unfocused or in a scanning state, etc.
[0209] Referring to FIG. 3C, instead of the microscope 310 of the non-limiting example 300 of FIG. 3A, the microscope 310' of the non-limiting example 300' of FIG. 3C can exclude the projection device 370 and the mirror 335e, but can otherwise be similar to the microscope 310 of FIG. 3A.
[0210] In particular, computing system 305 (similar to computing system 105a, remote computing system 105b, user device(s) 140, and / or processor 155 of FIG. 1 ) may be integrated within microscope 310′ (not shown) or may be external but communicatively coupled to microscope 310′ (shown in FIG. 3C ) and may control various operations of microscope 310′. As shown in FIG. 3C, a microscope slide 315 containing a first sample can be positioned on an adjustable microscope stage 320 (e.g., an XY stage or an XYZ stage similar to microscope stage 165a of FIG. 1), and light from a light source 325 (similar to light source 165 of FIG. 1) passes through stage 320, through microscope slide 315, through one of at least one objective or zoom lens 330 (similar to objective or zoom lens(es) 165f of FIG. 1), reflected from or passing through a plurality of mirrors, dichroic mirrors, and / or half mirrors 335, through an eyepiece 340 (similar to eyepiece 180 of FIG. 1), and projected onto at least one eye 345 of a user.
[0211] Microscope 310′ can include an FOV camera 350 (such as FOV camera 175 in FIG. 1 ) that can be used to capture image(s) or video(s) of a first sample contained in microscope slide 315 along a light beam 355 (shown in FIG. 3C as a medium-shaded line 355, for example). Light beam 355 can travel from light source 325 through stage 320, through the first sample contained in microscope slide 315, through at least one objective lens or one of zoom lenses 330, and be reflected off mirrors, dichroic mirrors, and / or half mirrors 335b and 335c to FOV camera 350. In other words, FOV camera 350 can capture image(s) or video(s) (along light beam 355) of the first sample contained in microscope slide 315 that is back-illuminated by light source 325. Eyepiece 340 can collect light for projected image(s) or video(s) of a first sample contained in microscope slide 315 projected by light source 325. Light beam 355 can travel from light source 325 through stage 320, through the first sample contained in microscope slide 315, through one of at least one objective lens or zoom lens 330, reflect off mirror 335c, through half mirror 335b, reflect off mirror 335a, through eyepiece 340, and into at least one of the user's eyes 345. In other words, the user can observe image(s) or video(s) (along light beam 355) of the first sample contained in microscope slide 315 that is back-illuminated by light source 325.
[0212] Microscope 310′ may further include an eye camera 360 (such as similar to eye camera 185 of FIG. 1 ) that can be used to capture image(s) or video(s) of at least one user's eye 345 along a light beam 365 (shown in FIG. 3C as a thick, darkly shaded line 365, for example). Light beam 365 may pass from at least one user's eye 345 through eyepiece 340 and be reflected off mirror 335a, dichroic mirror 335b, and / or half mirror 335d (i.e., mirror 335a, dichroic mirror 335b, and / or half mirror 335d) to eye camera 360. According to some embodiments, eye camera 360 may include one of, but is not limited to, an infrared (“IR”) camera, a back-reflecting IR camera, a visible color camera, a light source, a location photodiode, or the like.
[0213] In operation, similar to example 300 of FIG. 3A , microscope 310 can project an optical view of a first sample into eyepiece 340 through which at least one eye 345 of a user is observing. Line-of-sight camera 360 can capture at least one first image of at least one eye 345 of a user as the user views the optical view of the first sample observed through eyepiece 340 of microscope 310′. FOV camera 350 can capture at least one second image of the optical view of the first sample. Computing system 305 can analyze the captured at least one first image of at least one eye 345 of the user and the captured at least one second image of the optical view of the first sample to determine whether at least one eye 345 of the user is focusing on a particular region of the optical view of the first sample. Based on a determination that at least one eye 345 of the user is focused on a particular region of the optical view of the first sample, the computing system 305 may identify at least one particular portion of the at least one second image that corresponds to the particular region of the optical view of the first sample. The computing system 305 may collect attention data including the identified at least one particular portion of the at least one second image and store the collected attention data in a database (e.g., such as database(s) 110a or 110b of FIG. 1 ). According to some embodiments, the collection of attention data may occur without interrupting, slowing down, or disturbing the user as the user provides result data while diagnosing the first sample using the microscope.In some cases, the collected attention data may include, but is not limited to, at least one of: one or more coordinate locations of at least one particular portion of the optical view of the first sample; an attention duration during which a user focuses on at least one particular portion of the optical view of the first sample; a zoom level of the optical view of the first sample while the user focuses on at least one particular portion of the optical view of the first sample; etc. In some cases, the identified at least one particular portion of the at least one second image corresponding to a particular region of the optical view of the first sample may include, but is not limited to, at least one of: one or more particular cells, one or more particular tissues, one or more particular structures, one or more molecules, etc.
[0214] In some embodiments, the computing system 305 may generate at least one highlight field in the at least one second image that overlaps with the identified at least one particular portion of the at least one second image that corresponds to a particular region of the optical view of the first sample. In some cases, each of the at least one highlight field may include, without limitation, at least one of a color, a shape, or a highlight effect, etc., where the highlight effect may include, but is not limited to, at least one of an outlining effect, a shadowing effect, a patterning effect, a heat map effect, or a jet color map effect, etc.
[0215] Unlike example 300 of FIG. 3A , in which at least one generated highlight field is projected through eyepiece 340 via mirror, dichroic mirror, and / or half mirror 335 to at least one of the user's eyes 345, computing system 305 of example 300′ of FIG. 3C can display image(s) or video(s) of first sample 385 (shown in FIG. 3D ) on display screen 395a of display device 395. Similar to example 300 of FIG. 3B , the optical view of example 300′ of FIG. 3D can further include one or more generated highlight fields 390 (in this case, depicted or embodied by a jet color map or the like) that highlight the portion of first sample 385 on which the user's eye(s) are focused. For example, with respect to a jet color map embodiment, a red region of the color map may represent the highest occurrence or longest duration of eye focus or attention, while a yellow or orange region of the color map may represent the next highest occurrence or next longest duration of eye focus or attention, a green region of the color map may represent a lower occurrence or shorter duration of eye focus or attention, and a blue or purple region of the color map may represent the lowest occurrence or shortest duration of eye focus or attention, but statistically higher than when focus or attention is unfocused or in a scanning state, etc.
[0216] Similar to the display of image(s) or video(s) of first sample 245 on display screen 210a of display device 210 of Figure 2B, image(s) or video(s) of first sample 385 can be displayed in annotation data collection user interface ("UI") 380' displayed on display screen 395a of display device 395. Similar to the example of Figure 2B, annotation data collection UI 380' of Figure 3D can provide user interface inputs or icons (including, but not limited to, display control input or icon 380a', audio annotation control input or icon 380b', etc.). In some cases, the display control inputs or icons 380a' may include, but are not limited to, at least one of zoom in, zoom out, zoom scroll bar, focus in, focus out, direction shift control (e.g., shift up, shift down, shift right, shift left, shift right up, shift left up, shift right down, shift left down, etc.), autofocus, center out or center focus out, color map effect options or highlighting effect options, single screenshot or multiple screenshots, etc. In some cases, the audio annotation control inputs or icons 380b' may include, but are not limited to, at least one of record, play or pause, stop, mute, audio on, audio scroll bar, etc.
[0217] In some embodiments, the display of image(s) or video(s) of the first sample 385 on the display screen 395a of the display device 395 of FIG. 3D can be added to the optical view 380 of the first sample 385 observed through the eyepiece 340 of the microscope 310 of FIG. 3B.
[0218] 4A-4D (collectively "FIG. 4") are flow diagrams illustrating a method 400 for performing annotation data collection using gaze-based tracking, according to various embodiments. Method 400 in FIG. 4A continues in FIG. 4B following circular marker "A," and continues from FIG. 4A to FIG. 4C following circular marker "B." Method 400 in FIG. 4B continues in FIG. 4C following circular marker "C."
[0219] These techniques and procedures are shown and / or described in a particular order for illustrative purposes, but it should be understood that certain procedures may be rearranged and / or omitted within various embodiments. Moreover, although the method 400 illustrated by Figure 4 may be implemented by or using the systems, examples, or embodiments 100, 200, and 300 (or components thereof) of Figures 1, 2, and 3, respectively (in some cases, systems, examples, or embodiments 100, 200, and 300 are described below), such methods may also be implemented using any suitable hardware (or software) implementation. Similarly, although each of the systems, examples, or embodiments 100, 200, and 300 (or components thereof) of Figures 1, 2, and 3, respectively, can operate according to the method 400 illustrated by Figure 4 (e.g., by executing instructions embodied on a computer-readable medium), each of the systems, examples, or embodiments 100, 200, and 300 of Figures 1, 2, and 3, respectively, can also operate according to other modes of operation and / or perform other suitable procedures.
[0220] 4A , method 400 may include, at block 405, using a microscope to project an optical view of a first sample into an eyepiece through which at least one eye of a user is viewing. In some embodiments, the first sample may be contained in at least one of a microscope slide, a transparent sample cartridge, a vial, a tube, a capsule, a flask, a vessel, a receptacle, a microarray, a microfluidic chip, or the like. According to some embodiments, the microscope may include, but is not limited to, two or more of a plurality of mirrors, a plurality of dichroic mirrors, a plurality of half mirrors, or the like that reflect or transmit at least one of the optical view of the first sample viewed through the eyepiece or the optical view of at least one eye of the user viewed through the eyepiece and captured as at least one first image by a first camera.
[0221] The method 400 may further include capturing, with a first camera, at least one first image of at least one eye of the user while the user is viewing the optical view of the first sample through the eyepiece (block 410), and capturing, with a second camera, at least one second image of the optical view of the first sample (block 415).
[0222] At optional block 420, method 400 may include tracking, with the first camera, movement of at least one eye of the user. At optional block 425, method 400 may further include, with the computing system, simultaneously tracking at least one of: one or more coordinate locations of the identified at least one particular portion of the at least one second image; an attention duration during which the user focuses on a particular region of the optical view; a zoom level of the optical view of the first sample while the user focuses on the particular region of the optical view; etc. In some cases, the first camera may include, but is not limited to, one of an infrared (“IR”) camera, a back-reflective IR camera, a visible color camera, a light source, or a location photodiode, etc.
[0223] Method 400 may further include analyzing, using a computing system, at least one captured first image of at least one eye of the user and at least one captured second image of the optical view of the first sample to determine whether the at least one eye of the user is focused on a particular region of the optical view of the first sample (block 430), and based on a determination that the at least one eye of the user is focused on a particular region of the optical view of the first sample, identifying, using the computing system, at least one specific portion of the at least one second image viewed through the eyepiece that corresponds to the particular region of the optical view of the first sample (block 435). According to some embodiments, the identified at least one specific portion of the at least one second image that corresponds to the particular region of the optical view of the first sample may include, but is not limited to, at least one of one or more specific cells, one or more specific tissues, one or more specific structures, one or more molecules, or the like. In some embodiments, identifying the at least one particular portion of the at least one second image may include determining, with a computing system, a coordinate location within the at least one second image of the optical view that corresponds to the identified at least one particular portion of the at least one second image.
[0224] Method 400 may include, at block 440, using a computing system to collect attention data including the identified at least one particular portion of the at least one second image. At block 445, method 400 may include storing the collected attention data in a database. Method 400 may continue processing at optional block 450 of FIG. 4B following circular marker "A" or may continue processing at block 460 of FIG. 4C following circular marker "B."
[0225] At optional block 450 in FIG. 4B (following circular marker "A"), method 400 may include capturing one or more voice notes from the user using an audio sensor while the user is viewing the optical view of the first sample. Method 400 may further include mapping, using a computing system, the one or more voice notes captured from the user with at least one second image of the optical view of the first sample to match the one or more captured voice notes with the at least one second image of the optical view of the first sample (optional block 455). Method 400 may continue processing at block 465 in FIG. 4C (following circular marker "C").
[0226] 4C (following the circular marker "B"), method 400 may include, at block 460, receiving, using a computing system, result data provided by a user, the result data including at least one of a diagnosis of the first sample, a pathology score of the first sample, or a set of identification data corresponding to at least a plurality of portions of the first sample. Method 400 may further include, at block 465, training at least one of a neural network, a convolutional neural network ("CNN"), an artificial intelligence ("AI") system, or a machine learning system to generate a model used to generate the predictions based at least in part on at least one of an analysis of at least one captured first image of at least one eye of the user and at least one captured second image of the optical view of the first sample (and, in some cases, captured voice notes mapped to the captured at least one second image), or a joint analysis of collected attention data and the received result data. In some embodiments, the predictive value may include at least one of, but is not limited to, a predicted clinical outcome or a predicted attention data, etc. According to some embodiments, the collection of attention data may occur without interrupting, slowing down, or otherwise disturbing the user as the user provides result data while diagnosing the first sample using the microscope.
[0227] Referring to FIG. 4D , determining whether at least one eye of the user is focusing on a particular region of the optical view of the first sample (block 430) may include, in block 470, determining whether at least one eye of the user is focusing on a particular region of the optical view of the first sample based at least in part on at least one of the coordinate locations of the identified at least one particular portion of the at least one second image (block 470a), the attention duration during which the user focuses on the particular region of the optical view (block 470b), or the zoom level of the optical view of the first sample while the user is focusing on the particular region of the optical view (block 470c).
[0228] 5A-5D (collectively "FIG. 5") are flow diagrams illustrating a method 500 for performing annotation data collection using gaze-based tracking, according to various embodiments. The method 500 of FIG. 5B continues into FIG. 5C or FIG. 5D following circular marker "A," and returns from FIG. 5C or FIG. 5D to FIG. 5A following circular marker "B."
[0229] These techniques and procedures are shown and / or described in a particular order for illustrative purposes, but it should be understood that certain procedures may be rearranged and / or omitted within various embodiments. Moreover, although the method 500 illustrated by Figure 5 may be implemented by or using the systems, examples, or embodiments 100, 200, and 300 (or components thereof) of Figures 1, 2, and 3, respectively (in some cases, systems, examples, or embodiments 100, 200, and 300 are described below), such methods may also be implemented using any suitable hardware (or software) implementation. Similarly, although each of the systems, examples, or embodiments 100, 200, and 300 (or components thereof) of Figures 1, 2, and 3, respectively, can operate according to the method 500 illustrated by Figure 5 (e.g., by executing instructions embodied on a computer-readable medium), each of the systems, examples, or embodiments 100, 200, and 300 of Figures 1, 2, and 3, respectively, can also operate according to other modes of operation and / or perform other suitable procedures.
[0230] 5A , method 500 may include, at block 505, receiving, with a computing system, collected attention data corresponding to a user viewing an optical view of a first sample. At block 510, method 500 may include, with the computing system, receiving result data provided by the user, the result data including at least one of a diagnosis of the first sample, a pathology score of the first sample, or a set of identification data corresponding to at least a plurality of portions of the first sample. At block 515, method 500 may further include training at least one of a neural network, a convolutional neural network (“CNN”), an artificial intelligence (“AI”) system, or a machine learning system based at least in part on a joint analysis of the collected attention data and the received result data to generate a model used to generate the prediction.
[0231] In some embodiments, the first sample can be contained in at least one of a microscope slide, a clear sample cartridge, a vial, a tube, a capsule, a flask, a vessel, a receptacle, a microarray, or a microfluidic chip, etc. According to some embodiments, the predictive value can include, but is not limited to, at least one of a predicted clinical outcome or a predicted attention data, etc.
[0232] Referring to FIG. 5B, method 500 may further include tracking attention data using an eye tracking device (block 520) and simultaneously tracking using a computing system at least one of: one or more coordinate locations of the identified at least one particular portion of at least one second image of the optical view of the first sample; an attention duration during which the user focuses on a particular area of the optical view; or a zoom level of the optical view of the first sample while the user focuses on a particular area of the optical view (block 525).
[0233] In some cases, method 500 may further include capturing one or more voice notes from the user using an audio sensor while the user is viewing the optical view of the first sample (optional block 530), and mapping, using a computing system, the one or more voice notes captured from the user with at least one third image of the optical view of the first sample to match the one or more captured voice notes with the at least one third image of the optical view of the first sample (optional block 535). Method 500 may continue with the process at block 540 of FIG. 5C or the process at block 545 of FIG. 5D following circular marker "A."
[0234] At block 540 of FIG. 5C (following circular marker “A”), method 500 may include collecting attention data based on at least one first image of at least one eye of a user captured by a first camera while the user views an optical view of a first sample through an eyepiece of the microscope. In some embodiments, the microscope may include, but is not limited to, two or more of mirrors, dichroic mirrors, half mirrors, or the like that reflect or pass at least one of the optical view of the first sample observed through the eyepiece or the optical view of at least one eye of the user observed through the eyepiece and captured as at least one first image by the first camera. Method 500 may return to the process at block 505 of FIG. 5A following circular marker “B.”
[0235] Alternatively, at block 545 of FIG. 5D (following the circular marker "A"), the method 500 may include collecting attention data using an eye-tracking device while the user views a first image of the optical view of the first sample displayed on the display screen.
[0236] According to some embodiments, the collection of attention data can be performed without interrupting, slowing, or disturbing the user while the user is providing result data either while diagnosing the first sample using the microscope or while diagnosing the image of the first sample displayed on the display screen. In some embodiments, the collected attention data can include, but is not limited to, at least one of: one or more coordinate locations of at least one particular portion of the optical view of the first sample; an attention duration that the user focuses on at least one particular portion of the optical view of the first sample; a zoom level of the optical view of the first sample while the user focuses on at least one particular portion of the optical view of the first sample;
[0237] By way of example only, in some cases, method 500 may include: generating, with a computing system, at least one highlight field overlapping an identified at least one particular portion of at least one first image displayed on the display screen that corresponds to a particular region of the optical view of the first sample (optional block 550); displaying, with the computing system, the generated at least one highlight field on the display screen so as to overlap an identified at least one particular portion of the at least one first image displayed on the display screen that corresponds to the collected attention data (optional block 555); and displaying, with the computing system, at least one highlight field overlapping an identified at least one particular portion of the at least one first image displayed on the display screen that corresponds to the collected attention data (optional block 556); and The method may further include tracking attention data using a line tracking device (optional block 560) and matching, using a computing system, the tracked attention data with a display of at least one first image of the optical view of the first sample displayed on a display screen (optional block 565) based at least in part on at least one of: one or more coordinate locations of at least one particular portion of the optical view of the first sample; an attention duration during which the user focuses on at least one particular portion of the optical view of the first sample; or a zoom level of the optical view of the first sample while the user focuses on the at least one particular portion of the optical view of the first sample. In some cases, each of the at least one highlight field may include at least one of, but is not limited to, a color, a shape, or a highlight effect, etc. In some cases, the highlight effect may include, but is not limited to, at least one of, an outlining effect, a shadowing effect, a patterning effect, a heat map effect, a jet color map effect, etc.
[0238] The method 500 may return to the process at block 505 of FIG. 5A following the circular marker "B."
[0239] Figure 6 is a block diagram illustrating an exemplary computer or system hardware architecture according to various embodiments. Figure 6 provides a schematic illustration of one embodiment of service provider system hardware computer system 600 that can execute methods provided by various other embodiments as described herein and / or perform the functions of the computer or hardware systems (i.e., computing systems 105a, 105b, 205, and 305, microscopes 115, 310, and 310', display devices 120, 210, and 395, and user device(s) 140, etc.) as described above. Note that Figure 6 is only meant to provide a generalized illustration of the various components, and one or more of each component (or none of them) may be utilized, as appropriate. Thus, Figure 6 broadly illustrates how individual system elements may be implemented in a relatively separate or more integrated manner.
[0240] Computer or hardware system 600 may represent one embodiment of the computers or hardware systems described above with respect to Figures 1-5 (i.e., computing systems 105a, 105b, 205, and 305, microscopes 115, 310, and 310', display devices 120, 210, and 395, and user device(s) 140, etc.), and is shown to include hardware elements that may be electrically coupled (or otherwise in communication as appropriate) via a bus 605. The hardware elements may include one or more processors 610, including, but not limited to, one or more general-purpose processors and / or one or more special-purpose processors (e.g., microprocessors, digital signal processing chips, graphics acceleration processors, etc.), one or more input devices 615, including, but not limited to, a mouse, a keyboard, etc., and one or more output devices 620, including, but not limited to, a display device, a printer, etc.
[0241] The computer or hardware system 600 may further include (and / or be in communication with) one or more storage devices 625. The storage devices 625 may include, without limitation, local and / or network-accessible storage devices, and / or solid-state storage devices such as, without limitation, disk drives, drive arrays, optical storage devices, random access memory ("RAM") and / or read-only memory ("ROM"), which may be programmable, flash-updateable, etc. Such storage devices may be configured to implement any suitable data store, including, without limitation, various file systems, database structures, etc.
[0242] The computer or hardware system 600 may also include a communications subsystem 630. The communications subsystem 630 may include, but is not limited to, a modem, a network card (wireless or wired), an infrared communications device, a wireless communications device and / or chipset (such as a Bluetooth device, an 802.11 device, a WiFi device, a WiMax device, a WWAN device, a cellular communications facility, etc.). The communications subsystem 630 may enable the exchange of data with a network (such as the networks described below, to name but a few), other computers or hardware systems, and / or any other devices described herein. In many embodiments, the computer or hardware system 600 further comprises a working memory 635, which may include a RAM device or a ROM device as described above.
[0243] The computer or hardware system 600 may also include software elements shown as presently residing in working memory 635 and / or may be designed to implement methods and / or configure systems provided by other embodiments as described herein. These software elements include other code, such as an operating system 640, device drivers, executable libraries, and / or one or more application programs 645, which may include computer programs (including, but not limited to, hypervisors, VMs, etc.) provided by various embodiments. By way of example only, one or more procedures described with respect to the method(s) described above may be embodied as code and / or instructions executable by a computer (and / or a processor within a computer), in which case, in one aspect, such code and / or instructions may be used to configure and / or adapt a general-purpose computer (or other device) to perform one or more operations in accordance with the described methods.
[0244] These sets of instructions and / or code (i.e., instructions and / or code) may be coded and / or stored on a non-transitory computer-readable storage medium, such as storage device(s) 625 described above. In some cases, the storage medium may be incorporated within a computer system, such as system 600. In other embodiments, the storage medium may be separate from the computer system (i.e., a removable medium such as a compact disc) and / or may be provided in an installation package such that the storage medium can be used to configure and / or program a general-purpose computer with the stored instructions / code. These instructions may take the form of executable code that can be executed by computer or hardware system 600 and / or may take the form of source code and / or installable code that takes the form of executable code when compiled and / or installed on computer or hardware system 600 (e.g., using any of various publicly available compilers, installation programs, compression / decompression utilities, etc.).
[0245] It will be apparent to those skilled in the art that significant modifications can be made according to particular requirements. For example, customized hardware (such as programmable logic controllers, field programmable gate arrays, application specific integrated circuits, etc.) could also be used, and / or particular elements could be implemented in hardware, software (including portable software such as applets), or both. Furthermore, connection to other computing devices, such as network input / output devices, could be used.
[0246] As noted above, in one aspect, some embodiments may employ a computer or hardware system (such as computer or hardware system 600) to perform methods according to various embodiments of the present invention. According to one set of embodiments, some or all of the steps of such methods are performed by computer or hardware system 600 in response to processor 610 executing one or more sequences of one or more instructions (which may be embedded in other code, such as operating system 640 and / or application program 645) contained in working memory 635. Such instructions may be read into working memory 635 from another computer-readable medium, such as one or more of storage device(s) 625. By way of example only, execution of the sequences of instructions contained in working memory 635 may cause processor(s) 610 to perform one or more steps of the methods described herein.
[0247] As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any medium that participates in providing data that causes a machine to operate in a specific fashion. In one embodiment implemented using a computer or hardware system 600, various computer-readable media may be involved in providing instructions / code to the processor(s) 610 for execution and / or may be used to store and / or carry such instructions / code (e.g., as a signal). In many embodiments, computer-readable media are non-transitory, physical, and / or tangible storage media. In some embodiments, computer-readable media may take many forms, including, but not limited to, non-volatile media, volatile media, etc. Non-volatile media include, for example, optical and / or magnetic disks, such as storage device(s) 625. Volatile media include, but are not limited to, dynamic memory, such as working memory 635. In some alternative embodiments, the computer-readable medium may take the form of a transmission medium including, but not limited to, coaxial cables, copper wire, and fiber optics, including the wires that comprise the various components of the communications subsystem 630 (and / or the medium through which the communications subsystem 630 communicates with other devices) in addition to the bus 605. In one alternative set of embodiments, the transmission medium may also take the form of waves (including, but not limited to, radio waves, acoustic waves, and / or light waves, such as those generated during radio wave and infrared data communications).
[0248] Common forms of physical and / or tangible computer readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape or any other magnetic media, CD-ROMs, any other optical media, punch cards, paper tape, any other physical media with patterns of holes, RAM, PROMs, and EPROMs, flash EPROMs, any other memory chips or cartridges, carrier waves as described below, or any other medium from which a computer can read instructions and / or code.
[0249] Various forms of computer-readable media may be involved in carrying one or more sequences of one or more instructions to processor(s) 610 for execution. By way of example only, the instructions may initially be carried on a magnetic and / or optical disk of a remote computer. The remote computer may load the instructions into its dynamic memory and transmit the instructions as signals over a transmission medium to be received and / or executed by computer or hardware system 600. These signals, which may be in the form of electromagnetic, acoustic, optical, etc., are all examples of carrier waves that may encode instructions, according to various embodiments of the present invention.
[0250] The communications subsystem 630 (and / or its components) typically receives signals, and the bus 605 may then carry the signals (and / or the data, instructions, etc. carried by the signals) to the working memory 635, from which the processor(s) 605 retrieve and execute the instructions. The instructions received by the working memory 635 may optionally be stored on a storage device 625 either before or after execution by the processor(s) 610.
[0251] As described above, one set of embodiments includes methods and systems for performing annotation data collection, and more particularly, methods, systems, and apparatuses for performing annotation data collection using gaze-based tracking and / or training an artificial intelligence (“AI”) system (which may include, without limitation, at least one of a neural network, a convolutional neural network (“CNN”), a learning algorithm-based system, a machine learning system, or the like) based on annotation data collected using gaze-based tracking. FIG. 7 shows a schematic diagram of a system 700 that can be used in accordance with one set of embodiments. System 700 can include one or more user computers, user devices, or customer devices 705. User computers, user devices, or customer devices 705 can be general-purpose personal computers (including, by way of example only, desktop computers, tablet computers, laptop computers, handheld computers, etc. running any suitable operating system, some of which are available from vendors such as Apple Inc., Microsoft Corp.), cloud computing devices, server(s), and / or workstation computer(s) running any of a variety of commercially available UNIX or UNIX-like operating systems. The user computer, user device, or customer device 705 may also have one or more applications configured to perform the methods provided by various embodiments (e.g., as described above), as well as any of a variety of applications, including one or more office applications, database client and / or server applications, and / or web browser applications.Alternatively, user computers, user devices, or customer devices 705 may be any other electronic devices, such as thin client computers, Internet-enabled mobile phones, and / or personal digital assistants, capable of communicating over a network (e.g., network(s) 710, described below) and / or displaying and navigating web pages or other types of electronic documents. Although exemplary system 700 is shown with two user computers, user devices, or customer devices 705, any number of user computers, user devices, or customer devices may be supported.
[0252] Certain embodiments operate in a network environment that may include network(s) 710. Network(s) 710 may be any type of network familiar to those skilled in the art that is capable of supporting data communications using any of a variety of commercial (and / or free or proprietary) protocols, including, but not limited to, TCP / IP, SNA, IPX, AppleTalk, etc. By way of example only, network(s) 710 (similar to network(s) 150 of FIG. 1 , etc.) may each include, but are not limited to, a local area network ("LAN"), including a fiber network, an Ethernet network, a Token-Ring™ network, etc.; a wide area network ("WAN"); a wireless wide area network ("WWAN"); a virtual network such as a virtual private network ("VPN"); the Internet; an intranet; an extranet; a public switched telephone network ("PSTN"); an infrared network; a wireless network, including but not limited to a network operating under any of the IEEE 802.11 protocol suite, the Bluetooth protocol, and / or any other wireless protocol known in the art; and / or any combination of these and / or other networks. In particular embodiments, the network may include an access network of a service provider (e.g., an Internet service provider ("ISP")). In another embodiment, the network may include a core network of the service provider and / or the Internet.
[0253] An embodiment may also include one or more server computers 715. Each of the server computers 715 may be configured with an operating system, including, but not limited to, any of those mentioned above and any commercially available (or freely available) server operating systems. Each of the servers 715 may also run one or more applications, which may be configured to provide services to one or more clients 705 and / or other servers 715.
[0254] By way of example only, one of the servers 715 may be a data server, a web server, a cloud computing device(s), etc., as described above. The data server may, by way of example only, include (or be in communication with) a web server that may be used to process requests for web pages or other electronic documents from the user computers 705. The web server may also run a variety of server applications, including HTTP servers, FTP servers, CGI servers, database servers, Java servers, etc. In some embodiments of the present invention, the web server may be configured to serve web pages that may be run within web browsers on one or more of the user computers 705 to perform methods of the present invention.
[0255] The server computer 715, in some embodiments, may include one or more application servers, which may be configured with one or more applications accessible by clients running on one or more of the client computers 705 and / or other servers 715. By way of example only, the server(s) 715 may be one or more general-purpose computers capable of executing programs or scripts, including, but not limited to, web applications (which, in some cases, may be configured to perform methods provided by various embodiments), in response to the user computers 705 and / or other servers 715. By way of example only, web applications may be implemented as one or more scripts or programs written in any suitable programming language, such as Java, C, C#, or C++, and / or any scripting language, such as Perl, Python, or TCL, as well as any combination of programming and / or scripting languages. The application server(s) may also include a database server capable of processing requests from clients (including, depending on the configuration, dedicated database clients, API clients, web browsers, etc.) running on user computers, user devices, or customer devices 705 and / or another server 715, including, but not limited to, database servers commercially available from Oracle®, Microsoft®, Sybase®, IBM®, etc. In some embodiments, the application server may execute one or more of the processes for performing annotation data collection, and more particularly, may implement methods, systems, and apparatus for performing annotation data collection using gaze-based tracking and / or training AI systems based on annotation data collected using gaze-based tracking, as described in detail above.Data provided by the application server may be formatted as one or more web pages (e.g., including HTML, JavaScript, etc.) and / or may be forwarded to the user computer 705 via a web server (e.g., as described above). Similarly, the web server may receive web page requests and / or input data from the user computer 705 and / or forward the web page requests and / or input data to the application server. In some cases, the web server may be integrated with the application server.
[0256] According to further embodiments, one or more of the servers 715 may function as a file server and / or may contain one or more of the files (e.g., application code, data files, etc.) necessary to implement the various disclosed methods incorporated by applications running on the user computer 705 and / or another server 715. Alternatively, as will be appreciated by those skilled in the art, a file server may contain all necessary files that allow such applications to be launched remotely by the user computer, user device, or customer device 705 and / or server 715.
[0257] It should be noted that the functions described herein with respect to various servers (e.g., application server, database server, web server, file server, etc.) may be performed by a single server and / or multiple specialized servers, depending on implementation-specific needs and parameters.
[0258] In certain embodiments, the system may include one or more databases 720a-720n (collectively "databases 720"). The location of each of the databases 720 may be determined arbitrarily, and by way of example only, database 720a may reside on a storage medium local to (and / or resident on) server 715a (and / or user computer, user device, or customer device 705). Alternatively, database 720n may be located remotely from any or all of computers 705, 715, as long as it is capable of communicating with one or more of them (e.g., via network 710). In one particular set of embodiments, database 720 may reside on a storage-area network ("SAN") familiar to those skilled in the art. (Similarly, any files necessary to perform the functions attributed to computers 705, 715 may be stored locally and / or remotely on the respective computers, as appropriate.) In one set of embodiments, database 720 may be a relational database, such as an Oracle database, adapted to store, update, and retrieve data in response to SQL-formatted commands. This database may be controlled and / or maintained (i.e., controlled and / or maintained) by a database server, for example, as described above.
[0259] According to some embodiments, system 700 may further include a computing system 725 (such as computing systems 105a, 205, and 305 of FIGS. 1, 2A, and 3A) and corresponding database(s) 730 (such as database(s) 110a of FIG. 1). System 700 may further include a microscope 735 (such as microscopes 115 and 310 of FIGS. 1 and 3) and a display device 740 (such as display devices 120 and 210 of FIGS. 1 and 2) used to enable a user 745 to view an optical view of a first sample (e.g., as shown in FIGS. 2B and 3B), and a camera 750 capable of capturing an image of a user 745 (and, in some cases, capturing an image of at least one eye of the user 745) while the user 745 is within a field of view (“FOV”) 750a of the camera 750. In some cases, camera 750 may include, but is not limited to, one or more eye tracking sensors, one or more motion sensors, one or more tracking sensors, or the like. System 700 may further include one or more audio sensors 755 (optionally; similar to, e.g., audio sensor(s) 135 of FIG. 1 ; including, but not limited to, one or more microphones, one or more voice recorders, one or more audio recorders, or the like) and one or more user devices 760 (optionally; similar to, e.g., user device(s) 140 of FIG. 1 ; including, but not limited to, a smartphone, a mobile phone, a tablet computer, a laptop computer, a desktop computer, or a monitor, or the like). Alternatively or in addition to computing system 725 and corresponding database(s), system 700 may further include a remote computing system 770 (e.g., similar to remote computing system 105b of FIG. 1 ) and corresponding database(s) 775 (e.g., similar to database(s) 110b of FIG. 1 ).In some embodiments, the system 700 may further comprise an artificial intelligence (“AI”) system 780 .
[0260] During operation, microscope 735 can project an optical view of the first sample into eyepiece(s) through which at least one eye of user 745 is observing. Camera 750 (or an eye-tracking device) can capture at least one first image of at least one eye of user 745 as the user 745 views the optical view of the first sample. Computing system 725, user device 705a, user device 705b, user device(s) 760, server 715a or 715b, and / or remote computing system(s) 770 (collectively, “computing systems,” etc.) can analyze the at least one captured first image of at least one eye of user 745 and the at least one captured second image of the optical view of the first sample to determine whether at least one eye of user 745 is focused on a particular region of the optical view of the first sample. Based on a determination that at least one eye of the user 745 is focused on a particular region of the optical view of the first sample, the computing system can identify at least one particular portion of the at least one second image that corresponds to the particular region of the optical view of the first sample. The computing system can collect attention data including the identified at least one particular portion of the at least one second image and store the collected attention data in databases 720a-720n, 730, or 775. According to some embodiments, the collection of attention data can occur without interrupting, slowing, or disrupting the user while the user is providing result data either while diagnosing the first sample using microscope 735 or while diagnosing the image of the first sample displayed on display screen 740.In some cases, the collected attention data may include, but is not limited to, at least one of: one or more coordinate locations of at least one particular portion of the optical view of the first sample; an attention duration during which a user focuses on at least one particular portion of the optical view of the first sample; a zoom level of the optical view of the first sample while the user focuses on at least one particular portion of the optical view of the first sample; etc. In some cases, the identified at least one particular portion of the at least one second image corresponding to a particular region of the optical view of the first sample may include, but is not limited to, at least one of: one or more particular cells, one or more particular tissues, one or more particular structures, one or more molecules, etc.
[0261] In some embodiments, the computing system may generate at least one highlight field in the at least one second image that covers at least one identified portion of the at least one second image that corresponds to a particular region of the optical view of the first sample. In some cases, each of the at least one highlight field may include, without limitation, at least one of a color, a shape, or a highlight effect, etc., where the highlight effect may include, but is not limited to, at least one of an outlining effect, a shadowing effect, a patterning effect, a heat map effect, or a jet color map effect, etc.
[0262] According to some embodiments, the at least one second image can be displayed on a display screen (e.g., a display screen of the display device 740, etc.). Capturing at least one first image of at least one eye of the user 745 can include using a camera 750 to capture at least one first image of at least one eye of the user 745 while the user 745 views image(s) or video(s) of the optical view of the first sample displayed as the at least one second image on the display screen of the display device 740. An eye-tracking device can be used instead of the camera 750 to collect attention data while the user views the image(s) or video of the first sample displayed on the display screen of the display device 740. Identifying at least one specific portion of the at least one second image corresponding to a specific region of the optical view of the first sample can include using a computing system to identify at least one specific portion of the at least one second image displayed on the display screen that corresponds to a specific region of the optical view of the first sample. The computing system may display the at least one second image on a display screen (e.g., a display screen of the display device 740, etc.) with the generated at least one highlight field covering the identified at least one particular portion of the at least one second image corresponding to the particular region of the optical view of the first sample.
[0263] In some embodiments, the display of the at least one second image on the display screen can shift in response to a command by a user. In some cases, the shifting display of the at least one second image can include at least one of a horizontal shift, a vertical shift, panning, tilting, zooming in, or zooming out, etc., of the at least one second image on the display screen. The camera 750 can track the movement of at least one eye of the user 745 while the user 745 is viewing the shifting display of the at least one second image on the display screen. The computing system can match the tracked movement of the at least one eye of the user 745 with the shifting display of the at least one second image on the display screen based at least in part on one or more of the tracked movement of the at least one eye of the user 745, an identified at least one specific portion of the at least one second image corresponding to a specific region of the optical view of the first sample, or at least one of a horizontal shift, a vertical shift, panning, tilting, zooming in, or zooming out, etc., of the at least one second image on the display screen. Instead of using the camera 750, an eye-tracking device can be used to collect additional attention data when the user is viewing the shifting display of at least one second image on the display screen of the display device 740.
[0264] Alternatively, the microscope 735 can project the optical view of the first sample into an eyepiece through which at least one eye of the user 745 is observing. The second camera can capture at least one second image of the optical view of the first sample. In some cases, capturing the at least one first image of the at least one eye of the user 745 can include capturing the at least one first image of the at least one eye of the user 745 using the first camera while the user 745 is viewing the optical view of the first sample through the eyepiece. Identifying at least one specific portion of the at least one second image corresponding to a specific region of the optical view of the first sample can include identifying, using a computing system, at least one specific portion of the at least one second image being observed through the eyepiece that corresponds to a specific region of the optical view of the first sample. Generating at least one highlight field in the at least one second image overlying at least one identified portion of the at least one second image corresponding to a particular region of the optical view of the first sample can include using a computing system to generate at least one highlight field that overlaps with at least one identified portion of the at least one second image being viewed through the eyepiece that corresponds to a particular region of the optical view of the first sample. The computing system can use a projection device to project the generated at least one highlight field to overlap with at least one identified portion of the at least one second image being viewed through the eyepiece that corresponds to a particular region of the optical view of the first sample. Alternatively or additionally, the computing system can display the at least one second image on a display screen (e.g., a display screen of display device 740) with the generated at least one highlight field overlying at least one identified portion of the at least one second image that corresponds to a particular region of the optical view of the first sample.
[0265] In some cases, the first camera can be one of an infrared ("IR") camera, a back-reflecting IR camera, a visible color camera, a light source, a location photodiode, etc. In some cases, the microscope can include two or more of a plurality of mirrors, a plurality of dichroic mirrors, or a plurality of half mirrors that reflect or pass at least one of, but are not limited to, an optical view of the first sample observed through the eyepiece, an optical view of at least one eye of a user observed through the eyepiece and captured as at least one first image by the first camera, or a projection of the generated at least one highlighted field through the eyepiece onto at least one eye of a user, etc.
[0266] According to some embodiments, the projection of the optical view of the first sample onto the eyepiece can be shifted by at least one of adjusting an XY stage on which a microscope slide containing the first sample is mounted, changing an objective lens or a zoom lens, adjusting the focus of the eyepiece, etc. The camera 750 can track movement of at least one eye of the user 745 as the user 745 views the shifted projection of the optical view of the first sample onto the eyepiece. The computing system can match the tracked movement of the at least one eye of the user 745 with the shifted projection of the optical view of the first sample onto the eyepiece based at least in part on one or more of the tracked movement of the at least one eye of the user 745, identifying at least one specific portion of the at least one second image that corresponds to a specific region of the optical view of the first sample, or at least one of adjusting an XY stage on which a microscope slide containing the first sample is mounted, changing an objective lens or a zoom lens, adjusting the focus of the eyepiece, etc.
[0267] Alternatively or additionally, the one or more audio sensors 755 may capture one or more voice notes from the user 745 as the user 745 views the optical view of the first sample. The computing system may map the one or more voice notes captured from the user 745 with the at least one second image of the optical view of the first sample to match the one or more captured voice notes with the at least one second image of the optical view of the first sample.
[0268] According to some embodiments, the computing system can receive outcome data provided by a user. The outcome data can include at least one of a diagnosis of the first sample, a pathology score of the first sample, or a set of identification data corresponding to at least a plurality of portions of the first sample. The computing system can train an AI system 780 (which can generally include, but is not limited to, at least one of a neural network, a convolutional neural network (“CNN”), a learning algorithm-based system, a machine learning system, etc.) based at least in part on at least one of an analysis of at least one captured first image of at least one eye of the user and at least one captured second image of an optical view of the first sample, or a joint analysis of the collected attention data and the received outcome data, to generate a model used to generate a predicted value. In some embodiments, the predicted value can include, but is not limited to, at least one of a predicted clinical outcome, a predicted attention data, etc.
[0269] These and other features of system 700 (and its components) are described in more detail above with respect to FIGS.
[0270] Additional exemplary embodiments will now be described.
[0271] According to an aspect of some embodiments of the present invention, projecting, using a microscope, an optical view of the first sample into an eyepiece through which at least one eye of a user is viewing; capturing, with a first camera, at least one first image of the at least one eye of the user while the user is viewing the optical view of the first sample through the eyepiece; capturing at least one second image of the optical view of the first sample with a second camera; using a computing system to analyze the captured at least one first image of the at least one eye of the user and the captured at least one second image of the optical view of the first sample to determine whether the at least one eye of the user is focused on a particular region of the optical view of the first sample; based on determining that the at least one eye of the user is focused on a particular region of the optical view of the first sample, identifying, with the computing system, at least one particular portion of the at least one second image viewed through the eyepiece that corresponds to the particular region of the optical view of the first sample; collecting attention data including the identified at least one particular portion of the at least one second image using the computing system; and storing the collected attention data in a database; A method is provided, comprising:
[0272] Optionally, the first sample is contained in at least one of a microscope slide, a transparent sample cartridge, a vial, a tube, a capsule, a flask, a vessel, a receptacle, a microarray, or a microfluidic chip.
[0273] Optionally, the first camera is one of an infrared ("IR") camera, a back-reflective IR camera, a visible color camera, a light source, or a location photodiode.
[0274] Optionally, the microscope comprises two or more of a plurality of mirrors, a plurality of dichroic mirrors, or a plurality of half mirrors that reflect or pass at least one of the optical view of the first sample observed through the eyepiece or the optical view of the at least one eye of the user observed through the eyepiece and captured as the at least one first image by the first camera.
[0275] Optionally, the identified at least one specific portion of the at least one second image corresponding to the specific region of the optical view of the first sample includes at least one of one or more specific cells, one or more specific tissues, one or more specific structures, or one or more molecules.
[0276] Optionally, identifying the at least one particular portion of the at least one second image includes determining, using the computing system, a coordinate location within the at least one second image of the optical view that corresponds to the identified at least one particular portion of the at least one second image.
[0277] Optionally, receiving, with the computing system, result data provided by the user, the result data including at least one of a diagnosis of the first sample, a pathology score of the first sample, or a set of identification data corresponding to at least a plurality of portions of the first sample; training at least one of a neural network, a convolutional neural network ("CNN"), an artificial intelligence ("AI") system, or a machine learning system based at least in part on at least one of an analysis of the captured at least one first image of the at least one eye of the user and the captured at least one second image of the optical view of the first sample, or a joint analysis of the collected attention data and the received outcome data, to generate a model used to generate a prediction; Further includes:
[0278] Optionally, the predictive value comprises at least one of predicted clinical outcome or predicted attention data.
[0279] Optionally, collecting the attention data is done without interrupting, slowing down or disturbing the user while the user is providing the result data while diagnosing the first sample using the microscope.
[0280] Optionally, tracking movement of the at least one eye of the user with the first camera; using the computing system to simultaneously track at least one of one or more coordinate locations of the identified at least one particular portion of the at least one second image, an attention duration during which the user focuses on the particular region of the optical view, or a zoom level of the optical view of the first sample while the user focuses on the particular region of the optical view; Further includes:
[0281] Optionally, determining whether the at least one eye of the user is focusing on a particular region of the optical view of the first sample includes determining whether the at least one eye of the user is focusing on a particular region of the optical view of the first sample based at least in part on at least one of the one or more coordinate locations of the identified at least one particular portion of the at least one second image, the attention duration during which the user is focusing on the particular region of the optical view, or the zoom level of the optical view of the first sample while the user is focusing on the particular region of the optical view.
[0282] Optionally, capturing one or more voice notes from the user using an audio sensor while the user is viewing the optical view of the first sample; using the computing system to map the one or more voice notes captured from the user with the at least one second image of the optical view of the first sample to match the one or more captured voice notes with the at least one second image of the optical view of the first sample; Further includes:
[0283] According to an aspect of some embodiments of the present invention, a microscope configured to project an optical view of the first sample into an eyepiece through which at least one eye of a user is viewing; a first camera configured to capture at least one first image of the at least one eye of the user when the user views the optical view of the first sample through the eyepiece; a second camera configured to capture at least one second image of the optical view of the first sample; a computing system comprising at least one first processor and a first non-transitory computer-readable medium communicatively coupled to the at least one first processor; Equipped with The first non-transitory computer-readable medium has stored thereon computer software including a first set of instructions that, when executed by the at least one first processor, cause a computing system to: analyze the captured at least one first image of the at least one eye of the user and the captured at least one second image of the optical view of the first sample to determine whether the at least one eye of the user is focused on a particular region of the optical view of the first sample; based on a determination that the at least one eye of the user is focused on a particular region of the optical view of the first sample, identify at least one particular portion of the at least one second image being viewed through the eyepiece that corresponds to the particular region of the optical view of the first sample; collect attention data including the identified at least one particular portion of the at least one second image; and store the collected attention data in a database. A system is provided.
[0284] Optionally, said first set of instructions when executed by said at least one first processor: receiving result data provided by the user, the result data including at least one of a diagnosis of the first sample, a pathology score of the first sample, or a set of identification data corresponding to at least a plurality of portions of the first sample; training at least one of a neural network, a convolutional neural network ("CNN"), an artificial intelligence ("AI") system, or a machine learning system based at least in part on at least one of an analysis of the captured at least one first image of the at least one eye of the user and the captured at least one second image of the optical view of the first sample, or a joint analysis of the collected attention data and the received outcome data, to generate a model used to generate a prediction; The computing system further performs the steps of:
[0285] Optionally, the predictive value comprises at least one of predicted clinical outcome or predicted attention data.
[0286] Optionally, the first camera is further configured to track movement of the at least one eye of the user; the computing system is further configured to simultaneously track at least one of one or more coordinate locations, attention duration, or zoom level of the optical view of the first sample; Determining whether the at least one eye of the user is focused on a particular region of the optical view of the first sample includes determining whether the at least one eye of the user is focused on a particular region of the optical view of the first sample based at least in part on one or more of tracking the one or more coordinate locations of an attention gaze, tracking at least one of the movement and zoom level of the optical view of the first sample, or determining that the at least one eye of the user is continuing to look at a portion of the optical view of the first sample.
[0287] Optionally, an audio sensor configured to capture one or more voice notes from the user while the user is viewing the optical view of the first sample; The first set of instructions, when executed by the at least one first processor, further causes the computing system to map the captured one or more voice notes from the user with the at least one second image of the optical view of the first sample to match the captured one or more voice notes with the at least one second image of the optical view of the first sample.
[0288] According to an aspect of some embodiments of the present invention, receiving at least one first image of at least one eye of a user captured by a first camera while the user is viewing an optical view of a first sample through an eyepiece of a microscope; receiving at least one second image of the optical view of the first sample captured by a second camera; using a computing system to analyze the at least one first image and the at least one second image to determine whether the at least one eye of the user is focused on a particular region of the optical view of the first sample; using the computing system to track the user's attention based on the analysis; and using said computing system to collect attention data based on said tracking; A method is provided which includes:
[0289] According to an aspect of some embodiments of the present invention, receiving, with a computing system, collected attention data corresponding to a user viewing an optical view of the first sample; receiving, with the computing system, result data provided by the user, the result data including at least one of a diagnosis of the first sample, a pathology score of the first sample, or a set of identification data corresponding to at least a plurality of portions of the first sample; training at least one of a neural network, a convolutional neural network ("CNN"), an artificial intelligence ("AI") system, or a machine learning system based at least in part on the joint analysis of the collected attention data and the received outcome data to generate a model used to generate a prediction; A method is provided which includes:
[0290] Optionally, the first sample is contained in at least one of a microscope slide, a transparent sample cartridge, a vial, a tube, a capsule, a flask, a vessel, a receptacle, a microarray, or a microfluidic chip, or the like.
[0291] Optionally, the predictive value comprises at least one of predicted clinical outcome or predicted attention data.
[0292] Optionally, collecting the attention data is done without interrupting, slowing down or obstructing the user while the user is providing the result data either while diagnosing the first sample using a microscope or while diagnosing an image of the first sample displayed on a display screen.
[0293] Optionally, the collected attention data includes at least one of one or more coordinate locations of at least one particular portion of the optical view of the first sample, an attention duration during which the user focuses on the at least one particular portion of the optical view of the first sample, or a zoom level of the optical view of the first sample while the user focuses on the at least one particular portion of the optical view of the first sample.
[0294] Optionally, the attention data is collected based on at least one first image of the at least one eye of the user captured by a first camera while the user is viewing an optical view of the first sample through an eyepiece of a microscope.
[0295] Optionally, the microscope comprises two or more of a plurality of mirrors, a plurality of dichroic mirrors, or a plurality of half mirrors that reflect or pass at least one of the optical view of the first sample observed through the eyepiece or the optical view of the at least one eye of the user observed through the eyepiece and captured as the at least one first image by the first camera.
[0296] Optionally, the attention data is collected using an eye tracking device when the user is looking at a first image of the optical view of the first sample displayed on a display screen.
[0297] Optionally, generating, with the computing system, at least one highlight field overlapping with the identified at least one particular portion of the at least one first image displayed on the display screen that corresponds to a particular region of the optical view of the first sample. Further includes:
[0298] Optionally, using the computing system to display the generated at least one highlight field on the display screen so as to overlap the identified at least one particular portion of the at least one first image displayed on the display screen corresponding to the collected attention data; using the eye-tracking device to track the attention data while the user is viewing the first image of the optical view of the first sample displayed on the display screen; using the computing system, matching the tracked attention data with the representation of the at least one first image of the optical view of the first sample displayed on the display screen based at least in part on at least one of: one or more coordinate locations of at least one particular portion of the optical view of the first sample; an attention duration during which the user focuses on the at least one particular portion of the optical view of the first sample; or a zoom level of the optical view of the first sample while the user focuses on the at least one particular portion of the optical view of the first sample; Further includes:
[0299] Optionally, each of the at least one highlight field comprises at least one of a color, a shape, or a highlight effect, wherein the highlight effect comprises at least one of an outlining effect, a shadowing effect, a patterning effect, a heat map effect, or a jet color map effect.
[0300] Optionally, Tracking attention data using an eye-tracking device; using the computing system to simultaneously track at least one of: one or more coordinate locations of the identified at least one particular portion of at least one second image of the optical view of the first sample; an attention duration during which the user focuses on a particular region of the optical view; or a zoom level of the optical view of the first sample while the user focuses on the particular region of the optical view; Further includes:
[0301] Optionally, capturing one or more voice notes from the user using an audio sensor while the user is viewing the optical view of the first sample; using the computing system to map the one or more voice notes captured from the user with the at least one third image of the optical view of the first sample to match the one or more captured voice notes with the at least one third image of the optical view of the first sample; Further includes:
[0302] According to an aspect of some embodiments of the present invention, 1. An apparatus comprising: at least one processor; a non-transitory computer-readable medium communicatively coupled to the at least one processor; Equipped with The non-transitory computer-readable medium has stored thereon computer software including a set of instructions that, when executed by the at least one processor, receiving collected attention data corresponding to a user viewing an optical view of the first sample; receiving result data provided by the user, the result data including at least one of a diagnosis of the first sample, a pathology score of the first sample, or a set of identification data corresponding to at least a plurality of portions of the first sample; training at least one of a neural network, a convolutional neural network ("CNN"), an artificial intelligence ("AI") system, or a machine learning system based at least in part on the joint analysis of the collected attention data and the received outcome data to generate a model used to generate a prediction; An apparatus is provided that causes the apparatus to perform the steps of:
[0303] According to an aspect of some embodiments of the present invention, a first camera configured to capture at least one first image of the at least one eye of the user while the user is viewing the optical view of the first sample; a second camera configured to capture at least one second image of the optical view of the first sample; a computing system comprising at least one first processor and a first non-transitory computer-readable medium communicatively coupled to the at least one first processor; Equipped with The first non-transitory computer-readable medium has stored thereon computer software including a first set of instructions that, when executed by the at least one first processor, receiving collected attention data corresponding to a user viewing the optical view of the first sample; receiving result data provided by the user, the result data including at least one of a diagnosis of the first sample, a pathology score of the first sample, or a set of identification data corresponding to at least a plurality of portions of the first sample; training at least one of a neural network, a convolutional neural network ("CNN"), an artificial intelligence ("AI") system, or a machine learning system based at least in part on the joint analysis of the collected attention data and the received outcome data to generate a model used to generate a prediction; A system is provided that causes the computing system to perform the following:
[0304] While certain features and aspects have been described with reference to illustrative embodiments, those skilled in the art will recognize that numerous variations are possible. For example, the methods and processes described herein may be implemented using hardware components, software components, and / or any combination thereof (i.e., hardware components, software components, and / or any combination thereof). Furthermore, while various methods and processes described herein may be described with reference to particular structural and / or functional components for ease of explanation, methods provided by various embodiments are not limited to any structural and / or functional architecture (i.e., structural architecture and / or functional architecture), but rather may be implemented in any suitable hardware, firmware, and / or software configuration. Similarly, while certain functionality may be attributed to certain system components, unless the context dictates otherwise, this functionality may be distributed among various other system components according to some embodiments.
[0305] Moreover, although the steps of the methods and processes described herein are described in a particular order for ease of description, various steps can be rearranged, added, and / or omitted according to various embodiments, unless the context dictates otherwise. Moreover, steps described with respect to one method or process can be incorporated into other described methods or processes, and similarly, system components described according to a particular structural architecture and / or with respect to one system can be organized into alternative structural architectures and / or incorporated into other described systems. Thus, although various embodiments are described with or without certain features for ease of description and to illustrate exemplary aspects of those embodiments, various components and / or features described herein with respect to particular embodiments can be substituted, added, and / or subtracted from other described embodiments, unless the context dictates otherwise. Consequently, while several exemplary embodiments have been described above, it will be understood that the invention is intended to encompass all modifications and equivalents included within the scope of the following claims.
[0306] The description of various embodiments of the present invention has been presented for illustrative purposes and is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art that do not depart from the scope and spirit of the described embodiments. The terminology used herein has been chosen to best explain the principles, practical applications, or technical improvements over the art found in the marketplace of the embodiments, or to enable others skilled in the art to understand the embodiments disclosed herein.
[0307] It is expected that many relevant machine learning models will be developed during the life of the patent from the filing date of this application until expiration, and the scope of the term machine learning model will a priori include all such new technologies.
[0308] As used herein, the term "about" refers to ±10%.
[0309] The terms "comprising," "including," "having," and their conjugated variations mean "including, but not limited to." This term encompasses the terms "consisting of" and "consisting essentially of."
[0310] The phrase "consisting essentially of" means that a composition or method may include additional components and / or steps (i.e., components or steps, or both), but only if the additional components and / or steps do not materially alter the basic and novel characteristics of the claimed composition or method.
[0311] As used herein, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. For example, the phrase "a compound" or "at least one compound" can include plural compounds, including mixtures of plural compounds.
[0312] The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment described as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments and / or as excluding the incorporation of features from other embodiments.
[0313] The word "optionally" is used herein to mean "provided in some embodiments and not provided in others." Any particular embodiment of the present invention may include multiple "optional" features unless such features contradict each other.
[0314] Throughout this application, various embodiments of the invention may be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the invention. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges and individual numerical values within that range. For example, the description of a range such as 1 to 6 should be considered to have specifically disclosed subranges such as 1 to 3, 1 to 4, 1 to 5, 2 to 4, 2 to 6, 3 to 6, etc., as well as individual numerical values within that range, e.g., 1, 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range.
[0315] Whenever a range of numerical values is given herein, it is intended to include every recited number (fractional or integer) that falls within the given range. The phrases "a range between a first indicated number and a second indicated number" and "a range from a first indicated number to a second indicated number" are used interchangeably herein and are intended to include the first indicated number and the second indicated number and all fractional and integer numbers therebetween.
[0316] It will be appreciated that certain features of the invention, which are for clarity described in the context of separate embodiments, can also be provided in combination in a single embodiment. Conversely, various features of the invention, which are for brevity described in the context of a single embodiment, can also be provided separately or in any suitable subcombination, or as appropriate with any other described embodiment of the invention. Certain features described in the context of various embodiments should not be considered essential features of those embodiments, unless the embodiment is inoperable without those elements.
[0317] While the present invention has been described in conjunction with specific embodiments thereof, many alternatives, modifications, and variations will be apparent to those skilled in the art. Accordingly, it is intended to embrace all such alternatives, modifications, and variations that fall within the spirit and broad scope of the appended claims.
[0318] It is the intention of the applicant(s) that all publications, patents, and patent applications referenced herein, when incorporated by reference, be incorporated by reference in their entirety as if each individual publication, patent, and patent application was specifically and individually referenced. In addition, citation or identification of any reference in this application should not be construed as an admission that such reference is available as prior art to the present invention. To the extent section headings are used, they should not be construed as necessarily limiting. Additionally, any priority documents of this application are also incorporated by reference in their entirety. The claims as originally filed are as follows: Claim 1: 1. A computer-implemented method for automatically generating a training data set comprising a plurality of records, the method comprising: Here, one record is an image of a sample of the object; a display of monitored user actions on the sample; a ground truth representation of a monitored line of sight of the user observing the sample on a display or through an optical device, mapped to pixels of the image of the sample, wherein the monitored line of sight includes at least one location of the sample observed by the user and a time spent observing the at least one location; A computer-implemented method comprising: Claim 2: 10. The computer-implemented method of claim 1, wherein the sample of the object is selected from the group consisting of a biological sample, a live cell culture in a microwell plate, a slide of a pathology tissue sample for generating a pathology report, a 3D radiology image, and a manufactured microarray for identifying manufacturing defects. Claim 3: 3. The computer-implemented method of claim 1 or 2, further comprising training a machine learning model on the training dataset to generate predicted line-of-sight results for a target in response to input of a target image of a target sample of a target object. Claim 4: 4. The computer-implemented method of claim 1, wherein the ground truth representation of the monitored line of sight includes the total time that the monitored line of sight is mapped to each particular pixel of the image over an observation time interval. Claim 5: 5. The computer-implemented method of claim 4, wherein the ground truth representation of the monitored gaze includes at least one of: (i) a heat map corresponding to the image of the sample, wherein the intensity of each pixel of the heat map correlates with the total time that the monitored gaze is mapped to each pixel, and the pixels of the heat map are normalized to pixels that indicate different actual sizes of the sample at multiple zoom levels defined by the monitored operation and / or pixels located in different portions of the sample that are non-simultaneously visible on a display obtained by a panning operation of the monitored operation; and (ii) an overlay on the image of the sample, wherein features of the overlay correspond to the extent of the gaze and / or indicate the total time. Claim 6: 6. The computer-implemented method of claim 1, wherein the ground truth representation of the monitored line of sight includes an ordered time sequence that dynamically maps adaptations of the monitored line of sight of different fields of view being observed to different specific pixels over an observation time interval. Claim 7: 7. The computer-implemented method of claim 6, wherein the ground truth representation of the monitored gaze is shown as at least one of: (i) a directed line overlaid on pixels of the image of the sample showing dynamic adaptation of the monitored gaze; and (ii) presenting the ordered time sequence along with an indication of the time spent in each field of view. Claim 8: 8. The computer-implemented method of claim 1, wherein the recording of the training dataset further includes a ground truth representation of the user's actions to adjust the field of view of the sample, which is mapped to a ground truth representation of the monitored line of sight and the pixels of the image. Claim 9: 9. The computer-implemented method of claim 1, wherein the sample is observed as a magnified image thereof, and the user operations associated with the mapping of the monitored line of sight to specific pixels of the image are selected from the group including zooming in, zooming out, panning left, panning right, panning up, panning down, adjusting the light, adjusting the focus, and adjusting the zoom of the image. Claim 10: The sample is observed via a microscope; monitoring the gaze includes acquiring gaze data from at least one first camera that tracks the pupil of the user while the user is observing the sample under the microscope; the image of the sample being manipulated is captured by a second camera while the user is observing the sample under the microscope; The computer-implemented method comprises: obtaining a scanned image of the sample; aligning the scanned image of the sample with the image of the sample captured by the second camera; Further comprising: 10. The computer-implemented method of claim 1, wherein mapping comprises mapping the monitored line of sight to pixels of the scanned image using the alignment to the image captured by the second camera. Claim 11: The monitored gaze is represented as a weak annotation; The records of the training dataset include the following additional ground truth labels of the images of the samples: When the sample comprises a sample of tissue from a subject, a pathology report generated by the user observing the sample, a pathology diagnosis generated by the user observing the sample, a sample score indicating a pathological evaluation of the sample generated by the user observing the sample, at least one clinical parameter of the subject for which the sample is represented, historical parameters of the subject, and the results of a treatment administered to the subject; When the sample comprises a manufactured microarray, a user-provided indication of at least one manufacturing defect and a pass / fail indication of quality assurance testing; When the sample comprises a live cell culture, cell growth rate, cell density, cell homogeneity, and cell heterogeneity; data items provided by one or more other users; The computer-implemented method of claim 1 , further comprising at least one of: Claim 12: When the sample comprises a tissue sample of the subject, a target predicted pathology report and / or pathology diagnosis and / or sample score in response to input of a target image of a target biological sample of pathological tissue of the target individual and a target gaze of a target user; and when the sample includes the manufactured microarray, a pass / fail indication of target manufacturing defects and / or quality checks in response to input of a target image of the target manufactured microarray. When the sample comprises a live cell culture, the target cell growth rate, target cell density, target cell homogeneity, and target cell heterogeneity can be measured. 12. The computer-implemented method of claim 11, further comprising training a machine learning model on the training dataset to generate a result of Claim 13: 1. A computer-implemented method for assisting in the visual analysis of a sample of an object, comprising: providing a target image of the sample of the object to a machine learning model trained on a training dataset comprising a plurality of records; Here, one record is an image of a sample of the object; a display of monitored user actions on the sample; a ground truth representation of a monitored line of sight of the user observing the sample on a display or through an optical device, mapped to pixels of the image of the sample, wherein the monitored line of sight includes at least one location of the sample that the user is observing and a time spent observing the at least one location; and and obtaining, as a result of the machine learning model, an indication of a predicted monitored line of sight of pixels of the target image; 10. A computer-implemented method comprising: Claim 14: 14. The computer-implemented method of claim 13, wherein the results include a heat map of multiple pixels mapped to pixels of the target image, the intensity of the pixels of the heat map being correlated to the predicted time of gaze, and the pixels of the heat map being normalized to pixels that indicate different actual sizes of the sample at multiple zoom levels defined by the monitored operation and / or pixels located in different portions of the sample that are non-simultaneously visible on a display obtained by a panning operation of the monitored operation. Claim 15: 15. The computer-implemented method of claim 13 or 14, wherein the results include a time series showing dynamic gaze mapped to pixels of the target image over a time interval, and the computer-implemented method further includes monitoring in real time the gaze of a user observing the target image, comparing a difference between the real-time monitoring and the time series, and generating an alert when the difference exceeds a threshold. Claim 16: 16. The computer-implemented method of claim 13, wherein the records of the training dataset further include a ground truth representation of the user's actions that are mapped to a ground truth representation of the monitored gaze and the pixels of the image, and the results include a prediction of an action to be taken with respect to a presentation of the target image. Claim 17: 16. The computer-implemented method of claim 15, further comprising: monitoring a user's manipulation of the sample presentation in real time; comparing a difference between the real-time monitoring of manipulation and a prediction of manipulation; and generating an alert when the difference exceeds a threshold. Claim 18: 1. A computer-implemented method for assisting in the visual analysis of a sample of an object, comprising: providing the sample target images to a machine learning model; obtaining a sample score as a result of the machine learning model, the sample score being indicative of a visual evaluation of the sample; Including, The machine learning model is trained on a training dataset including a plurality of recordings, wherein one recording includes an image of a sample of an object, a representation of a monitored user's manipulation of a presentation of the sample, and a ground truth representation of a monitored gaze of the user observing the sample on a display or through an optical device mapped to pixels of the image of the sample, wherein the monitored gaze includes at least one location of the sample observed by the user and a time spent observing the at least one location, and a ground truth representation of a sample visual evaluation score assigned to the sample. A computer-implemented method. Claim 19: An eye-tracking component integrated with the microscope between the objective lens and the eyepiece, comprising: an optical device that directs a first set of electromagnetic frequencies reflected back from each eye of a user viewing a sample under a microscope to a respective first camera that generates an indication of the user's tracked line of sight, and simultaneously directs a second set of electromagnetic frequencies from the sample under the microscope to a second camera that captures an image indicative of the field of view being viewed by the user; A component comprising: Claim 20: The first set of electromagnetic frequencies are IR frequencies generated by an infrared (IR) source, the first camera comprises a near-IR camera, the second set of electromagnetic frequencies comprises the visible light spectrum, the second camera comprises a red-green-blue (RGB) camera, and the optical device directs the first set of electromagnetic frequencies from the IR source to an eyepiece where the eye of the user is located, directs the back-reflected first set from the eye of the user through the eyepiece to the NIR camera, and directs the second set of electromagnetic frequencies from the sample to the near-IR camera. 20. The component of claim 19, wherein the optical device including a beam splitter directing the electromagnetic light waves from a single optical path after reflection from two eyes into two optical paths to two of the first cameras is selected from the group consisting of: using an infrared spectrum light source along with polarizers and / or wave plates directing different polarizations into different optical paths, and / or dichroic mirrors and spectral filters, and / or applying amplitude modulation at different frequencies to each optical path for heterodyne detection.
Claims
1. A computer-implemented method for automatically generating a training data set, comprising: (a) automatically generating the training data set using a processor, wherein at least one record of a plurality of records comprises: an image of a sample of the object; a display of monitored user actions on the sample; a ground truth representation of a monitored line of sight of the user observing the sample on a display or through an optical device, mapped to pixels of the image of the sample, wherein the monitored line of sight includes at least one location of the sample observed by the user and a time spent observing the at least one location; Including, The monitored gaze is represented as a weak annotation; The at least one record may include the following additional ground truth labels of the image of the sample: When the sample comprises a sample of tissue from a subject, a pathology report generated by the user observing the sample, a pathology diagnosis generated by the user observing the sample, a sample score indicating a pathological evaluation of the sample generated by the user observing the sample, at least one clinical parameter of the subject for which the sample is represented, historical parameters of the subject, and the results of a treatment administered to the subject; When the sample comprises a manufactured microarray, a user-provided indication of at least one manufacturing defect and a pass / fail indication of quality assurance testing; When the sample comprises a live cell culture, the growth rate, cell density, homogeneity parameters, and / or heterogeneity parameters of the live cell culture may be measured. The computer-implemented method further comprising at least one of:
2. 10. The computer-implemented method of claim 1, wherein the sample of the object is selected from the group consisting of a biological sample, a live cell culture in a microwell plate, a slide of a pathology tissue sample for generating a pathology report, a 3D radiology image, and a manufactured microarray for identifying manufacturing defects.
3. A computer-implemented method as described in claim 1, further comprising (b) training a machine learning model on the training dataset to generate predicted line-of-sight results for a target in response to input of a target image of a target sample of a target object.
4. The computer-implemented method of claim 1 , wherein the ground truth representation of the monitored line of sight includes a total time that the monitored line of sight is mapped to each particular pixel of the image over an observation time interval.
5. A computer-implemented method for automatically creating a training data set, comprising: (a) automatically generating the training data set using a processor, wherein at least one record of a plurality of records comprises: an image of a sample of the object; a display of monitored user actions on the sample; a ground truth representation of a monitored line of sight of the user observing the sample on a display or through an optical device, mapped to pixels of the image of the sample, wherein the monitored line of sight includes at least one location of the sample observed by the user and a time spent observing the at least one location; Including, the ground truth representation of the monitored line of sight includes a total time that the monitored line of sight is mapped to each particular pixel of the image over an observation time interval; A computer-implemented method in which the ground truth representation of the monitored gaze includes at least one of: (i) a heat map corresponding to the image of the sample, wherein the intensity of each pixel of the heat map is correlated with the total time the monitored gaze is mapped to each pixel, and the pixels of the heat map are normalized to pixels indicating different actual sizes of the sample at multiple zoom levels determined by the monitored operation and / or pixels located in different portions of the sample that are non-simultaneously visible on a display obtained by a panning operation of the monitored operation; and (ii) an overlay on the image of the sample, wherein features of the overlay correspond to the extent of the monitored gaze and / or indicate the total time.
6. The computer-implemented method of claim 1 , wherein the ground truth representation of the monitored gaze includes an ordered time sequence that dynamically maps the monitored gaze adaptations of different fields of view being observed to different specific pixels over an observation time interval.
7. 7. The computer-implemented method of claim 6, wherein the ground truth representation of the monitored gaze is shown as at least one of: (i) a directed line overlaid on pixels of the image of the sample showing dynamic adaptation of the monitored gaze; and (ii) presenting the ordered time sequence along with an indication of the time spent in each field of view.
8. 8. The computer-implemented method of claim 6 or 7, wherein the records of the training dataset further include a ground truth representation of the user's manipulations to adjust the field of view of the sample, which are mapped to a ground truth representation of the monitored line of sight and the pixels of the image.
9. 2. The computer-implemented method of claim 1, wherein the sample is observed as a magnified image thereof, and the monitored user operations associated with the mapping of the monitored line of sight to specific pixels of the image are selected from the group including zooming in, zooming out, panning left, panning right, panning up, panning down, adjusting the light, adjusting the focus, and adjusting the zoom of the image.
10. The sample is observed via a microscope; monitoring gaze includes acquiring gaze data from at least one first camera that tracks the pupils of the user while the user is observing the sample under the microscope; the image of the sample being manipulated is captured by a second camera while the user is observing the sample under the microscope; The computer-implemented method comprises: obtaining a scanned image of the sample; aligning the scanned image of the sample with the image of the sample captured by the second camera; Further comprising:
2. The computer-implemented method of claim 1, wherein mapping comprises mapping the monitored line of sight to pixels of the scanned image using the alignment to the image captured by the second camera.
11. When the sample comprises a tissue sample of the subject, a target predicted pathology report and / or pathology diagnosis and / or sample score in response to input of a target image of a target biological sample of pathological tissue of the target individual and a target gaze of a target user; and when the sample includes the fabricated microarray, a pass / fail indication of target manufacturing defects and / or quality checks in response to input of a target image of the target fabricated microarray. When the sample comprises a live cell culture, the target growth rate, target cell density, homogeneity parameter, and / or heterogeneity parameter of the live cell culture may be determined.
2. The computer-implemented method of claim 1, further comprising training a machine learning model on the training dataset to generate a result of
12. 1. A computer-implemented method for assisting in the visual analysis of a sample of an object, comprising: inputting a target image of the sample of the object into a machine learning model trained on a training dataset comprising a plurality of records; wherein at least one record of the plurality of records comprises: an image of a sample of the object; a display of monitored user actions on the sample; a ground truth representation of a monitored line of sight of the user observing the sample on a display or through an optical device, mapped to pixels of the image of the sample, wherein the monitored line of sight includes at least one location of the sample that the user was observing or is observing, and a time spent observing the at least one location; and obtaining as output of the machine learning model an indication of predicted monitored gaze of pixels of the target image and a heat map of pixels mapped to pixels of the target image, wherein the intensity of pixels of the heat map correlates with predicted times of gaze, and the pixels of the heat map are normalized to pixels representing different actual sizes of the sample at multiple zoom levels defined by the monitored operation and / or pixels located in different parts of the sample that are non-simultaneously visible on a display resulting from a panning operation of the monitored operation; 10. A computer-implemented method comprising:
13. 13. The computer-implemented method of claim 12, wherein the output includes a time series showing dynamic gaze mapped to pixels of the target image over a time interval, the computer-implemented method further including: monitoring in real time the gaze of a user observing the target image; comparing a difference between the real-time monitoring and the time series; and generating an alert when the difference exceeds a threshold.
14. 13. The computer-implemented method of claim 12, wherein the records of the training dataset further include a ground truth representation of the user's actions mapped to a ground truth representation of the monitored gaze and the pixels of the image, and the output includes a prediction of an action for a presentation of the target image.
15. 14. The computer-implemented method of claim 13, further comprising: monitoring a user's manipulation of the sample presentation in real time; comparing a difference between the real-time monitoring of manipulation and a prediction of manipulation; and generating an alert when the difference exceeds a threshold.
16. A computer-implemented method as described in claim 1, wherein the image of the sample includes a sample of tissue from a subject, and the at least one record includes a pathology report created by the user observing the sample, a pathology diagnosis created by the user observing the sample, a sample score indicating a pathology evaluation of the sample created by the user observing the sample, at least one clinical parameter of the subject for which the sample is indicated, historical parameters of the subject, and / or results of a treatment administered to the subject.
17. The computer-implemented method of claim 1, wherein the image of the subject includes a live cell culture, and the at least one record includes at least one label indicating growth rate, cell density, homogeneity parameter, and / or heterogeneity parameter of the live cell culture.
Citation Information
Patent Citations
Optical microscope system and method capable of tracking viewing positions in real time
CN110441901A
Image-recording device, image-recording method, and image-recording program
JP2007293818A
System and method to teach and evaluate image grading performance using prior learned expert knowledge base
US20180268737A1