System and method for real-time processing of medical images - Patents.com
Patent Information
- Application Number
- JP2023580547
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-07-04
- Filing Date
- 2022-07-04
- Publication Date
- 2025-07-10
AI Technical Summary
Current medical imaging systems, particularly endoscopy, suffer from high misdiagnosis rates due to manual and time-consuming visual inspection of video images, which are costly and difficult to standardize, leading to inefficiencies and inaccuracies in diagnosing conditions like colorectal cancer.
A system and method for real-time processing of medical images using machine learning models to analyze and annotate images during procedures, incorporating speech-to-text conversion for user input, allowing for immediate classification and reporting of objects of interest with confidence scoring and bounding boxes.
Enhances diagnostic accuracy by reducing human error, providing real-time annotation and reporting of medical images, thereby improving the efficiency and reliability of medical procedures.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 218,357, filed July 4, 2021, the entire contents of which are incorporated herein by reference.
[0002] Described herein generally are various embodiments of a system for processing medical images in real time, as well as methods and computer program products thereof. [Background technology]
[0003] The following paragraphs are provided as background to the present disclosure, but are not to be construed as an admission that anything discussed herein is part of the prior art or the knowledge of those skilled in the art.
[0004] Medical imaging provides the input necessary to confirm the diagnosis of disease, monitor a patient's response to treatment, and in some cases provide therapeutic treatment. Several different medical imaging modalities can be used for various medical diagnostic procedures. Examples of medical imaging modalities include gastrointestinal (GI) endoscopy, x-ray, MRI, CT scan, ultrasound, ultrasonography, echocardiography, cystography, laparoscopy, etc. Each of these requires analysis to ensure a proper diagnosis. The current state of the art can result in a rate of misdiagnosis that could be improved.
[0005] For example, endoscopy is the gold standard for confirming the diagnosis of gastrointestinal diseases, monitoring patient response to treatment, and in some cases providing curative treatment. Endoscopy videos collected from patients during clinical trials are typically reviewed by independent clinicians to reduce bias and increase accuracy. However, these analyses require visual review of video images and either manually recording the results or manually annotating the images, which is costly, time-consuming, and difficult to standardize.
[0006] Every year, millions of patients are misdiagnosed, with nearly half of them suffering from early-stage cancer. Colorectal cancer (CRC) is the third leading cause of cancer deaths globally, but if detected early, it can be successfully treated. Currently, clinicians manually report diagnoses after visually analyzing endoscopy / colonoscopy video images. The misdiagnosis rate in endoscopy is over 28%, which is mainly due to human error. Misdiagnosis is therefore a major problem for healthcare systems and patients, as well as having significant socio-economic impacts.
[0007] Traditional systems display the video generated by the endoscope during the endoscopy procedure, record the video (in rare cases), and offer no further functionality. In some cases, researchers save the images to their desktop and use offline programs to manually draw lines around polyps or other objects of interest. However, this analysis is done after the endoscopy has been performed, so clinicians cannot rescan areas of the colon after the procedure has concluded, even if there are inconclusive results.
[0008] What is needed is a system and method that addresses the above-mentioned problems and / or shortcomings. Summary of the Invention
[0009] In accordance with the teachings herein, various embodiments of systems and methods for processing medical images in real time, and computer products for use therewith, are provided.
[0010] In one broad aspect, in accordance with the teachings herein, in at least one embodiment, there is provided a system for analyzing medical image data for a medical procedure, the system comprising: a non-transitory computer readable medium having stored thereon program instructions for analyzing medical image data for a medical procedure; and, when the program instructions are executed, the system performs the following steps: receiving at least one image from a series of images; determining when at least one object of interest (OOI) is present in the at least one image; and, when the at least one OOI is present, determining a classification of the at least one OOI, both determining using at least one machine learning model; determining, during the medical procedure, at least and at least one processor configured to: display at least one image and any determined OOI to a user on a display; receive an input audio signal comprising speech from the user during the medical procedure and recognize the speech; when the speech is recognized as a comment on the at least one image during the medical procedure, convert the speech into at least one text string using a speech-to-text conversion algorithm; match the at least one text string with the at least one image to which the speech from the user was provided; and generate at least one annotated image in which the at least one text string is linked to the corresponding at least one image.
[0011] In at least one embodiment, the at least one processor is further configured to, when the utterance is recognized as a request for at least one reference image having an OOI classified with the same classification as the at least one OOI, display the at least one reference image and receive input from a user confirming or rejecting the classification of the at least one OOI.
[0012] In at least one embodiment, the at least one processor is further configured to receive an input from a user indicating a user classification for the at least one image having an undetermined OOI when the at least one OOI is classified as suspicious.
[0013] In at least one embodiment, the at least one processor is further configured to automatically generate a report that includes the at least one annotated image.
[0014] In at least one embodiment, the at least one processor is further configured to, for a given OOI in the given image, identify bounding box coordinates of a bounding box associated with the given OOI in the given image, calculate a confidence score based on a probability distribution of a classification of the given OOI, and when the confidence score is higher than a confidence threshold, overlay the bounding box on the at least one image with the bounding box coordinates.
[0015] In at least one embodiment, the at least one processor is configured to determine a classification of the OOI by applying a convolutional neural network (CNN) to the OOI by performing convolution, activation, and pooling operations to generate a matrix, processing the matrix using the convolution, activation, and pooling operations to generate a feature vector, and performing classification of the OOI based on the feature vector.
[0016] In at least one embodiment, the at least one processor is further configured to, when generating the at least one annotated image, overlay a timestamp on the corresponding at least one image.
[0017] In at least one embodiment, the at least one processor is further configured to indicate the confidence score of the at least one image in real time on a display or in a report.
[0018] In at least one embodiment, the at least one processor is configured to receive input audio during the medical procedure by starting to receive an audio stream of input audio from a user upon detection of a first user action including pausing the display of the series of images, taking a snapshot of a given image in the series of images, or providing a first voice command, and terminating reception of the audio stream upon detection of a second user action including remaining silent for a predetermined length of time, pressing a designated button, or providing a final voice command.
[0019] In at least one embodiment, the at least one processor is further configured to store a series of images upon receiving input audio during the medical procedure, thereby designating at least one image to receive annotation data for generating the corresponding at least one annotated image.
[0020] In at least one embodiment, the at least one processor is further configured to generate a report of the medical procedure by capturing a set of patient information data for addition to the report, loading a subset of the series of images comprising the at least one annotated image, and combining the set of patient information data and the subset of the series of images comprising the at least one annotated image into the report.
[0021] In at least one embodiment, the at least one processor is further configured to perform training of the at least one machine learning model by applying the encoder to the at least one training image to generate at least one feature vector for training OOI in the at least one training image, selecting a class of the training OOI by applying the at least one feature vector to the at least one machine learning model, and reconstructing, using a decoder, a labeled training image by associating the at least one feature vector with the at least one training image and the selected class for training the at least one machine learning model.
[0022] In at least one embodiment, the classes are a healthy tissue class, an unhealthy tissue class, a suspect tissue class, or an unfocused tissue class.
[0023] In at least one embodiment, the at least one processor is further configured to train the at least one machine learning model using a training dataset including labeled training images, unlabeled training images, or a mixture of labeled and unlabeled training images, where the images include examples categorized by healthy tissue, unhealthy tissue, suspicious tissue, and out-of-focus tissue.
[0024] In at least one embodiment, the at least one processor is further configured to train the at least one machine learning model using supervised, unsupervised, or semi-supervised learning.
[0025] In at least one embodiment, the training data set further includes subcategories for each of the unhealthy and suspicious tissues.
[0026] In at least one embodiment, the at least one processor is further configured to create at least one machine learning model by receiving training images as input to the encoder; projecting the training images using the encoder onto features that are part of the feature space; mapping the features to a set of target classes using a classifier; identifying morphological characteristics of the training images to generate a new training dataset, the new training dataset having data linking parameters to the training images; and determining based on the morphological characteristics whether there are one or more mapped classes or no mapped classes.
[0027] In at least one embodiment, the at least one processor is further configured to determine a classification of the at least one OOI by receiving one or more features as input to a decoder, mapping one of the features to an unlabeled dataset using a deconvolutional neural network, and reconstructing a new training image from one of the features using the decoder to train the at least one machine learning model.
[0028] In at least one embodiment, the at least one processor is further configured to train a speech-to-text algorithm using the speech dataset to compare new speech data with the speech dataset to identify matches to the ground truth text, where the speech dataset includes the ground truth text and speech data of the ground truth text.
[0029] In at least one embodiment, a speech-to-text algorithm maps at least one OOI to one of a plurality of OOI medical terms.
[0030] In at least one embodiment, the medical image data is obtained from one or more endoscopic procedures, one or more MRI scans, one or more CT scans, one or more X-rays, one or more ultrasound photographs, one or more nuclear medicine images, or one or more histological images.
[0031] In another broad aspect, in accordance with the teachings herein, in at least one embodiment, a system is provided for training at least one machine learning model and a speech-to-text conversion algorithm for use in analyzing medical image data for a medical procedure, the system including a non-transitory computer-readable medium having stored thereon program instructions for training the machine learning model, and at least one processor configured, upon execution of the program instructions, to: apply an encoder to at least one training image to generate at least one feature for a training object of interest (OOI) in the at least one training image; select a class of the training OOI by applying the at least one feature to the at least one machine learning model; reconstruct, using a decoder, a labeled training image by associating the at least one feature with the training image and the selected class for training the at least one machine learning model; train the speech-to-text conversion algorithm to identify matches between new speech data and the ground truth text using a speech dataset including ground truth text and speech data for the ground truth text, thereby generating at least one text string; and overlaying the training OOI and the at least one text string on an annotated image.
[0032] In at least one embodiment, the classes are a healthy tissue class, an unhealthy tissue class, a suspect tissue class, or an unfocused tissue class.
[0033] In at least one embodiment, the at least one processor is further configured to train the at least one machine learning model using a training dataset including labeled training images, unlabeled training images, or a mixture of labeled and unlabeled training images, where the images include examples categorized by healthy tissue, unhealthy tissue, suspicious tissue, and out-of-focus tissue.
[0034] In at least one embodiment, the at least one processor is further configured to train the at least one machine learning model using supervised, unsupervised, or semi-supervised learning.
[0035] In at least one embodiment, the training data set further includes subcategories for each of the unhealthy and suspicious tissues.
[0036] In at least one embodiment, the at least one processor is further configured to create the at least one machine learning model by receiving training images as input to the encoder; projecting the training images using the encoder into a feature space including features; mapping the features to a set of target classes using the classifier; and identifying morphological characteristics of the training images to generate a training dataset, the training dataset having data linking parameters to the training images; and determining based on the morphological characteristics whether there are one or more mapped classes or no mapped classes.
[0037] In at least one embodiment, the at least one processor is further configured to receive one or more features as input to a decoder, map one of the features to an unlabeled dataset using a deconvolutional neural network, and reconstruct a new training image from one of the features using the decoder for training the at least one machine learning model.
[0038] In at least one embodiment, a speech-to-text algorithm maps at least one OOI to one of a plurality of OOI medical terms.
[0039] In at least one embodiment, the at least one processor is further configured to generate at least one new training image from an object of interest (OOI) detected while analyzing the medical image data when the at least one text string associated with the OOI is determined to be a ground truth for the OOI based on a speech-to-text algorithm that generates input speech that matches the at least one text string.
[0040] In at least one embodiment, the at least one processor is further configured to generate at least one new training image from an object of interest (OOI) detected while analyzing the medical image data when it is determined that the at least one text string associated with the OOI is not a ground truth for that OOI based on a speech-to-text algorithm that generates input speech that matches the at least one text string.
[0041] In at least one embodiment, training is performed on medical image data obtained from one or more endoscopic procedures, one or more MRI scans, one or more CT scans, one or more X-rays, one or more ultrasound images, one or more nuclear medicine images, or one or more histological images.
[0042] In another broad aspect, in accordance with the teachings herein, in at least one embodiment, there is provided a method for analyzing medical image data for a medical procedure, the method including: receiving at least one image from a series of images; determining when at least one object of interest (OOI) is present in the at least one image and, when the at least one OOI is present, determining a classification of the at least one OOI, both determinings being performed using at least one machine learning model; displaying the at least one image and any determined OOI to a user on a display during the medical procedure; receiving an input audio signal including an utterance from the user during the medical procedure and recognizing the utterance; when the utterance is recognized as a comment on the at least one image during the medical procedure, converting the utterance into at least one text string using a speech-to-text conversion algorithm; matching the at least one text string with at least one image provided with the utterance from the user; and generating at least one annotated image in which the at least one text string corresponds to the at least one image.
[0043] In at least one embodiment, the method further includes, when the utterance is recognized as including a request for at least one reference image including a classification, displaying at least one reference image having an OOI classified with the same classification as the at least one OOI, and receiving input from the user confirming or rejecting the classification of the at least one OOI.
[0044] In at least one embodiment, the method further includes receiving an input from a user indicating a user classification for the at least one image having an undetermined OOI when the at least one OOI is classified as suspicious.
[0045] In at least one embodiment, the method further includes automatically generating a report that includes the at least one annotated image.
[0046] In at least one embodiment, the method further includes, for a given OOI in a given image, identifying bounding box coordinates of a bounding box associated with the given OOI in the given image, calculating a confidence score based on a probability distribution of a classification of the given OOI, and overlaying the bounding box on at least one image with the bounding box coordinates when the confidence score is higher than a confidence threshold.
[0047] In at least one embodiment, the method further includes determining a classification of the OOI by applying a convolutional neural network (CNN) to the OOI by performing convolution, activation, and pooling operations to generate a matrix, processing the matrix using the convolution, activation, and pooling operations to generate a feature vector, and performing classification of the OOI based on the feature vector.
[0048] In at least one embodiment, the method further includes, when generating the at least one annotated image, overlaying a timestamp onto the corresponding at least one image.
[0049] In at least one embodiment, the method further includes indicating the confidence score of the at least one image in real time on a display or in a report.
[0050] In at least one embodiment, receiving input audio during the medical procedure includes starting to receive an audio stream of input audio from a user upon detection of a first user action including pausing the display of the series of images, taking a snapshot of a given image in the series of images, or providing a first voice command, and ending receiving the audio stream upon detection of a second user action including remaining silent for a predetermined length of time, pressing a designated button, or providing a final voice command.
[0051] In at least one embodiment, the method further includes storing a series of images upon receiving input audio during the medical procedure, thereby designating at least one image to receive annotation data for generating at least one corresponding annotated image.
[0052] In at least one embodiment, the method further includes generating a report of the medical procedure by capturing a set of patient information data for addition to the report, loading a subset of the series of images including at least one annotated image, and combining the set of patient information data and the subset of the series of images including at least one annotated image into the report.
[0053] In at least one embodiment, the method further includes performing training of the at least one machine learning model by applying an encoder to the at least one training image to generate at least one feature vector for a training OOI in the at least one training image, selecting a class of the training OOI by applying the at least one feature vector to the at least one machine learning model, and reconstructing, using a decoder, a labeled training image by associating the at least one feature vector with the at least one training image and the selected class for training the at least one machine learning model.
[0054] In at least one embodiment, the classes are a healthy tissue class, an unhealthy tissue class, a suspect tissue class, or an unfocused tissue class.
[0055] In at least one embodiment, the method further includes training at least one machine learning model using a training dataset including labeled training images, unlabeled training images, or a mixture of labeled and unlabeled training images, where the images include examples categorized by healthy tissue, unhealthy tissue, suspicious tissue, and out-of-focus tissue.
[0056] In at least one embodiment, the method further includes training at least one machine learning model using supervised, unsupervised, or semi-supervised learning.
[0057] In at least one embodiment, the training data set further includes subcategories for each of the unhealthy and suspicious tissues.
[0058] In at least one embodiment, the method further includes creating at least one machine learning model by receiving training images as input to an encoder, projecting the training images using the encoder onto features that are part of a feature space, and mapping the features to a set of target classes using a classifier, and identifying morphological characteristics of the training images to generate a new training dataset, the new training dataset having data linking parameters to the training images, and determining based on the morphological characteristics whether there are one or more mapped classes or no mapped classes.
[0059] In at least one embodiment, the method further includes determining a classification of the at least one OOI by receiving one or more features as input to a decoder, mapping one of the features to an unlabeled dataset using a deconvolutional neural network, and reconstructing a new training image from one of the features using the decoder to train at least one machine learning model.
[0060] In at least one embodiment, the method further includes training a speech-to-text algorithm using the speech dataset to compare new speech data with the speech dataset to identify matches to the ground truth text, where the speech dataset includes the ground truth text and speech data of the ground truth text.
[0061] In at least one embodiment, a speech-to-text algorithm maps at least one OOI to one of a plurality of OOI medical terms.
[0062] In at least one embodiment, the medical image data is obtained from one or more endoscopic procedures, one or more MRI scans, one or more CT scans, one or more X-rays, one or more ultrasound photographs, one or more nuclear medicine images, or one or more histological images.
[0063] In another broad aspect, in accordance with the teachings herein, in at least one embodiment, there is provided a method for training at least one machine learning model and a speech-to-text conversion algorithm for use in analyzing medical image data for a medical procedure, the method including: applying an encoder to at least one training image to generate at least one feature for a training object of interest (OOI) in the at least one training image; selecting a class of the training OOI by applying the at least one feature to the at least one machine learning model; reconstructing labeled training images using a decoder by associating the at least one feature with the training image and the selected class for training the at least one machine learning model; training the speech-to-text conversion algorithm to identify matches between new speech data and the ground truth text using a speech dataset including ground truth text and speech data for the ground truth text, thereby generating at least one text string; and overlaying the training OOI and the at least one text string on an annotated image.
[0064] In at least one embodiment, the classes are a healthy tissue class, an unhealthy tissue class, a suspect tissue class, or an unfocused tissue class.
[0065] In at least one embodiment, the method further includes training at least one machine learning model using a training dataset including labeled training images, unlabeled training images, or a mixture of labeled and unlabeled training images, where the images include examples categorized by healthy tissue, unhealthy tissue, suspicious tissue, and out-of-focus tissue.
[0066] In at least one embodiment, training the at least one machine learning model includes using supervised learning, unsupervised learning, or semi-supervised learning.
[0067] In at least one embodiment, the training data set further includes subcategories for each of the unhealthy and suspicious tissues.
[0068] In at least one embodiment, the method further includes creating at least one machine learning model by receiving training images as input to the encoder, projecting the training images using the encoder into a feature space including features, mapping the features to a set of target classes using a classifier, and identifying morphological characteristics of the training images to generate a training dataset, the training dataset having data linking parameters to the training images, and determining based on the morphological characteristics whether there are one or more mapped classes or no mapped classes.
[0069] In at least one embodiment, the method further includes receiving one or more features as input to a decoder, mapping one of the features to an unlabeled dataset using a deconvolutional neural network, and reconstructing a new training image from one of the features using the decoder to train at least one machine learning model.
[0070] In at least one embodiment, a speech-to-text algorithm maps at least one OOI to one of a plurality of OOI medical terms.
[0071] In at least one embodiment, the method further includes generating at least one new training image from an object of interest (OOI) detected while analyzing the medical image data when the at least one text string associated with the OOI is determined to be a ground truth for the OOI based on a speech-to-text algorithm that generates input speech that matches the at least one text string.
[0072] In at least one embodiment, the method further includes generating at least one new training image from an object of interest (OOI) detected while analyzing the medical image data when it is determined that the at least one text string associated with the OOI is not a ground truth for that OOI based on a speech-to-text algorithm that generates input speech that matches the at least one text string.
[0073] In at least one embodiment, training is performed on medical image data obtained from one or more endoscopic procedures, one or more MRI scans, one or more CT scans, one or more X-rays, one or more ultrasound images, one or more nuclear medicine images, or one or more histological images.
[0074] Other features and advantages of the present application will become apparent from the following detailed description taken in conjunction with the accompanying drawings. It should be understood, however, that the detailed description and specific examples, while indicating preferred embodiments of the present application, are given by way of illustration only, since various changes and modifications within the spirit and scope of the present application will become apparent to those skilled in the art from this detailed description.
[0075] For a better understanding of the various embodiments described herein, and to show more clearly how these may be put into practice, reference is made by way of example to the accompanying drawings, in which at least one exemplary embodiment is shown and which are now described, and which are not intended to limit the scope of the teachings described herein. [Brief description of the drawings]
[0076] [Figure 1] 1 is a block diagram of an exemplary embodiment of a system for processing medical procedure images, such as, but not limited to, endoscopic images, in real time. [Diagram 2] 2A-2C are diagrams of an exemplary endoscopy apparatus setup and an alternative exemplary embodiment of an endoscopy image analysis system for use with the system of FIG. 1. [Diagram 3] FIG. 3 is a block diagram of an exemplary embodiment of hardware components and data flow of a computing device for use with the endoscopic image analysis system of FIG. [Figure 4] FIG. 2 is a block diagram of an exemplary embodiment of the interaction between input speech and a real-time annotation process. [Figure 5A] FIG. 2 is a block diagram of an exemplary embodiment of a method for processing an input audio stream and a sequence of input images in a real-time annotation process. [Figure 5B] 5B is a block diagram of an exemplary embodiment of a method for initiating and terminating capture of the input audio stream of FIG. 5A. [Figure 5C] 1 is a block diagram of an exemplary embodiment of a method for processing an input audio stream using a speech recognition algorithm. [Figure 6] 3 is a block diagram of an exemplary embodiment of a method for performing image analysis during an endoscopy procedure using the system of FIG. 2. [Figure 7] FIG. 2 is a block diagram of an exemplary embodiment of an image analysis training algorithm. [Figure 8A] FIG. 1 is a block diagram of a first exemplary embodiment of a U-net architecture used by an object detection algorithm. [Figure 8B] FIG. 13 is a detailed block diagram of a second exemplary embodiment of the U-net architecture used by the object detection algorithm. [Figure 9] FIG. 1 shows an example of an endoscopic image with healthy morphological characteristics. [Figure 10] FIG. 1 shows an example of an endoscopic image with unhealthy morphological characteristics. [Figure 11] FIG. 1 shows an example of an unlabeled video frame image from an exclusive dataset. [Figure 12] FIG. 2 is a block diagram of an example embodiment of a report generation process. [Figure 13] FIG. 2 is a block diagram of an example embodiment of a method for processing an input video stream using video processing and annotation algorithms. [Figure 14] 11 is a chart of training results showing the rate of positive speech recognition results versus the true positive value. [Figure 15] FIG. 2 is a block diagram of an exemplary embodiment of a speech recognition algorithm. [Figure 16] FIG. 2 is a block diagram of an example embodiment of an object detection algorithm that may be used by the image analysis algorithm. [Figure 17] FIG. 1 illustrates an example embodiment of a report including annotated images. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0077] Further aspects and features of the exemplary embodiments described herein will become apparent from the following description taken in conjunction with the accompanying drawings.
[0078] Various embodiments according to the teachings of the present specification are described below to provide an example of at least one embodiment of the claimed subject matter. Any embodiment described herein does not limit the claimed subject matter. The claimed subject matter is not limited to a device, system, or method having all of the features of any one of the devices, systems, or methods described below, or features common to more than one or all of the devices, systems, or methods described herein. There may be a device, system, or method described herein that is not an embodiment of any claimed subject matter. Any subject matter described herein that is not claimed herein may be the subject of another means of protection, for example, a continuing patent application, and the applicant, inventor, or owner does not intend to abandon, disclaim, or offer to the public any such subject matter by its disclosure herein.
[0079] It should be understood that for simplicity and clarity of description, where deemed appropriate, reference numerals may be repeated among the figures to indicate corresponding or similar elements. Additionally, numerous specific details have been described to provide a thorough understanding of the embodiments described herein. However, it will be understood by those skilled in the art that the embodiments described herein may be practiced without these specific details. In other instances, well-known methods, procedures, and components have not been described in detail so as not to obscure the embodiments described herein. Additionally, the description should not be considered as limiting the scope of the embodiments described herein.
[0080] It should also be noted that the terms "coupled" or "couple" as used herein can have several different meanings depending on the context in which the terms are used. For example, the terms coupled or couple can have a mechanical or electrical meaning. For example, as used herein, the terms coupled or couple can indicate that two elements or devices can be directly connected to each other, or can be connected to each other via one or more intermediate elements or devices, via electrical signals, electrical connections, or mechanical elements, depending on the particular context.
[0081] It should also be noted that, as used herein, the term "and / or" is intended to represent an inclusive "or." That is, "X and / or Y" is intended to mean, for example, X or Y or both. As a further example, "X, Y, and / or Z" is intended to mean X, Y, Z, or any combination thereof.
[0082] It should be noted that terms of degree, such as "substantially," "about," and "approximately," as used herein, refer to a reasonable amount of deviation from the modified term such that the end result is not materially altered. These terms of degree may also be construed to include deviations from the modified term, such as 1%, 2%, 5%, or 10%, if this deviation does not negate the meaning of the term it modifies.
[0083] Additionally, the recitation of numerical ranges herein by endpoints includes all numbers and fractions subsumed within that range (e.g., 1 to 5 includes 1, 1.5, 2, 2.75, 3, 3.90, 4, and 5). It is also understood that all such numbers and fractions are presumed to be modified by the term "about," which refers to a variation of the referenced number by a certain amount, such as, for example, 1%, 2%, 5%, or 10%, if the end result would not be significantly altered.
[0084] Also, it should be noted that the use of the term "window" in describing the operation of any system or method described herein is intended to be understood as describing a user interface for performing initialization, configuration, or other user operations.
[0085] Exemplary embodiments of the device, system, or method described by the teachings herein may be implemented as a combination of hardware and software. For example, the embodiments described herein may be implemented, at least in part, by using one or more computer programs to run on one or more programmable devices including at least one processing element and at least one storage element (i.e., at least one volatile memory element and at least one non-volatile memory element (memory elements may also be referred to as memory units herein)). The hardware may include input devices including at least one of a touch screen, a touch pad, a microphone, a keyboard, a mouse, a button, a key, a slider, an electroencephalogram (EEG) input device, an eye tracking device, etc., as well as one or more of a display, a printer, etc., depending on the hardware implementation.
[0086] It should also be noted that some elements used to implement at least some of the embodiments described herein may be implemented via software written in a high-level procedural language, such as object-oriented programming. ++ , C#, JavaScript, Python, or any other suitable programming language and may include modules or classes as known to those skilled in the art of object-oriented programming. Alternatively, or in addition, some of these elements that are implemented via software may be written in assembly language, machine language, or firmware, as appropriate. In either case, the language may be a compiled or interpreted language.
[0087] At least some of these software programs may be stored on a computer readable medium, such as, but not limited to, a ROM, a magnetic disk, an optical disk, a USB key, or on the cloud, that is readable (or accessible) by a device having a processor, an operating system, and associated hardware and software necessary to implement the functionality of at least one of the embodiments described herein. The software program code, when read by the device, configures the device to operate in a new specific predefined manner (e.g., an application specific computer) to perform at least one of the methods described herein.
[0088] At least some of the programs associated with the device, system, and method embodiments described herein may be dispersible in a computer program product that includes a computer-readable medium carrying computer usable instructions, such as program code, for one or more processing units. The medium may be provided in a variety of forms, including, but not limited to, one or more diskettes, compact discs, tapes, chips, and non-transitory forms, such as magnetic and electronic storage devices. In alternative embodiments, the medium may be transitory in nature, such as, but not limited to, wired transmissions, satellite transmissions, Internet transmissions (e.g., downloads), media, digital signals, and analog signals. The computer usable instructions may also be in a variety of formats, including compiled and non-compiled code.
[0089] In accordance with the teachings herein, various embodiments are provided for systems and methods, and computer products for use therewith, for processing medical images of various modalities. The processing may be performed in real-time.
[0090] In at least one embodiment of the system, the system provides an improvement over conventional systems that analyze medical imaging data for a medical procedure to generate an annotated image from a sequence of images, such as, for example, a video feed, taken during the medical procedure. The medical procedure can be a medical diagnostic procedure. For example, the system receives an image, which may be a video frame from a sequence of video frames or may be taken from a sequence of images, such as, for example, one or more images of one or more corresponding CT slices or MRI slices. The system determines when an object of interest (OOI) is present in the image and, when an OOI is present, determines a classification of the OOI. The system makes both of these determinations using at least one machine learning model. The system displays the image and any determined OOI to a user on a display during the medical procedure. The system also receives input speech from the user during the medical procedure. The system recognizes speech from the input speech and converts the speech to a text string using a speech-to-text conversion algorithm. In some cases, the system matches the text string to a corresponding image. The system generates an annotated image in which the text string is linked (e.g., overlaid) to the corresponding image. In at least one alternative embodiment, the text string may include a command to view an image (which may be referred to as a reference image) from a library or database that has been classified similarly to the OOI and may be displayed to enable a user to compare a given image from a series of images (e.g., a series of video frames or a series of images from CT or MRI slices) to the reference image to determine whether the automatic classification of the OOI is correct. Medical Imaging Technology
[0091] Various embodiments of the systems and methods for processing medical images in real time described herein are applicable to various medical imaging technologies. One advantage of the embodiments described herein includes providing speech recognition to generate text in real time that can be used to (a) identify / mark areas of interest in an image, where the areas of interest may be anomalies, areas of structural damage, areas of physiological changes, or therapeutic targets, and / or (b) mark / tag areas of interest in an image for next steps in therapy or treatment. Another advantage includes the ability to generate instant reports (e.g., images may be included in the report based on the identification / marking / tagging as well as the generated text or portions thereof). Another advantage includes displaying previously annotated or characterized images in real time that are similar to the OOI identified by the operator to enhance and support the operator's diagnostic capabilities.
[0092] Various embodiments described herein can also be applied to voice-to-text technology during a procedure, such as the opportunity to provide real-time, time-stamped documentation of intraprocedural events for quality assurance and clinical notes. In an endoscopy procedure, for example, this would include documentation of patient symptoms (e.g., pain), administration of pain medication, repositioning of the patient, etc. These data can then be recorded simultaneously with other monitoring information, patient physiological parameters (e.g., pulse, BP, oximetry), instrument manipulation, etc.
[0093] Table 1 below provides a non-exhaustive list of examples of clinical applications for using various embodiments of the systems and methods for processing medical images described herein. [Table 1-1] [Table 1-2]
[0094] The additional clinical applications in Table 1 reflect the fact that "endoscopic" techniques are used in many other specialties where there is a need for real-time identification and real-time documentation of abnormalities by an operator busy with the visuomotor requirements of performing the procedure. Most "endoscopic" procedures are primarily diagnostic, although therapeutic interventions are becoming more prevalent.
[0095] In contrast, surgical laparoscopy is primarily therapeutic, although it is based on precise identification of therapeutic targets. Many operations are long, leaving little opportunity for comprehensive documentation of procedural events and interventions, which must instead be done post-procedure from memory.
[0096] It should be noted that most specialists incorporate histopathology into their management plans, but histopathology diagnosis, reporting, etc. is performed by the histopathologist. One advantage of the embodiments described herein is that it provides a mechanism for the histopathologist to identify, localize and annotate images or OOIs in real time during the examination, generate subsequent reports, and access comparable images / OOIs from a databank.
[0097] Another advantage of the embodiments described herein is that it provides an option to mark the location of an OOI in an image using voice control / annotation, which may be applied to radiology and histopathology: a radiologist or pathologist can identify a lesion as an OOI and at the same time annotate the OOI with voice-to-text technology using a standardized vocabulary.
[0098] Intra-procedural image or video annotation, potentially including OOI localization using voice-text, is a means of documenting or reporting a procedure (eg, based on a video recording of a laparoscopic surgical procedure). Endoscopy applications
[0099] Various embodiments of the systems and methods for processing medical images described in accordance with the teachings herein are described with images obtained from GI endoscopy for illustrative purposes. It should therefore be understood that the systems and methods described herein may be used with medical images generated from different types of endoscopy applications, or other medical applications in which images are obtained using other imaging modalities, such as the examples shown in Table 1. Various endoscopy applications in which the systems and methods described herein may be used may include, but are not limited to, respiratory, ENT, obstetrics and gynecology, cardiology, urology, neurology, orthopedics, and general surgery. Respiratory system:
[0100] Endoscopy applications include, but are not limited to, flexible bronchoscopy and medical thoracoscopy, such as endobronchial ultrasound and navigational bronchoscopy, based on the use of standardized endoscopic platforms, with or without narrow band imaging (NBI). ENT:
[0101] Endoscopy applications include, but are not limited to, surgical procedures to address hearing complications, such as myringotomy or other ENT surgery, surgical procedures to address laryngeal diseases affecting the epiglottis, tongue, or vocal cords, surgical procedures on the maxillary sinuses, nasal polyps, or other clinical or structural evaluations integrated into the ENT doctor's decision support system. Obstetrics and Gynecology:
[0102] Endoscopy applications include, but are not limited to, structural and pathological evaluation and diagnosis of conditions related to obstetrics and gynecology, such as, for example, minimally invasive surgery (including robotic surgical techniques), laparoscopic surgery, etc. Cardiology:
[0103] Endoscopy applications include, but are not limited to, structural and pathological evaluation and diagnosis of diseases related to cardiology, such as minimally invasive surgery (including robotic surgical techniques). Urology:
[0104] Endoscopy applications include procedures used in the diagnosis and treatment of renal disease, renal structural and pathological evaluation, and therapeutic procedures (including robotic and minimally invasive surgery), as well as applications including, but not limited to, the treatment of kidney stones, cancer, and the like, as localized and / or surgical treatments. Neurology (Central Nervous System / Spine):
[0105] Endoscopy applications include, but are not limited to, structural and pathological evaluation of the spine, such as minimally invasive spine surgery based on standardized techniques or 3D imaging. Orthopedics:
[0106] Endoscopy applications include, but are not limited to, joint surgery.
[0107] Referring initially to FIG. 1 , a block diagram of an exemplary embodiment of an automated system 100 for detecting morphological characteristics in a medical procedure and annotating one or more images in real time is shown. The medical procedure may be a medical diagnostic procedure. When used in the context of endoscopy, the system 100 may be referred to as an endoscopic image analysis (EIA) system. However, as previously mentioned, the system 100 may be used in conjunction with other imaging modalities and / or medical diagnostic procedures. The system 100 may communicate with at least one user device 110. In some embodiments, the system 100 may be implemented by a server. The user device 110 and the system 100 may communicate over a communication network 105, which may be, for example, wired or wireless. The communication network 105 may be, for example, the Internet, a wide area network (WAN), a local area network (LAN), WiFi, Bluetooth, etc.
[0108] The user device 110 may be a computing device operated by a user. The user device 110 may be, for example, a smartphone, a smartwatch, a tablet computer, a laptop, a virtual reality (VR) device, or an augmented reality (AR) device. The user device 110 may be a combination of computing devices operating together, such as, for example, a smartphone and a sensor. The user device 110 may also be, for example, a device remotely operated by a user, in which case the user device 110 may be operated by the user via, for example, a personal computing device (such as a smartphone). The user device 110 may be configured to execute applications (e.g., mobile apps) that communicate with some parts of the system 100.
[0109] System 100 can run on a single computer. System 100 includes a processor unit 124, a display 126, a user interface 128, an interface unit 130, input / output (I / O) hardware 132, a network unit 134, a power supply unit 136, and a memory unit (also called a "data store") 138. In other embodiments, system 100 can have more or fewer components, but generally function in a similar manner. For example, system 100 can be implemented using multiple computing devices or systems.
[0110] The processor unit 124 may include a standard processor, such as an Intel Xeon processor. Alternatively, there may be multiple processors used by the processor unit 124, which may work in parallel to perform certain functions. The display 126 may be, but is not limited to, a computer monitor or an LCD display, such as for a tablet device. The user interface 128 may be an application programming interface (API) or web-based application accessible via the network unit 134. The network unit 134 may be a standard network adapter, such as an Ethernet or 802.11x adapter.
[0111] The processor unit 124 may operate with a prediction engine 152, which may be implemented using one or more standalone processors, such as a graphical processing unit (GPU), that functions to provide predictions using the machine learning models 146 stored in the memory unit 138. The prediction engine 152 may build one or more prediction algorithms by applying training data to one or more machine learning algorithms. The training data may include, for example, image data, video data, audio data, and text. The prediction may involve first identifying an object in the image and then determining its classification. For example, the training may be based on morphological characteristics of an OOI, such as, for example, a polyp or at least one other physiological structure that may be encountered in other medical diagnostic / surgical applications or other imaging modalities, and then during image analysis, the image analysis software first identifies whether a newly acquired image has an OOI that matches the morphological characteristics of the image of a polyp, and if so, predicts that the OOI is a polyp or at least one other physiological structure. This may include determining a confidence score that the OOI has been correctly identified.
[0112] The processor unit 124 may also execute software instructions for a graphical user interface (GUI) engine 154, which is used to generate the various GUIs. The GUI engine 154 provides data according to a fixed layout for each user interface, as well as receiving data or control input from a user. The GUI engine 154 may then use the input from the user to change the data shown on the display 126 or to change the operation of the system 100, which may include presenting a different GUI.
[0113] The memory unit 138 may store program instructions for an operating system 140, program code 142 for other applications (also referred to as "programs 142"), an input module 144, multiple machine learning models 146, an output module 148, a database 150, and a GUI engine 154. The machine learning models 146 may include, but are not limited to, image recognition and classification algorithms based on deep learning models and other approaches. The database 150 may be, for example, a local database stored in the memory unit 138, or in other embodiments, an external database, such as a database on the cloud, multiple databases, or a combination thereof.
[0114] In at least one embodiment, the machine learning model 146 includes a convolutional neural network (CNN), a recurrent neural network (RNN), and / or other suitable implementation of predictive modeling (e.g., a multi-layer perceptron). A CNN is designed to recognize images and patterns. A CNN performs convolution operations and can be used, for example, to classify regions of an image and look at edges of objects recognized within an image region. Since an RNN can be used to recognize sequences such as text, speech, time evolution, etc., an RNN can be applied to a sequence of data to predict what will happen next. Thus, a CNN can be used to detect what is happening or to detect at least one physiological structure on a given image at a given time, and an RNN can be used to provide an information message (e.g., classification of an OOI).
[0115] The programs 142 comprise program code that, when executed, configures the processor unit 124 to operate in a particular manner to implement various functions and tools for the system 100. The programs 142 include program code that may be used for a variety of algorithms, including image analysis algorithms, speech recognition algorithms, text matching algorithms, and term correction algorithms.
[0116] Referring to FIG. 2, a diagram of an exemplary setup 200 of a system for acquiring and processing medical images in real time is shown. The setup 200 shown in FIG. 2 illustrates a system for acquiring and processing endoscopic images as a specific example of medical images, but may be used for other medical applications and / or medical imaging modalities. The setup 200 includes an endoscope system and an endoscopic image analysis (EIA) system 242. The endoscope system includes five main components: an endoscope platform 210, a main image processor 215, an endoscope 220, a handheld controller 225, and an endoscope monitor 240. The endoscopic image analysis system includes elements 245-270.
[0117] The main image processor 215 receives input via the endoscope 220. The endoscope 220 may be any endoscope suitable for insertion into a patient. In other embodiments, for other medical applications and / or imaging modalities, the endoscope is replaced with another imaging device and / or sensor as described below to acquire images, such as the examples shown in Table 1. The main image processor 215 also receives input from a user when the endoscope 220 is inserted into the digestive tract or other body site and the camera of the endoscope 220 is used to capture images (e.g., image signals). The main image processor 215 receives image signals from the endoscope 220 that can be processed to be displayed or output. For example, the main image processor 215 sends images captured by the endoscope 220 to the endoscope monitor 240 for display. The endoscope monitor 240 may be any monitor suitable for an endoscopic procedure that is compatible with the endoscope 220 and the main image processor 215. For other medical imaging modalities, the main image processor 215 may receive images from other devices / platforms, such as a CT scanning device, an ultrasound device, an MRI scanner, an X-ray device, a nuclear medicine imaging device, a histology imaging device, etc., and accordingly, the output from the endoscope 220 is replaced by the output from each of these devices / platforms in their applications, such as the examples shown in Table 1.
[0118] The image processing unit 235 controls the processing of image signals from the endoscope 220. The image processing unit 235 includes a main image processor 215 that is used to receive image signals from the endoscope 220 and then process the image signals in a manner consistent with conventional image processing performed by a camera. The main image processor 215 then controls the display of the processed images on the endoscope monitor 240 by sending image data and control signals to the endoscope monitor 240 via a connecting cable 236.
[0119] The endoscope 220 is connected to a handheld control panel 225 consisting of programmed buttons 230. The handheld control panel 225 and the programmed buttons 230 may be part of the input module 144. The programmed buttons 230 may be pressed to send an input signal to control the endoscope 220. The programmed buttons 230 may be actuated by a user (which may be a clinician, gastroenterologist, or other medical professional) to send an input signal to the main image processor 215, which may be used to instruct the main image processor 215 to pause the display of a series of images (e.g., a video stream or a sequence of video frames) or to take a snapshot of a given image in the series of images (e.g., a video frame of a video stream or a video frame in a sequence of video frames). The input signal may temporarily interrupt the display of a series of images (e.g., a video stream displayed on the endoscope monitor 240), allowing the server 120 to detect a particular image (e.g., a video frame) to be annotated.
[0120] In at least one embodiment, the endoscope 220 is replaced with an imaging device that produces another type of image that may or may not together form a video (e.g., slices produced by an MRI device). In such a case, the sequence of images is a series of those images (e.g., a series of slices).
[0121] The EIA system 242 provides an analysis platform, such as an AI-based analysis platform, that includes one or more components used to analyze images acquired by the endoscope 220 and provide corresponding annotated versions of those images, as well as other functions. The EIA system 242 may be considered to be an alternative exemplary embodiment of the system 100. More generally, the EIA system 242 may be considered to be an alternative exemplary embodiment of the system 100 when used with other medical imaging modalities. In such cases, any reference to endoscopy, endoscope, or endoscopic images may be replaced with other medical imaging procedures, imaging modalities, imaging devices, or medical images, respectively, such as the examples shown in Table 1.
[0122] In this exemplary embodiment, the EIA system 242 includes a microcomputer 255 that can be connected to the endoscope monitor 240 via, for example, an HDMI cable 245 to receive the endoscope images. The HDMI cable 245 can be any standard HDMI cable. A conversion key 250 allows an HDMI port of the endoscope monitor 240 to be connected to a USB port of the microcomputer 255. The microcomputer 255 is communicatively coupled to one or more memory devices, such as the memory unit 138, in which the programs 142, the prediction engine 152, and the machine learning models 146 are collectively stored. The microcomputer 255 executes image analysis software program instructions to apply image analysis algorithms to image signals collected by the endoscope 220.
[0123] The microcomputer 255 may be, for example, an NVIDIA Jetson microcomputer with a CPU and a GPU, along with one or more memory elements. In addition, the image analysis algorithm includes an object detection algorithm that may be based on YOLOv4, which uses a convolutional neural network (e.g., as shown in FIG. 16) to perform certain functions. The YOLOv4 object detection algorithm may be advantageous because the EIA system may analyze images faster. The YOLOv4 object detection algorithm may be implemented by, for example, an NVIDIA Jetson microcomputer with a software accelerator, such as TensorRT, Raspberry Pi, or TensorFlow.
[0124] The software accelerator TensorRT may be advantageous because it allows the EIA system 242 to train the machine learning model 146 at a faster rate using a GPU, such as an NVIDIA GPU. The software accelerator TensorRT may provide additional benefits to the EIA system 242 by allowing modification of the machine learning model 146 without affecting the performance of the EIA system 242. The software accelerator TensorRT may achieve these advantages of the EIA system 242 using certain functions such as layer fusion, block fusion, and float-to-int converter. When the EIA system 242 uses YOLOv4, the software accelerator TensorRT may improve the performance speed of YOLOv4.
[0125] The microcomputer 255 may be connected to a microphone 270 via a USB connection 268. The microphone 270 receives acoustic signals that may include user input, such as during a medical procedure (e.g., a medical diagnostic procedure), and converts the acoustic signals into input voice signals. The microphone 270 may be considered to be part of the I / O hardware 132. One or more processors of the microcomputer 255 may receive the input voice signals captured by the microphone 270 by operation of the input module software 144. The microcomputer 255 may then apply a speech recognition algorithm to the input voice signals collected by the microphone 270. The speech recognition algorithm may be implemented using one or more of the programs 142, the prediction engine 152, and the machine learning model 146.
[0126] The image analysis monitor 265 may be connected to the microcomputer 255 via an HDMI connection using a standard HDMI cable 260. The microcomputer 255 displays the results of the image analysis algorithms and the speech recognition algorithms on the image analysis monitor 265. For example, for a given image, the image analysis monitor 265 may display one or more OOIs with bounding boxes placed around each OI, and optionally, color indicators may be used for the bounding boxes to indicate specific information about the elements contained within the bounding boxes. The annotations generated by the speech recognition and voice-to-text algorithms may be stored in the database 150 or some other data store. The voice-to-text algorithms may be implemented using one or more of the programs 142, the prediction engine 152, and the machine learning models 146. The microcomputer 255 displays the annotations on the image analysis monitor 265.
[0127] It should be noted that in at least one embodiment described herein, the confidence score may also be generated by the image analysis software. This may be done by comparing each pixel of the determined bounding box of the OOI determined for a given image (i.e., a given video frame) to the ground truth of the object based on the classification of the object, e.g., a polyp. The confidence score may be defined, for example, as a decimal between 0 and 1, which may be interpreted as a percentage of confidence. The confidence score may then represent the level of agreement between multiple contributors and indicate a "confidence" in the validity of the result. The aggregated result may be selected based on the most confident answer. The confidence score is then compared to a pre-set confidence threshold, which may be adjusted over time to improve performance. If the confidence score is greater than the confidence threshold, the bounding box, classification, and optionally the confidence score may be displayed to the user along with the given image during the medical procedure. Alternatively, if the confidence score is less than the confidence threshold, the image analysis system may label the given image as suspicious and display this label to the user along with the given image. In at least one implementation, the confidence score is the output of the network. In such a case, the object detection model may output the object class, the object location, and / or the confidence score. The confidence score may be generated by the neural network by performing convolution, activation, and pooling operations. An example of how the confidence score is generated may be shown in FIG. 16.
[0128] 3, a block diagram of an exemplary embodiment of hardware components and data flows 300 of a computing device for use with the microcomputer 255 of the EIA system 242 is shown. As described herein with reference to FIG. 3, the hardware components and data flows 300 may be used with the EIA system 242 in the context of endoscopy. More generally, however, the EIA system 242 may be considered an alternative exemplary embodiment of the system 100 when used for other medical imaging applications and imaging modalities. In such cases, any subsequent references to endoscopy, endoscope, and endoscopic images may be replaced with other medical imaging procedures, imaging modalities, imaging devices, or medical images, respectively, such as the examples shown in Table 1.
[0129] The microcomputer 255 is implemented on an electronic board 310 with various input and output ports. The microcomputer 255 generally comprises a CPU 255C, a GPU 255G, and a memory unit 255M. For example, the microcomputer 255 may be hardware designed for high-performance AI systems such as medical equipment, high-resolution sensors, or automated optical inspection, with a GPU 255G with NVIDIA CUDA cores, and a CPU 255C with NVIDIA Camel ARM, Vision Accelerator, Video Encode, and Video Decode. The data flow 300 consists of input signals provided to the microcomputer 255 and output signals generated by the microcomputer and sent to one or more output devices, storage devices, or remote computing devices. The conversion key 250 receives the video input signal and directs the video input signal to the microcomputer USB video input port 370. Alternatively, the video input signal may be provided via a USB cable, in which case the conversion key 250 is not required and the microcomputer USB video input port 370 receives the video input signal. The microcomputer USB video input port 370 allows the microcomputer 255 to receive a real-time video input signal from the endoscope 220 .
[0130] The microcomputer 255 receives potential user input by directing an input audio signal from the microphone 270 to the microcomputer audio USB port 360. The microcomputer 255 then receives the input audio signal from the microcomputer audio USB port 360 for use in the speech recognition algorithm. Additional input devices may also be connected to the microcomputer 255 via optional USB connections 380. For example, the microcomputer 255 may be connected to two optional USB connections 380 (e.g., for a mouse and keyboard).
[0131] The microcomputer CPU 255C and GPU 255G work in combination to execute one or more of the program 142, the machine learning model 146, and the prediction engine 152. The microcomputer 255 may be configured to first store all output files in the memory unit 255M and then store all output files in an external memory. The external memory may be a USB memory card connected to the data output port 330. Alternatively or additionally, the external memory may be provided by the user device 110. Alternatively or additionally, the microcomputer 255 may provide the output data to another computer (or computing device) for storage. For example, the microcomputer 255 may store the output data in a secure cloud server. As another example, the microcomputer 255 may store and output data to the user device 110, which may be a smartphone with a compatible application.
[0132] The microcomputer 255 may have buttons 340 that allow a user to select one or more pre-programmed functions. The buttons 340 may be configured to provide control inputs for specific functions associated with the microcomputer 255. For example, one of the buttons 340 may be configured to turn on the microcomputer CPU 255C and / or GPU 255G, turn off the microcomputer CPU 255C and / or GPU 255G, initiate operation of a quality control process on the microcomputer 255, execute a GUI showing endoscopic images, including annotated images, and start and end annotations. The button 340 may also have an LED light 341 or other similar visual output device. The microcomputer 255 receives power through a power cable port 350. The power cable port 350 provides power to the various components of the microcomputer 255, allowing them to operate.
[0133] The microcomputer processor 255C can display the image analysis results on the monitor 265 via the microcomputer USB video output port 320. The monitor 265 can be connected to the microcomputer 255 via the microcomputer HDMI video output port 320 using an HDMI connection.
[0134] Referring to FIG. 4, a block diagram of an exemplary embodiment of a method 400 for processing input audio and video signals using a real-time annotation process 436 is shown. Although the method 400 and subsequent methods and processes are described as being performed by the EIA system 242, it should be noted that this is for illustrative purposes and it should be understood that the system 100 or another suitable processing system may be used. More generally, however, the EIA system 242 can be considered an alternative exemplary embodiment of the system 100 when used for other medical imaging applications and imaging modalities. In such cases, any reference to endoscopy, endoscope, and endoscopic images can be replaced with other medical imaging procedures, imaging modalities, imaging devices, or medical images, respectively, such as the examples shown in Table 1. The method 400 can be performed by the CPU 255C and the GPU 255G.
[0135] Method 400 can provide annotation process 436 in real time because EIA system 242 has GPU 255G and CPU 255C with high performance capabilities and the way the object detection algorithm is constructed. Alternatively or additionally, method 400 and object detection algorithm may be run on the cloud using AWS GPUs and a user can upload an endoscopy video and use a process similar to real-time annotation process 436 (e.g., simulating an endoscopy in real time or allowing pausing of the video).
[0136] At 405, before performing the real-time annotation process 436, the EIA system 242 places the speech recognition algorithm 410 in a standby state. While waiting, the speech recognition algorithm 410 waits for an input audio signal from the input module 144. The speech recognition algorithm 410 may be implemented using one or more of the program 142, the machine learning model 146, and the prediction engine 152.
[0137] At 420, the EIA system 242 receives a start signal 421 from a user at a first signal receiver to begin the real-time annotation process 436. The EIA system 242 receives an input audio signal via the microphone 270. For example, the signal receiver may be one of the buttons 340.
[0138] At 422, the EIA system 242 captures an input voice signal and converts the input voice signal into speech data by using a speech recognition algorithm 410, which may be implemented using the program 142. The speech data is then processed by a speech-to-text conversion algorithm to convert the speech data into one or more text strings that are used to create the annotation data. The EIA system 242 then uses an image and annotation data matching algorithm to determine which images the annotation data should be added to.
[0139] At 430, the image and annotation data matching algorithm determines from the input image series (e.g., the input video signal) the given image to which the text string in the annotation data corresponds, and then links the annotation data to the given image. Linking the annotation data to the given image may include, for example, (a) overlaying the annotation data on the given image, (b) superimposing the annotation data on the given image, (c) providing a hyperlink on the given image that links to a web page having the annotation data, (d) providing a pop-up window having the annotation data that pops up upon hovering a cursor over the given image or a relevant portion thereof, or (e) any equivalent link known to those skilled in the art. The image and annotation data matching algorithm may make this determination, for example, using matching timestamps for the capture of the image being annotated and the receipt of the annotation data. The input image series may be, for example, an input video signal from a video input stream acquired using the endoscope 220. In other imaging modalities, the input video signal may instead be a series of images as previously described.
[0140] At 432, the second signal receiver receives and processes the termination signal 422. For example, the second signal receiver may be another button 340 or may be the same as the first signal receiver. Upon receiving the termination signal 422, the EIA system 242 terminates the real-time annotation process 436. When the termination signal 422 is not received, the EIA system 242 continues the real-time annotation process 436 by continuing the operation of the speech recognition algorithm 410, the annotation capture, and the matching algorithm 430.
[0141] At 434, the EIA system 242 outputs one or more annotated images, which may (a) be displayed on a monitor or display, (b) be incorporated into a report, (c) be stored in a data storage element / device, and / or (d) be transmitted to another electronic device.
[0142] The microcomputer 255 includes an internal storage 440, such as a memory unit 255M. The internal storage 440 can be used to store data, such as a complete video or a portion thereof of an endoscopic procedure, one or more annotated images, and / or audio data. For example, the microcomputer 255 can capture audio data during the real-time annotation process 436 and store it in the internal storage 440. Alternatively or in addition, the microcomputer 255 can store annotated images in the internal storage 440.
[0143] Referring to FIG. 5A, a block diagram of an exemplary embodiment of a method 500 for processing an input audio stream and an input stream of a series of images (e.g., an input video stream) in a real-time annotation process 436 is shown. The method 500 may be executed by the CPU 255C and / or the GPU 255G. The method 500 is initiated by a start command signal 423 received as input by the EIA system 242. The speech recognition algorithm 410 receives the input audio signal and begins processing to start speech recognition. The EIA system 242 records the audio data determined by the speech recognition algorithm 410. The speech recognition algorithm 410 stops processing the input audio signal when a stop command signal 422 is received.
[0144] The speech-to-text algorithm 520 may be implemented using one or more of the programs 142, the prediction engine 152, and the machine learning models 146. For example, the speech-to-text algorithm 520 may be an open-source pre-trained algorithm such as Wav2vec 2.0, or any other suitable speech recognition algorithm. The speech-to-text algorithm 520 takes the speech data determined by the speech recognition algorithm 410 and converts the speech data into text 525 using an algorithm that may be a convolutional neural network (e.g., as shown in FIG. 15 ).
[0145] The text 525 is then processed by the term correction algorithm 530. The term correction algorithm 530 may be implemented using one or more of the program 142 and the prediction engine 152. The term correction algorithm 530 uses a string matching algorithm and a custom vocabulary to correct mistakes made by the speech-to-text conversation algorithm 520. The term correction algorithm 142 may be an open source algorithm such as Fuzzywuzzy. The text 525 is cross-referenced with each term in the custom vocabulary. The term correction algorithm 142 then calculates a match score based on how closely the text 525 matches the terms in the custom vocabulary. The term correction algorithm determines whether the matching score is higher than a threshold matching score. The term correction algorithm 530 replaces the text 525, or a portion thereof, with a term in the custom vocabulary if the matching score is higher than the threshold matching score.
[0146] The speech recognition output 540 may be referred to as annotation data, which includes annotations that the user adds to a given image that the user has commented on. The speech recognition output 540 is sent to the matching algorithm 430. The matching algorithm 430 may be implemented using the program 142 or the machine learning model 146. The matching algorithm 430 determines the matching image to which the annotation data corresponds (i.e., the image that the user verbally commented on, converted into annotation data), and overlays the annotation data from the speech recognition output 540 onto the matching image captured from the input stream of the sequence of images 510 from the endoscope 220 (e.g., a video input stream) to generate the annotated image output 434. The annotated image output 434 may be a key image 434-1 (e.g., with OOI) overlaid by the speech recognition output 540. The annotated image output 434 may be a video clip 434-2 overlaid by the speech recognition output 540. The key image 434-1 and the video clip 434-2 may be output by the server 120 and stored at 440.
[0147] In at least one embodiment, endoscope 220 is replaced with an imaging device that produces other types of images (e.g., slices produced by an MRI device). In such cases, key image 434-1 may be a different type of image (e.g., a slice) and video clip 434-2 may be replaced with a sequence of images (e.g., a sequence of slices).
[0148] The speech-to-text algorithm 520 can be trained using a speech data set that includes ground truth text and audio data of the ground truth text. New audio data can be compared to the new speech data set to identify matches with the ground truth text. Ground truth text and audio data of the ground truth text can be obtained for various medical applications and imaging modalities, examples of which are shown in Table 1.
[0149] Referring to FIG. 5B, a block diagram of an exemplary embodiment of a method 550 for initiating and terminating the capture of an input audio stream to be processed by the speech recognition algorithm 410 of FIG. 5A is shown. The method 550 may be executed by the CPU 255C. The EIA system 242 initiates the speech recognition algorithm 410 in response to an input start signal 423 (provided, for example, by user interaction), which may include a video pause command 560, a snapshot take command 562, or a voice start command 564. When the input signal provides the video pause command 560, the EIA system 242 pauses the input video stream. When the input signal 421 provides the snapshot take command 562, the EIA system 242 takes a snapshot of the input video stream, which involves capturing the particular image displayed when the snapshot take command 562 is received. When the input signal 421 provides a voice start command 564, such as "start annotation," the EIA system 242 begins annotation. For other medical applications and / or imaging modalities, other control actions may be performed, as known to those skilled in the art.
[0150] In at least one embodiment, EIA system 242 is replaced by an equivalent system for analyzing images obtained from an imaging device that produces other types of images (e.g., slices produced by an MRI machine). In such a case, video pause command 560 is replaced with a command to pause the display of a series of images (e.g., a series of slices).
[0151] The EIA system 242 terminates operation of the speech recognition algorithm 410 in response to an input termination signal 424 (e.g., generated by a user), which may include a silence input 570, a button press input 572, or a voice termination command 574. The silence input 570 may be, for example, an inaudible input, or an input voice below a threshold volume level. The silence input 570 may last, for example, at least five seconds to successfully terminate operation of the speech recognition algorithm 410. The button press input 572 may be the result of a user pressing a designated button, such as one of the buttons 340. A voice termination command 574, such as "stop annotation," may be used to stop annotation of the image.
[0152] Referring to FIG. 5C, a block diagram of a method 580 for processing an input audio stream, such as an audio signal 582, using a speech-to-text algorithm, such as a speech recognition and speech-to-text algorithm 520, cross-referenced with a custom vocabulary 584 is shown. The method 580 may be executed by one or more processors of the EIA system 242. The custom vocabulary 584 may be constructed before the EIA system 242 is operated and optionally updated from time to time. In other embodiments, the custom vocabulary 584 may be constructed for other medical applications and / or medical imaging modalities. The speech-to-text algorithm 520 receives the audio signal 582, which is typically a user-recorded input to the microphone 270. The ground truth 586 may be a set of terms specific to the medical procedure being performed, such as a gastrointestinal endoscopy, or another type of endoscopic procedure, or other medical procedure using another imaging modality as previously described. The ground truth 586 may be a database file stored in a database (such as the database 150). There may be multiple ground truth data sets for different categories of terms, such as stomach, colon, liver, etc. The ground truth 586 may initially consist of pre-determined terms specific to gastrointestinal endoscopy, or other medical applications and / or imaging modalities. Thus, the ground truth allows the speech-to-text conversion algorithm to map at least one OOI to one of multiple OOI medical terms. For example, one OOI may be mapped to multiple medical terms, as multiple features may occur, such as polyp and bleeding. The ground truth 586 may be advantageous because it allows for updates and accuracy analysis of the speech recognition algorithm 520. The EIA system 242 may receive user input from a keyboard and / or microphone that updates the ground truth 586. A user may provide terms, for example, by typing terms and / or speaking into the microphone 270, to update the ground truth 586. The custom vocabulary 584 is a dictionary of key-value pairs.The “key” is the output string 525 of the speech recognition algorithm 520 and the “value” is the corresponding text from the ground truth 586.
[0153] Referring to Figure 6, a block diagram of an exemplary embodiment of a method 600 for performing image analysis during an endoscopy procedure using the system of Figure 2 is shown. The method 600 may be implemented by the CPU 255C and GPU 255G of the EIA system 242, allowing the EIA system 242 to continually adapt to the user and generate effective image analysis output for each OOI. Some steps of the method 600 may be performed using the CPU 255C and GPU 255G of the microcomputer 255 and the main image processor 215 of the endoscopy platform 210.
[0154] Method 600 begins with the initiation of an endoscopy procedure at 610. The initiation of an endoscopy may begin when the endoscopy device is powered on (or started) at 620. In parallel with this, microphone 270 and the AI platform (e.g., EIA system 242) are turned on at 650. Method 600 includes two branches that run in parallel with each other.
[0155] Following the branch of the method 600 beginning at 620, the processor 215 of the endoscopy platform 210 receives a signal that an operational endoscopy device 220 is present.
[0156] At 622, the processor 215 performs a diagnostic check to determine that an operational endoscopy device 220 is properly connected to the processor 210. Step 622 may be referred to as an endoscopy quality assurance (QA) step. The processor 215 sends a confirmation to the monitor 240 to indicate to the user that the QA step was successful or unsuccessful. If the processor 215 sends an error message to the monitor 240, the user must resolve the error before continuing with the procedure.
[0157] Referring to the other branch of method 600 beginning with step 650, after step 650 is executed, method 600 proceeds to step 652 where EIA system 242 performs a diagnostic check to determine that microcomputer 255 and microphone 270 are properly connected, which may be referred to as an AI platform quality assurance (QA) step. The AI platform QA step includes checking the algorithms. If there are errors, EIA system 252 generates an error message that is displayed on monitor 265 to notify the user that one or more issues associated with the error message must be resolved before continuing with the video stream capture execution.
[0158] If the QA step is performed successfully, the method 600 proceeds to step 654, where the EIA system 242 captures an input video stream including images provided by the endoscopy device 220. Image data from the input video stream may be received by the input module 142 for processing by image analysis algorithms. When the input video stream is being received, or a series of images for other medical imaging modality applications are being input, the microcomputer 255 may activate the LED light 341 to indicate that the EIA system 242 is operating (e.g., by showing a steady green light).
[0159] Referring back to the left branch again, at 624, which is the start of the endoscopy procedure, the processor 215 checks the patient information by asking the user to input the patient information (e.g., via the input module 144) or by directly downloading the patient information from the medical chart. The patient information may consist of the patient demographic information, the user (e.g., of the EIA system 242), the procedure type, and any unique identifier. The microcomputer 255 inputs a specific frame / image from the start of the endoscopy procedure. The specific image may be used by the EIA system 242 to generate a second output. The second output is used in a DICOM report that includes the specific image from the start of the endoscopy procedure, which may be used to capture the patient information for the DICOM report. Alternatively or additionally, medical diagnosis (e.g., endoscopy diagnosis) information data may be captured. To ensure privacy, the server 120 may prevent the patient information from being stored in any other data file.
[0160] After both the initiation of the endoscopy procedure and the capture of the video stream by the EIA system 242 at 626, the EIA system 242 then waits to receive an input signal to begin recording audio. This marks the start of process A 632 and process B 660. Upon receiving the input start signal 421, the EIA system 242 begins process A 632 and process B 660.
[0161] At 628, the EIA system 242 receives user input as speech in an input audio signal. The EIA system 242 continues recording the input audio signal until it receives an end of input signal 424.
[0162] At 630, after receiving the input end signal 424, the EIA system 242 ends recording the input audio signal, which indicates the end of process A632. However, the EIA system 242 may repeat process A632 at later times when audio start and stop commands are provided until the endoscopic procedure is completed and the endoscopy device 220 is powered off.
[0163] Once the endoscopic procedure is complete, method 600 proceeds to 634 where the processor 215 receives a signal that the endoscopic procedure is complete.
[0164] At 638, the processor 215 powers off the endoscopy platform 210. Alternatively or additionally, the EIA system 242 receives a signal indicating that the endoscopy platform 210 has been powered off.
[0165] Referring again to the right branch of method 600, process B 660 executes in parallel with process A 632 and includes all steps of process A 632, performing speech recognition and speech-to-text algorithms to generate annotation data at 656 and match the image to the annotation data at 658. The EIA system 242 may repeatedly perform process B 660 until an input signal including a user command to power off the endoscopy device is received by the EIA system 242.
[0166] At 656, the EIA system 242 initiates the speech recognition and speech-to-text conversion processes to generate annotation data. This may be done using the speech recognition algorithm 410, the speech-to-text conversion algorithm 520, the term correction algorithm 530, and the real-time annotation process 436.
[0167] At 658, the EIA system 242 matches the images with the annotations. This can be done using the matching algorithm 430.
[0168] At 662, the real-time annotation process 436 receives a command signal from a user to prepare a data file for output generation and storage. For example, image data, audio signal data, annotated images, and / or a series of images (such as a video clip) may be marked for storage. An output file may be generated using the annotated images in a particular data format, such as the DICOM format.
[0169] At 664, the EIA system 242 sends a message that the output file is ready, which may occur a set time (e.g., 20 seconds or less) after the EIA system 242 receives a data file ready command signal from the user. At this point, the output file may be displayed on a monitor, stored in a storage element, and / or transmitted to a remote device. Reports may also be printed.
[0170] At 666, the EIA system 242 powers down the operational AI platform and microphones at the end of the procedure. Alternatively, the EIA system 242 receives a signal indicating that the AI platform and microphones have been powered down. The EIA system 242 can be powered down by the user entering a software command that initiates a system shutdown and disables power from the power supply unit 136.
[0171] Referring to FIG. 7, a diagram of an example embodiment of an image analysis training algorithm 700 is shown. The encoder 720 receives an input X 790 (e.g., via the input module 144). The input X 790 is at least one image from a sequence of images provided by a medical imaging device (e.g., the endoscope 220). The encoder 720 compresses the input X 790 into a feature vector 730 using at least one convolutional neural network (CNN). The feature vector 730 may be an n-dimensional vector or matrix of numerical features that describe the input X 790 for the purposes of pattern recognition. The encoder 720 may perform the compression by allowing only the maximum value (i.e., max pooling) of a 2x2 patch to propagate towards the feature layer of the CNN at multiple points.
[0172] The feature vector 730 is then input to a decoder 770. The decoder 770 reconstructs a high-resolution image 780 from the low-resolution feature vector 730.
[0173] The classifier 740 maps the feature vector 730 to a distribution over target classes 750. For labeled (i.e., annotated with a category or classification) input images, the classifier 740 may be trained together with the encoder 720 and the decoder 770. This can be advantageous because it facilitates the encoder 720 and the decoder 770 to learn features that are useful for classification while jointly learning how to classify those features.
[0174] The classifier 740 may be constructed from two convolutional layers that reduce the channel dimension by half and then to one, followed by a fully connected (FC) linear layer that projects the hidden state into a real-valued vector with a size equal to the number of categories. The result is mapped using a mapping function, e.g., softmax, to represent the category distribution across the target classes. Between the convolutional layers, a Swish activation function (e.g., x*sigmoid(x)) can be used. The output of the classifier 740 provides the probability that the model would assign to each category given the OOI in the input image.
[0175] The encoder 720, decoder 770, and classifier 740 enable the EIA system 242 to perform semi-supervised training. Semi-supervised training is advantageous because it allows the EIA system 242 to build image analysis algorithms with a small number of labeled training data sets.
[0176] Given an image Xj, we define the autoencoder loss (LAE) for maximum likelihood (ML) learning of the parameters according to: LAE(xj)=(p(x=xj)log p(x=xj|h=Eθ(x))+(1-p(x=xj))log(1-p(x=xj|h=Eθ(x)))) where p(x=xj) is for the input image and p(x=xj|h=Eθ(x)) is for the reconstructed image (i.e., the probability that the reconstructed image from the decoder is the same as the input image), both interpreted as Bernoulli distributions over the channel-wise and pixel-wise representations of the color image. The Bernoulli distribution provides a measure of consistency between the input and reconstructed images. Each image pixel contains three channels (red, green, and blue). Each channel holds a real value in the range [0,...,1] that represents the intensity of the corresponding color, with 0 representing no intensity and 1 representing maximum intensity. Since the range is [0,...,1], the intensity value can be used as a probability for LAE(xj), which is the binary cross entropy (BCE) between the model and the data distribution of the samples. Minimizing the LAE using stochastic gradient descent involves a learning procedure. Minimizing the LAE encourages learning a feature vector that captures the information in the image. We do this using only the encoded feature vector to reconstruct the input image. In other words, minimizing the LAE encourages the learning of informative features that can be used for classification when labels are available. The LAE can be trained in an unsupervised manner, which means that the EIA system 242 does not require a labeled training dataset to be built.
[0177] Given a labeled image (xi, yi), the EIA system 242 defines a classifier loss (LCLF) for maximum likelihood (ML) learning of parameters according to: LCLF(xi, yi)=log p(y=yi|h=Eθ(x)) Where p(y = yi|h = Eθ(x)) is the probability of category yi, and LCLF(xi, yi) is the discrete cross-entropy (CE) between the model and the categorical distribution of the samples. LCLF encourages the learned features to be useful for classification and provides the probability for each category when an input image used in the analysis pipeline is given. LCLF is trained in a supervised manner, which means that the server 120 needs a labeled training dataset for construction. LCLF can be considered as a loss that quantifies the consistency between the prediction from the model and the ground truth labels provided in the training data. If LCLF is a standard cross-entropy loss, this would mean that the model uses the log softmax probability for the correct class.
[0178] The semi-supervised loss for the dataset D is defined as follows. LCLF(D) = λ1N(ΣiLCLF(xi, yi)) + 1M(ΣjLAE(xj)) Where λ controls the weight of the classification component, N is the number of labeled images, M is the number of unlabeled images, and generally, N << M (M is much larger than N). The semi-supervised loss enables learning useful features from a large number of unlabeled images and learning a powerful classifier (more precisely, one that can be trained more accurately and quickly) from a small number of labeled images. The weight can force the learning of features more suitable for classification at the expense of a worse reconstruction. An appropriate value for λ could be, for example, 10,000. The weight may provide a way to form a single loss as a linear combination of the loss of the autoencoder and the loss of the classifier, which can be determined using some form of cross-validation.
[0179] A series of medical images (e.g., an endoscopy video stream) may be analyzed for object detection to determine OOIs in the images using different algorithms. Multiple open source datasets and / or dedicated medical diagnostic procedure datasets may be used to train the algorithms. For example, for colonoscopy, the dataset includes images classified with different classes of OOIs: healthy, unhealthy, and unlabeled colonoscopy images, all examples of which are shown in Figures 9, 10, and 11. The algorithms (e.g., image analysis algorithms, object detection algorithms) look at the morphological properties of the tissue to classify it, and if the tissue cannot be clearly identified, it can be assigned to an "out of focus tissue" (or blurry) class. Thus, images of the out of focus tissue class are poor and / or low quality images such that object detection and / or classification cannot be performed accurately. For other medical applications and / or imaging modalities, other classes may be used based on the object of interest to be located and classified.
[0180] The system 100, or the EIA system 242 (in the context of endoscopy), may combine supervised 710 and unsupervised 760 methods during training of the machine learning methods used to classify OOIs. This panel of algorithms (e.g., two or more algorithms working together) may use a U-net architecture (e.g., as shown in FIG. 8A or FIG. 8B). Although the training is described in the context of gastrointestinal endoscopy, it should be understood that the training may be performed for other types of endoscopy, other types of medical applications, and / or other imaging modalities by using a training set of images having various objects desired to be detected and classified.
[0181] An annotated image dataset 790 (e.g., an annotated endoscopic image dataset) can also be used to train the supervised method 710. In this case, an encoder (E) 720 projects a given image into a latent feature space and constructs an algorithm / feature vector 730 that allows a classifier (C) 740 to map features to a distribution over target classes and discriminate multiple classes based on morphological properties of diseases / tissues in the training images 750.
[0182] By using the unlabeled image, the auxiliary decoder (G) 770 maps the features to a distribution on the image using a reconstruction method 780. To implement the reconstruction method 780 in the U-net architecture, the image can be decomposed into pixels and an initial pressure distribution obtained from the detected signal using an image reconstruction algorithm (e.g., as shown diagrammatically on the right side of the U-net architecture). The unsupervised method 760 can add value by allowing the features to use a smaller number of annotated images per class.
[0183] Referring to FIG. 8A, a block diagram of a first exemplary embodiment of a U-net architecture 800 that may be used by the image analysis algorithm (which may be stored in the program 142) is shown.
[0184] The convolution block 830 receives an input image 810 (e.g., via the input module 144). The convolution block 830 consists of a convolution layer, an activation layer, and a pooling layer (e.g., in series). The convolution block 830 generates features XXX. An example of this is shown for the first convolution block 830 in the top left of FIG. 8A.
[0185] The deconvolution block receives the features generated by one of the convolution blocks and the previous deconvolution block. For example, the deconvolution block 820 in the top right of FIG. 8A receives the features XXX generated by the convolution block 830, as well as the output of the preceding (i.e., next lower) deconvolution block. The deconvolution block 840 consists of a convolution layer, a transposed convolution layer, and an activation layer. The deconvolution block 840 generates output features 820. The output features 820 can be, for example, an array of numerical values. The deconvolution block 840 adds information to the provided features, allowing the reconstruction of an image given the corresponding features.
[0186] The classifier block 850 consists of a convolutional layer, an activation layer, and a fully connected layer. The classifier block 850 receives the features XXX generated by the last convolutional block in the series of convolutional blocks. The classifier block 850 generates classes of one or more objects in the image being analyzed. For example, each image or region of an image may be labeled with one or more classes, such as "is a polyp" or "is not a polyp" in the GI endoscopy example, although other classes may be used for other types of endoscopic procedures, medical procedures, and / or imaging modalities.
[0187] Referring to FIG. 8B, a block diagram of a second exemplary embodiment of a U-net architecture 860 that may be used by the image analysis algorithm (which may be stored in the program 142) is shown.
[0188] At 864, the first convolutional layer receives an input image (e.g., via the input module 144). The various convolutional layers at this level linearly blend the input image and only the linear part of the convolution is used (e.g., in the case of a 3x3 convolution, one pixel order is lost) to learn concise features (i.e., a representation) of the input image. This can be done by a 3x3 convolution, ReLu operation. After each subsequent 3x3 convolution, ReLu operation, the resolution of the layer decreases. For example, the resolution of the layer can go from 572x572 (having 3 channels) to 570x570 (having 64 channels) to 568x568 (having 64 channels). At the final layer, a max pooling 2x2 operation can be applied to generate a convolved layer for the next convolutional layer (868). In addition, a copy-and-crop operation can be applied to the convolved layer for deconvolution (896).
[0189] At 868, the subsequent convolutional layer receives the convolved layer from the convolutional layer above (from 864). Only the linear part of the convolution is used so that the various layers linearly blend the input images and learn concise features (i.e., a representation) of the input images. This is done by a 3x3 convolution, ReLu operation. After each subsequent 3x3 convolution, ReLu operation, the resolution of the layer decreases. For example, the resolution of the layer can go from 284x284 (with 64 channels) to 282x282 (with 128 channels) to 280x280 (with 128 channels). At the final layer, a max pooling 2x2 operation is applied to generate the convolved layer for the next convolutional layer (872). In addition, a copy and crop operation is applied to the convolved layer for deconvolution (892).
[0190] At 872, another subsequent convolutional layer receives the convolved layers from the previous convolutional layer above (from 868). Only the linear part of the convolution is used so that the various layers at this level linearly blend the input images and learn concise features (i.e., a representation) of the input images. This is done by a 3x3 convolution, ReLu operation. After each subsequent 3x3 convolution, ReLu operation, the resolution of the layer decreases. For example, the resolution of the layer can go from 140x140 (with 128 channels) to 138x138 (with 256 channels) to 136x136 (with 256 channels). At the final layer, a max pooling 2x2 operation is applied to generate the convolved layers for the next convolutional layer (876). In addition, a copy and crop operation is applied to the convolved layers for deconvolution (888).
[0191] At 876, the convolutional layer receives the convolved layer from the previous convolutional layer above (from 872). Only the linear part of the convolution is used so that the various layers linearly blend the input image and learn concise features (i.e., a representation) of the input image. This is done by a 3x3 convolution, ReLu operation. After each subsequent 3x3 convolution, ReLu operation, the resolution of the layer decreases. For example, the resolution of the layer can go from 68x68 (with 256 channels) to 66x66 (with 512 channels) to 64x64 (with 512 channels). At the final layer, a max pooling 2x2 operation is applied to generate the convolved layer for the next convolutional layer (880). In addition, a copy and crop operation is applied to the convolved layer for deconvolution (884).
[0192] At 880, the convolutional layer receives features from the convolutional layer above (from 876). Only the linear part of the convolution is used so that the various layers linearly blend the input image and learn concise features (i.e., a representation) of the input image. This is done by a 3x3 convolution, ReLu operation. After each subsequent 3x3 convolution, ReLu operation, the resolution of the layer decreases. For example, the resolution of a layer can go from 32x32 (with 512 channels) to 30x30 (with 1024 channels) to 28x28 (with 512 channels). At the final layer, an ascending convolution pooling 2x2 operation is applied to the convolved layer for deconvolution (884).
[0193] The decoder 770 then performs deconvolution at 884, 888, 892, and 896. The decoder 770 reconstructs an image from the features by adding dimension to the features using a series of linear transformations (upward convolutions) that map a single dimension to a 2x2 patch. The reconstructed image is represented using RGB channels (red, green, blue) for each pixel, with each value in the range [0,...,1]. A value of 0 means no intensity and a value of 1 means maximum intensity. The reconstructed image is identical in dimensions and format to the input image.
[0194] At 884, the deconvolution layer receives features from the lower convolution layer (from 880) and the cropped image from the previous convolution (from 876). These steps build a high-resolution segmentation map by a sequence of ascending convolutions and concatenation of high-resolution features from the shrinking pass. This ascending convolution uses a learned kernel to map each feature vector to an output window of 2X2 pixels, followed by a nonlinear activation function. For example, the layer resolution can go from 56x56 (with 1024 channels) to 54x54 (with 512 channels) to 52x52 (with 512 channels). At the final layer, an ascending convolution pooling 2x2 operation is applied to the deconvolved layer for the next deconvolution layer (888).
[0195] At 888, the deconvolution layer receives the deconvolved layer from the lower deconvolution layer (from 884) and the cropped image from the previous convolution (from 872). These steps build a high-resolution segmentation map by a sequence of ascending convolutions and concatenation with high-resolution features from the shrinking pass. This ascending convolution uses a learned kernel to map each feature vector to an output window of 2X2 pixels, followed by a nonlinear activation function. For example, the layer resolution can go from 104x104 (with 512 channels) to 102x102 (with 256 channels) to 100x100 (with 256 channels). At the final layer, an ascending convolution pooling 2x2 operation is applied to the deconvolved layer for the next deconvolution layer (892).
[0196] At 892, the deconvolution layer receives the deconvolved layer from the deconvolution layer below (from 888) and the cropped image from the previous convolution (from 868). These steps build a high-resolution segmentation map by a sequence of ascending convolutions and concatenation with high-resolution features from the shrinking pass. This ascending convolution uses a learned kernel to map each feature vector to an output window of 2X2 pixels, followed by a nonlinear activation function. For example, the layer resolution can go from 200x200 (with 256 channels) to 198x198 (with 128 channels) to 196x196 (with 128 channels). At the final layer, an ascending convolution pooling 2x2 operation is applied to the deconvolved layer for the next deconvolution layer (896).
[0197] At 896, the deconvolution layer receives the deconvolved layer (from 892) from the deconvolution layer below (e.g., via input module 144) and the cropped image from the previous convolution (from 864). These steps build a high-resolution segmentation map by a sequence of upward convolutions and concatenation with high-resolution features from the contraction pass. This upward convolution uses a learned kernel to map each feature vector to an output window of 2X2 pixels, followed by a nonlinear activation function. For example, the resolution of the layer can go from 392x392 (with 128 channels) to 390x390 (with 64 channels) to 388x388 (with 64 channels). At the last layer, a convolution 1x1 operation is applied to the deconvolved layer, the reconstructed image (898).
[0198] At 898, the reconstructed image is output along with the features obtained from the convolution. The reconstructed image is identical in size and format to the input image. For example, the resolution of the reconstructed image may be 572x572 (with 3 channels).
[0199] Although FIG. 8B shows a U-net architecture with three convolutional layers, the U-net architecture may be structured so that there are more convolutional layers (e.g., for images of different sizes, or for different depths of analysis).
[0200] 9, an example of an endoscopy image with healthy morphological characteristics 900 is shown. The endoscopy image with healthy morphological characteristics 900 consists of, from left to right, a normal cecum, a normal pylorus, and a normal Z-line. These colonoscopy images with healthy morphological characteristics 900 are obtained from the Kvasir dataset. The endoscopy images with healthy morphological characteristics 900 may be used by the EIA system 242 to train image analysis algorithms in a supervised or semi-supervised manner.
[0201] Referring to FIG. 10, an example of an endoscopic image with unhealthy morphological characteristics 1000 is shown. The endoscopic image with unhealthy morphological characteristics 1000 consists of, from left to right, a stained protruding polyp, a stained resection margin, esophagitis, polyps, and ulcerative colitis. These endoscopic images with unhealthy morphological characteristics 1000 are obtained from the Kvasir dataset. The endoscopic images with unhealthy morphological characteristics 1000 can be used by the EIA system 242 to train image analysis algorithms in a supervised or semi-supervised manner. Alternatively or additionally, medical images with healthy or unhealthy morphological characteristics can be obtained from other devices / platforms, such as, but not limited to, CT scanners, ultrasound devices, MRI scanners, X-ray devices, nuclear medicine imaging devices, histology imaging devices, etc., to adapt the methods and systems described herein for use in other types of medical applications.
[0202] 11, an example of unlabeled video frame images from a dedicated dataset 1100 is shown. The unlabeled video frame images from the dedicated dataset 1100 include both healthy and unhealthy tissue. The unlabeled video frame images from the dedicated dataset 1100 are used by the EIA system 242 to train image analysis algorithms in a semi-supervised manner.
[0203] Referring to FIG. 12, a block diagram of an exemplary embodiment of a report generation process 1200 is shown. The report may be generated in a particular format, such as, for example, a DICOM report format. It should be noted that while the process 1200 is described as being performed by the EIA system 242, this is for illustrative purposes, and it should be understood that the system 100 or another suitable processing system may be used. More generally, however, the EIA system 242 may be considered an alternative exemplary embodiment of the system 100 when used for other medical imaging applications and imaging modalities. In such cases, any reference to endoscopy, endoscopy, or endoscopic images may be replaced with other medical imaging procedures, imaging modalities, imaging devices, or medical images, respectively, such as the examples shown in Table 1, and the process 1200 may be used with these other medical imaging procedures, imaging modalities, imaging devices, and medical images.
[0204] At 1210, the EIA system 242 loads a patient demographics frame. The patient demographics frame may consist of patient identifiers such as the name, date of birth, gender, medical number, etc. of the patient undergoing the endoscopic procedure. The EIA system 242 can display the patient demographics frame on the endoscope monitor 240. The EIA system 242 can collect patient data using still images from the endoscope monitor 240.
[0205] At 1220, the EIA system 242 executes an optical character recognition algorithm, which may be stored in the program 142. The EIA system 242 reads the patient demographic frame using the optical character recognition algorithm. The optical character recognition algorithm may use a set of codes that can identify text characters in specific locations of the image. In particular, the optical character recognition algorithm can see the borders of the image that show the patient information.
[0206] At 1230, the EIA system 242 extracts the read patient information and uses the information in generating a report.
[0207] At 1240, the EIA system 242 loads key images (i.e., video frames or images from a sequence of images) and / or video clips along with annotations (e.g., from database 150), if applicable, for report generation. Key frames may be those identified by an image and annotation data matching algorithm.
[0208] At 1250, the EIA system 242 generates a report which may be output to a display via the output module 148, for example, and / or transmitted to an electronic health record system or electronic medical record system via a network unit.
[0209] 13, a block diagram of an exemplary embodiment of a method 1300 for processing a series of images using image processing and annotation algorithms that may be used by the EIA system 242 is shown. While the method 1300 is described as being performed by the EIA system 242, it should be noted that this is for illustrative purposes and it should be understood that the system 100 or another suitable processing system may be used. More generally, however, the EIA system 242 may be considered an alternative exemplary embodiment of the system 100 when used for other medical imaging applications and imaging modalities. In such cases, any reference to endoscopy, endoscopy, or endoscopic images may be replaced with other medical imaging procedures, imaging modalities, imaging devices, or medical images, respectively, such as the examples shown in Table 1, and the process 1300 may be used with these other medical imaging procedures, imaging modalities, imaging devices, and medical images.
[0210] At 1310, the EIA system 242 receives the sequence of images 1304 and crops an image from the sequence of images, such as endoscopic images from an input video stream. For example, the cropping can be performed using an image processing library such as OpenCV (an open source library). The EIA system 242 can input the raw geometry and x min, x max, y min, and y max. OpenCV can then generate the cropped image.
[0211] At 1320, the EIA system 242 detects one or more objects in the cropped endoscopic image. Once one or more objects are detected, their locations are determined, and then a classification and confidence score for each object is determined. This may be done using a trained object detection algorithm. The architecture of this object detection algorithm may be YOLOv4. The object detection algorithm may be trained, for example, using a public database or Darknet.
[0212] Acts 1310 and 1320 may be repeated for several images from the image series 1305.
[0213] At 1330, the EIA system 242 receives a signal (560, 562, 564) to initiate annotation for one or more images from the image series 1305. The EIA system 242 then performs speech recognition, speech-to-text conversion, and generates annotation data 1335, which may be performed as described above.
[0214] The method 1300 then proceeds to 1340 where the annotation data is added to the matching image to create an annotated image. Again, this may be repeated for multiple images from the image series 1305 based on commands and comments provided by the user. The annotated images may be output to an output video stream 1345.
[0215] Table 2 below shows the results of classifying tissues using supervised and unsupervised methods. [Table 2]
[0216] Referring now to FIG. 14, a chart 1400 of training results for YOLOv4 is shown, depicting the accuracy of the speech recognition algorithm used by the EIA system 242, showing the rate of positive speech recognition results (P) versus the true positive (TP) value. The x-axis of the chart represents the number of training iterations (one iteration is one mini-batch of images consisting of 32 images), and the y-axis represents the TP detection rate for polyp detection using the validation set. The chart 1400 shows that the TP rate starts at 0.826 at iteration 500 and increases to 0.922 after iteration 1000. Over the course of iterations 1000-3000, the TP rate generally remains at a level of approximately 0.92-0.93. TP can reach 0.93 after 3000 iterations.
[0217] The accuracy of classification provided by an AI algorithm was selected as the analytical metric to evaluate the accuracy of object detection or speech recognition. The term false positive (FP) refers to the error of a machine learning model predicting a "true" value despite the actual observed value being "false". On the other hand, a false negative (FN) indicates the error of a machine learning model outputting a "false" predicted value despite the actual observed value being "true". FP is a major factor that reduces the reliability of software classification platforms in the medical field when using machine learning models. As a result, the trained object and speech recognition algorithms described herein have been validated using metrics such as accuracy.
[0218] 15, there is shown a block diagram of an exemplary embodiment of a speech recognition algorithm 1500. The speech recognition algorithm 1500 may be implemented using one or more of the program 142, the prediction engine 152, and the machine learning model 146. It should be understood that in other embodiments, the speech recognition algorithm 1500 may be used with other medical imaging procedures, imaging modalities, imaging devices, or medical images, such as the examples shown in Table 1.
[0219] The speech recognition algorithm 1500 receives raw audio data 1510 acquired via the microphone 270. The speech recognition algorithm 1500 includes a convolutional neural network block 1520 and a converter block 1530. The convolutional neural network block 1520 receives the raw audio data 1510. The convolutional neural network block 1520 extracts features from the raw audio data 1510 to generate a feature vector. Each convolutional neural network in the convolutional neural network block 1520 may be identical, including the weights used. The number of convolutional neural network blocks 1520 in the speech recognition algorithm 1500 may depend on the length of the raw audio data 1510.
[0220] The converter block 1530 receives the feature vector from the convolutional neural network block 1520. The converter block 1530 generates a character corresponding to the user input by extracting features from the feature vector.
[0221] 16, a block diagram of an example embodiment of a data flow 1600 for an object detection algorithm 1620 that may be used by an image analysis algorithm is shown. The object detection algorithm 1620 may be implemented using one or more of the program 142, the prediction engine 152, and the machine learning model 146. It should be understood that in other embodiments, the object detection algorithm 1620 may be used with other medical imaging procedures, imaging modalities, imaging devices, or medical images, such as the examples shown in Table 1.
[0222] The object detection algorithm 1620 receives the processed image 1610. The processed image 1610 may be a cropped and resized version of the original image.
[0223] The processed image 1610 is input into CPSDarknet53 1630, a convolutional neural network capable of extracting features from the processed image 1610.
[0224] The output of the CSPDarknet53 1630 is provided to a spatial pyramid pooling operator 1640 and a path aggregation network 1650.
[0225] The spatial pyramid pooling operator 1640 is a pooling layer that can remove the fixed size constraint of the CSPDarknet53 1630. The output of the spatial pyramid pooling operator 1640 is provided to a path aggregation network 1650.
[0226] The path aggregation network 1650 processes the output from the CSPDarknet53 1630 and the spatial pyramid pooling operator 1640 by extracting features of different depths from the output of the CSPDarknet53 1630. The path aggregation network 1650 outputs to the Yolo head 1660.
[0227] Yolo head 1660 predicts and generates a class 1670 of the OOI, a bounding box 1680, and a confidence score 1690. The class 1670 is the classification of the OOI. Figures 9-11 show various examples of images containing classified objects. For example, the class 1670 could be a polyp. However, if the classification 1690 is not determined with a high enough confidence score, the image may be classified as suspicious.
[0228] 17, an exemplary embodiment of a report 1700 including annotated images generated in accordance with the teachings herein is shown. The report 1700 includes various information collected during image and audio capture occurring during a medical procedure (e.g., a medical diagnostic procedure such as an endoscopy procedure) in accordance with the teachings herein. The report 1700 generally includes various elements including, but not limited to, (a) patient data (i.e., name, date of birth, etc.), (b) information regarding the medical procedure (e.g., date of procedure, if any biopsy was performed, if any treatment was performed, etc.), (c) a description field to provide a description of the procedure and any findings, (d) one or more annotated images, and (e) a recommendation field that includes the text of any recommendations for further treatment / follow-up for the patient. In other embodiments, some of the elements other than the annotated image may be optional. In some cases, the annotated image may be included in the report along with the bounding box, annotation data, and confidence score. In other cases, the bounding box, annotation data, and / or confidence score may not be included in the report.
[0229] In at least one embodiment described herein, the EIA system 242 or system 100 may be configured to perform several functions. For example, a given image may be displayed in which OOI is detected and classified and the classification is included in the given image. A user may then provide a comment by speech if they may disagree with the automatic classification provided by the EIA system 242. In this case, the user's comment is converted into a text string that matches the given image. Annotation data is generated using the text string, and the annotation data is linked (e.g., overlaid or superimposed) to the given image.
[0230] In at least one embodiment, a given image may be displayed in which an OOI is detected and automatically classified, and the automatic classification is included in the given image. A user may want to view the given image and double-check whether the automatic classification is correct. In such a case, the user may provide a command to display other images having OOIs of the same classification as the automatic classification. The user's speech may include this command. Thus, when speech-to-text conversion is performed, the text may be inspected to determine whether it includes a command such as a request for a reference image having OOIs classified with the same classification as the at least one OOI. The EIA system 242 or a processor of the system 100 may then retrieve the reference image from the data store, display the reference image, and receive a subsequent input from the user via speech that confirms or rejects the automatic classification of the at least one OOI. Annotation data may be generated and overlaid on the given image based on this subsequent input.
[0231] In at least one embodiment described herein, the EIA system 242 or system 100 may be configured to perform several functions. For example, a given image may be displayed in which OOI is detected and classified and the classification is included in the given image. A user may then provide a comment by speech if they may disagree with the automatic classification provided by the EIA system 242. In this case, the user's comment is converted into a text string that matches the given image. Annotation data is generated using the text string, and the annotation data is linked (e.g., overlaid or superimposed) to the given image.
[0232] In at least one embodiment described herein, the EIA system 242 or system 100 may be configured to perform several functions. For example, a given image may be displayed if an OOI is detected but the confidence score associated with the classification is not sufficient to reliably classify the OOI. In such a case, the given image may be displayed and indicated as suspicious, in which case input may be received from a user indicating a user classification for at least one image having pending OOI. The given image may then be annotated with the user classification.
[0233] In at least one embodiment described herein, the EIA system 242 or system 100 may be configured to overlay a timestamp when generating an annotated image, the timestamp indicating the time the image was originally acquired by a medical imaging device (e.g., endoscope 220).
[0234] While applicants' teachings have been combined with various embodiments for illustrative purposes, it is not intended that applicants' teachings described herein be limited to such embodiments, as the embodiments described herein are intended to be examples. Rather, applicants' teachings as described and illustrated herein encompass various alternatives, modifications, and equivalents without departing from the embodiments described herein, the general scope of which is defined in the appended claims.
Claims
1. A system for analyzing medical image data for a medical procedure, the system comprising: A non-transitory computer-readable medium storing program instructions for analyzing medical image data for the medical procedure; At least one processor, wherein when the at least one processor executes the program instructions, Receiving at least one image from a series of images; Determining when at least one object of interest (OOI) exists in the at least one image, and when the at least one OOI exists, determining a classification of the at least one OOI, both determinations being performed using at least one machine learning model; During the medical procedure, using a bounding box to display the at least one image and any determined OOIs to a user on a display; Receiving an input audio signal including speech from the user during the medical procedure and recognizing the speech; When the speech is recognized as a comment on the at least one image during the medical procedure, using a speech-to-text conversion algorithm and a term correction algorithm to convert the speech into at least one text string; Collating the at least one text string with at least one image provided with the speech from the user; During the medical procedure, generating at least one annotated image in which the at least one text string is linked to at least one corresponding image; At least one processor configured to perform the above; A system comprising the above.
2. The system according to claim 1, wherein the at least one processor is further configured to display the at least one reference image when the speech is recognized as a request for at least one reference image having an OOI classified in the same classification as the at least one OOI during the medical procedure, and to receive from the user an input to confirm or reject the classification for the at least one OOI in order to update the at least one machine learning model.
3. The system according to claim 1 or claim 2, wherein the at least one processor is further configured to receive from the user an input indicating a user classification for at least one image having the undetermined OOI when the at least one OOI is classified as being suspicious.
4. The system according to any one of claims 1 to 2, wherein the at least one processor is further configured to automatically generate a report including the at least one annotated image.
5. For a given OOI within a given image, the at least one processor is identifying the bounding box coordinates of the bounding box, wherein the bounding box is associated with the given OOI within the given image, calculating a confidence score based on a probability distribution of classifications for the given OOI, overlaying the bounding box on the at least one image at the bounding box coordinates when the confidence score is higher than a confidence threshold, overlaying custom vocabulary on the at least one image when receiving confirmation from the user during the medical procedure, The system according to any one of claims 1 to 2, further configured to perform.
6. The at least one processor is applying a convolutional neural network (CNN) to the OOI by performing convolutional operations, activation operations, and pooling operations to generate a matrix, generating a feature vector by processing the matrix using the convolutional operations, activation operations, and pooling operations, performing a classification for the OOI based on the feature vector, The system according to any one of claims 1 to 2, configured to determine the classification of the OOI thereby.
7. The system according to any one of claims 1 to 2, wherein the at least one processor is further configured to overlay at least one timestamp of an in-procedure event and a timestamped document on the corresponding at least one image when generating the at least one annotated image during the medical procedure.
8. The system according to claim 5, wherein the at least one processor is further configured to display a confidence score of the at least one image in real time on a display during the medical procedure.
9. The at least one processor is to start receiving an audio stream of the input voice from the user upon detection of a first user action, wherein the first user action is to temporarily pause the display of the series of images, to take a snapshot of a given image within the series of images, or to provide a first voice command and is configured to receive the input voice during the medical procedure by to end the reception of the audio stream upon detection of a second user action, wherein the second user action is to remain silent for a predetermined length, to press a designated button, or to provide a last voice command and is configured to receive the input voice during the medical procedure by The system according to any one of claims 1 to 2.
10. The system according to any one of claims 1 to 2, wherein the at least one processor is further configured to store the series of images when receiving the input voice during the medical procedure, thereby designating the at least one image to receive annotation data for generating at least one corresponding annotated image.
11. The at least one processor is to capture a set of patient information data generated during the medical procedure to be added to the report, to load a subset of the series of images including the at least one annotated image or at least one OOI identified by the bounding box, to combine the set of patient information data and the subset of the series of images including the at least one annotated image into the report, The system according to claim 4, which is further configured to generate a report of the medical procedure by
12. The at least one processor is to apply an encoder to at least one training image to generate at least one feature vector for a training OOI within the at least one training image, Selecting a class of the training OOI by applying the at least one feature vector to the at least one machine learning model; Reconstructing a labeled training image using a decoder by associating the at least one feature vector with the at least one training image and a selected class for training the at least one machine learning model; The system according to any one of claims 1 to 2, further configured to perform training of the at least one machine learning model thereby.
13. The system according to claim 12, wherein the class is a healthy tissue class, an unhealthy tissue class, a suspicious tissue class, or an out-of-focus tissue class.
14. The at least one processor is Training the at least one machine learning model using a training data set including labeled training images, unlabeled training images, or a mixture of labeled and unlabeled training images, wherein the images include examples categorized by healthy tissue, unhealthy tissue, suspicious tissue, and out-of-focus tissue The system according to claim 12, further configured to perform.
15. The system according to claim 12, wherein the at least one processor is further configured to train the at least one machine learning model using supervised learning, unsupervised learning, or semi-supervised learning.
16. The system according to claim 14, wherein the training data set further includes subcategories for each of the unhealthy tissue and the suspicious tissue.
17. The at least one processor is Receiving a training image as an input to the encoder; Projecting the training image into features that are part of a feature space using the encoder; Mapping the features to a set of target classes using a classifier; Identifying morphological characteristics of the training image to generate a new training data set, wherein the new training data set has data linking parameters to the training image Based on the morphological characteristics, determining whether there is one or more mapped classes or no mapped classes; The system according to claim 12, further configured to create the at least one machine learning model thereby.
18. The at least one processor Receiving one or more of the features as input to the decoder; Using a deconvolutional neural network to map one of the features to an unlabeled dataset; Reconstructing a new training image from one of the features using the decoder to train the at least one machine learning model The system according to claim 17, further configured to determine a classification for the at least one OOI thereby.
19. The at least one processor is further configured to train the speech-text conversion algorithm using the utterance dataset to identify a match with the ground truth text by comparing new speech data with the utterance dataset, the utterance dataset including the ground truth text and the speech data of the ground truth text, the system according to any one of claims 1 to 2.
20. The system according to any one of claims 1 to 2, wherein the speech-text conversion algorithm maps the at least one OOI to one of a plurality of OOI medical terms.
21. The system according to any one of claims 1 to 2, wherein the medical image data is obtained from one or more endoscopic procedures, one or more MRI scans, one or more CT scans, one or more X-rays, one or more ultrasound images, one or more nuclear medicine images, or one or more histological images.
22. A system for training at least one machine learning model for use in analyzing medical image data for a medical treatment and a speech-text conversion algorithm, the system A non-transitory computer-readable medium storing program instructions for training the machine learning model therein; At least one processor, wherein when the at least one processor executes the program instructions, Applying an encoder to at least one training image to generate at least one feature for a target training object (OOI) within the at least one training image; Selecting a class of the training OOI by applying the at least one feature to the at least one machine learning model; Reconstructing a labeled training image using a decoder by associating the at least one feature with the selected class for training the training image and the at least one machine learning model; Training the speech-text conversion algorithm to identify a match between new speech data and the ground truth text using a speech data set including ground truth text and speech data for the ground truth text, thereby generating at least one text string; Overlaying the training OOI and the at least one text string on an annotated image; At least one processor configured to perform; A system comprising.
23. The system according to claim 22, wherein the class is a healthy tissue class, an unhealthy tissue class, a suspicious tissue class, or an out-of-focus tissue class.
24. The at least one processor is Training the at least one machine learning model using a training data set including labeled training images, unlabeled training images, or a mixture of labeled and unlabeled training images, wherein the images include examples categorized by healthy tissue, unhealthy tissue, suspicious tissue, and out-of-focus tissue. The system according to claim 22 or claim 23, further configured to perform.
25. The system according to any one of claims 22 to 23, wherein the at least one processor is further configured to train the at least one machine learning model using supervised learning, unsupervised learning, or semi-supervised learning.
26. The system of claim 24, wherein the training data set further includes subcategories for each of the unhealthy tissue and the suspicious tissue. **Claim 27** The at least one processor receives a training image as input to the encoder; uses the encoder to project the training image into a feature space including features; uses a classifier to map the features to a set of target classes; identifies morphological characteristics of the training image to generate a training data set, wherein the training data set has data linking parameters to the training image; determines whether there is one or more mapped classes or no mapped classes based on the morphological characteristics and is further configured to create the at least one machine learning model thereby, according to any one of claims 22 to 23. **Claim 28** The at least one processor receives one or more of the features as input to the decoder; uses a transposed convolutional neural network to map one of the features to an unlabeled data set; uses the decoder to reconstruct a new training image from one of the features to train the at least one machine learning model and is further configured to perform the above, according to claim 27. **Claim 29** The system according to any one of claims 22 to 23, wherein the speech-to-text conversion algorithm maps the at least one OOI to one of a plurality of OOI medical terms. **Claim 30** The at least one processor is further configured to generate at least one new training image from an object of interest (OOI) detected during analysis of the medical image data when, based on the speech-to-text conversion algorithm that generates input speech matching the at least one text string, it is determined that the at least one text string associated with the OOI is the ground truth for that OOI, according to any one of claims 22 to 23. **Claim 31** The system according to any one of claims 22 to 23, wherein when the at least one processor determines that at least one text string associated with the OOI is not the ground truth for that OOI based on the speech-to-text conversion algorithm that generates input speech matching the at least one text string, the at least one processor is further configured to generate at least one new training image from an object of interest (OOI) detected during analysis of the medical image data.
32. The system according to any one of claims 22 to 23, wherein the training is performed on medical image data obtained from one or more endoscopic procedures, one or more MRI scans, one or more CT scans, one or more X-rays, one or more ultrasound images, one or more nuclear medicine images, or one or more histological images.
33. A method for analyzing medical image data for a medical procedure, the method comprising: Receiving at least one image from a series of images; Determining when at least one object of interest (OOI) is present in the at least one image and, when the at least one OOI is present, determining a classification of the at least one OOI, both determinations being performed using at least one machine learning model; During the medical procedure, using a bounding box to display the at least one image and any determined OOIs to a user on a display; Receiving an input audio signal including speech from the user during the medical procedure and recognizing the speech; When the speech is recognized as a comment on the at least one image during the medical procedure, converting the speech into at least one text string using a speech-to-text conversion algorithm and a term correction algorithm; Collating the at least one text string with at least one image provided with speech from the user; During the medical procedure, generating at least one annotated image in which the at least one text string is linked to at least one corresponding image; A method comprising.
34. When the utterance is recognized as including a request for at least one reference image including the classification, during the medical procedure, display at least one reference image having an OOI classified in the same classification as the at least one OOI, and to update the at least one machine learning model, further comprising receiving from the user an input to confirm or reject the classification for the at least one OOI, the method according to claim 33.
35. The method according to claim 33 or claim 34, further comprising receiving from the user an input indicating a user classification for at least one image having the undetermined OOI when the at least one OOI is classified as suspicious.
36. The method according to any one of claims 33 to 34, further comprising automatically generating a report including the at least one annotated image.
37. For a given OOI within a given image, identifying the bounding box coordinates of the bounding box, wherein the bounding box is associated with the given OOI within the given image, calculating a confidence score based on the probability distribution of the classification for the given OOI, overlaying the bounding box on the at least one image at the bounding box coordinates when the confidence score is higher than a confidence threshold, overlaying custom vocabulary on the at least one image when receiving confirmation from the user during the medical procedure, The method according to any one of claims 33 to 34, further comprising.
38. Determining the classification for the OOI is applying a convolutional neural network (CNN) to the OOI by performing convolutional operations, activation operations, and pooling operations to generate a matrix, generating a feature vector by processing the matrix using the convolutional operations, activation operations, and pooling operations, performing a classification for the OOI based on the feature vector, The method according to any one of claims 33 to 34, comprising.
39. The method according to any one of claims 33 to 34, further comprising overlaying at least one timestamp of an in-treatment event and a timestamped document on the corresponding at least one image when generating the at least one annotated image during the medical treatment.
40. The method according to claim 34, further comprising displaying a confidence score of the at least one image on a display in real time during the medical treatment.
41. Receiving the input voice during the medical treatment is starting to receive an audio stream for the input voice from the user upon detection of a first user action, wherein the first user action is temporarily stopping the display of the series of images, taking a snapshot of a given image within the series of images, or providing a first voice command including; ending the reception of the audio stream upon detection of a second user action, wherein the second user action is remaining silent for a predetermined length, pressing a designated button, or providing a last voice command including; The method according to any one of claims 33 to 34, including.
42. The method according to any one of claims 33 to 34, further comprising storing the series of images when receiving the input voice during the medical treatment, thereby designating the at least one image to receive annotation data for generating a corresponding at least one annotated image.
43. Capturing a set of patient information data generated during the medical treatment to be added to the report; loading a subset of the series of images including the at least one annotated image or at least one OOI identified by the bounding box; combining the set of patient information data and the subset of the series of images including the at least one annotated image into the report; The method according to any one of claims 33 to 34, further comprising generating a report of the medical treatment thereby.
44. Apply an encoder to at least one training image to generate at least one feature vector for a training OOI within the at least one training image; Select a class of the training OOI by applying the at least one feature vector to the at least one machine learning model; Reconstruct a labeled training image using a decoder by associating the at least one feature vector with the at least one training image and a selected class for training the at least one machine learning model; The method according to any one of claims 33 to 34, further comprising performing training of the at least one machine learning model by.
45. The method according to claim 44, wherein the class is a healthy tissue class, an unhealthy tissue class, a suspicious tissue class, or an out-of-focus tissue class.
46. Training the at least one machine learning model using a training data set including labeled training images, unlabeled training images, or a mixture of labeled training images and unlabeled training images, wherein the images include examples categorized by healthy tissue, unhealthy tissue, suspicious tissue, and out-of-focus tissue; The method according to claim 44, further comprising.
47. The method according to claim 44, wherein training the at least one machine learning model includes using supervised learning, unsupervised learning, or semi-supervised learning.
48. The method according to claim 46, wherein the training data set further includes subcategories for each of the unhealthy tissue and the suspicious tissue.
49. Receiving a training image as an input to the encoder; Projecting the training image onto features that are part of a feature space using the encoder; Mapping the features to a set of target classes using a classifier; Identifying morphological characteristics of the training image to generate a new training data set, wherein the new training data set has data linking parameters to the training image; Based on the morphological characteristics, determining whether there is one or more mapped classes or no mapped classes The method according to claim 44, further comprising creating the at least one machine learning model by doing so.
50. Determining the classification for the at least one OOI is Receiving one or more of the features as input to the decoder Using a convolutional neural network to map one of the features to an unlabeled dataset Reconstructing a new training image from one of the features using the decoder to train the at least one machine learning model The method according to claim 49, comprising.
51. Training the speech-to-text conversion algorithm using the speech dataset to identify a match with the ground truth text by comparing new speech data with the speech dataset, the speech dataset further comprising the ground truth text and the speech data of the ground truth text.
52. The method according to claim 43, wherein the speech-to-text conversion algorithm maps the at least one OOI to one of a plurality of OOI medical terms.
53. The method according to any one of claims 33 to 34, wherein the medical image data is obtained from one or more endoscopic procedures, one or more MRI scans, one or more CT scans, one or more X-rays, one or more ultrasound images, one or more nuclear medicine images, or one or more histological images.
54. A method for training at least one machine learning model and a speech-to-text conversion algorithm for use in analyzing medical image data for medical treatment, the method comprising Applying an encoder to at least one training image to generate at least one feature for a target training object (OOI) within the at least one training image Selecting the class of the training OOI by applying the at least one feature to the at least one machine learning model Reconstructing the labeled training image using a decoder by associating the at least one feature with a selected class for training the training image and the at least one machine learning model, Training the speech-to-text conversion algorithm to identify a match between new speech data and the ground truth text using a speech data set including the ground truth text and the speech data for the ground truth text, thereby generating at least one text string, Overlaying the training OOI and the at least one text string onto an annotated image, A method comprising.
55. The method according to claim 54, wherein the class is a healthy tissue class, an unhealthy tissue class, a suspicious tissue class, or an unfocused tissue class.
56. Training the at least one machine learning model using a training data set including labeled training images, unlabeled training images, or a mixture of labeled and unlabeled training images, wherein the images include examples categorized by healthy tissue, unhealthy tissue, suspicious tissue, and unfocused tissue The method according to claim 54 or claim 55, further comprising.
57. The method according to any one of claims 54 to 55, wherein training the at least one machine learning model includes using supervised learning, unsupervised learning, or semi-supervised learning.
58. The method according to claim 56, wherein the training data set further includes subcategories for each of the unhealthy tissue and the suspicious tissue.
59. Receiving a training image as an input to the encoder, Projecting the training image into a feature space including features using the encoder, Mapping the features to a set of target classes using a classifier, Identifying morphological characteristics of the training image to generate a training data set, wherein the training data set has data linking parameters to the training image Based on the morphological characteristics, determining whether there are one or more mapped classes or no mapped classes; The method according to any one of claims 54 to 55, further comprising creating the at least one machine learning model thereby.
60. Receiving one or more of the features as input to the decoder; Using a transposed convolutional neural network to map one of the features to an unlabeled dataset; Using the decoder to reconstruct a new training image from one of the features for training the at least one machine learning model; The method according to claim 59, further comprising.
61. The method according to any one of claims 54 to 55, wherein the speech-to-text conversion algorithm maps the at least one OOI to one of a plurality of OOI medical terms.
62. Based on the speech-to-text conversion algorithm that generates input speech matching the at least one text string, when it is determined that at least one text string associated with the OOI is the ground truth for that OOI, generating at least one new training image from the object of interest (OOI) detected during the analysis of the medical image data. The method according to any one of claims 54 to 55, further comprising.
63. Based on the speech-to-text conversion algorithm that generates input speech matching the at least one text string, when it is determined that at least one text string associated with the OOI is not the ground truth for that OOI, generating at least one new training image from the object of interest (OOI) detected during the analysis of the medical image data. The method according to any one of claims 54 to 55, further comprising.
64. The training is performed on medical image data obtained from one or more endoscopic procedures, one or more MRI scans, one or more CT scans, one or more X-rays, one or more ultrasound images, one or more nuclear medicine images, or one or more histological images. The method according to any one of claims 54 to 55.