Slide imaging apparatus and method
The apparatus and method for slide imaging address the challenge of out-of-focus errors by using an optical system to capture images, identify focus patterns, and extrapolate focal distances, resulting in improved image quality and reduced errors.
Patent Information
- Application Number
- JP2024187029
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-10-28
- Filing Date
- 2024-10-23
- Publication Date
- 2025-06-23
- Estimated Expiration
- 2044-10-23
AI Technical Summary
In slide scanning, out-of-focus errors can lead to significant penalties, especially when decisions are based on the slide. Existing methods struggle to accurately determine the best-focus Z-plane during a single scan, particularly when encountering dust, tissue, or pen marks.
An apparatus and method for slide imaging that includes an optical system with an optical sensor, a slide port, a processor, and memory. The system captures a first image, identifies a focus pattern, extrapolates a focal distance, and captures a second image at the extrapolated focal distance, allowing for improved focus determination across the slide.
This approach enables more accurate focus determination and image capture, reducing the impact of out-of-focus errors and improving the quality of scanned images.
Smart Images

Figure 0007697125000007 
Figure 0007697125000008 
Figure 0007697125000009
Abstract
Description
Technical Field
[0001] The present invention generally relates to the field of medical image processing. In particular, the present invention relates to an apparatus and method for slide imaging.
Background Art
[0002] In slide scanning, the penalty for out-of-focus errors can be significant, especially when decisions are made based on that slide. In an approach that digitizes a slide in a single pass, it is necessarily required to determine the best-focus Z-plane during a single scan, thus forcing a decision in a situation where there may be tissue, dust, pen marks, etc. on the slide. For example, when encountering dust, the wrong Z-reference plane will be selected in that scan, and the tissue in front of the dust will not be in focus. Similarly, pen marks also completely defocus the tissue. This is because the Z-plane of the pen mark is on top of the cover glass.
Summary of the Invention
Means for Solving the Problems
[0003] In one aspect, an apparatus for imaging a slide includes at least one optical system including an optical sensor, a slide port configured to hold the slide, at least one processor, and a memory communicatively connected to the at least one processor, the memory storing instructions to configure the at least one processor to receive at least one region of interest, use the at least one optical system to capture a first image of the slide at a first position within the at least one region of interest, identify a focus pattern as a function of the first image and the first position, extrapolate a focal distance at a second position as a function of the focus pattern, and use the at least one optical system to capture a second image of the slide at the second position and the focal distance.
[0004] In another aspect, a method of imaging a slide includes receiving at least one region of interest using at least one processor, capturing a first image of the slide at a first position within the at least one region of interest using at least one processor and at least one optical system, identifying a focus pattern as a function of the first image and the first position using at least one processor, extrapolating a focal distance at a second position as a function of the focus pattern using at least one processor, and capturing a second image of the slide at the second position and the focal distance using at least one processor and the at least one optical system.
[0005] These and other aspects and features of non-limiting embodiments of the present invention will become apparent to those skilled in the art upon review of the following description of specific non-limiting embodiments of the present invention in conjunction with the accompanying drawings.
Brief Description of the Drawings
[0006] For the purpose of illustrating the present invention, the drawings show aspects of one or more embodiments of the present invention. However, it should be understood that the present invention is not limited to the exact arrangements and instrumentalities shown in the drawings.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6A
Figure 6B
Figure 6C
Figure 7
Figure 8
Best Mode for Carrying Out the Invention
[0007] At a high level, aspects of the present disclosure relate to an apparatus and method for slide imaging. The apparatus described herein can generate an image of a slide and / or a sample on the slide. In one embodiment, the apparatus may capture a first image and identify one or more regions of interest. The regions of interest may include features such as writing, dust, samples, etc. The sample may include tissue. In one embodiment, the apparatus can identify a focus pattern for the region of interest. For example, the apparatus can identify the focus pattern as a plane based on a plurality of points at which an optimal focus is determined. In one embodiment, the apparatus may determine which regions contain a sample. In some embodiments, which regions contain a sample may be determined after one or more of the other steps described herein. In some embodiments, delaying the determination of which regions contain a sample can make the slide imaging process more efficient. For example, since it is difficult to run a sophisticated model on a scanning device during scanning, efficiency may be improved. In another example, when classifying a region as dust or an annotation and skipping the scan of the region, the risk of false positives is minimized, so efficiency may be improved. In another example, since it may be useful to scan the annotation, it may be optimal to perform the steps in this order. Exemplary embodiments showing aspects of the present disclosure are described below in the context of several specific examples.
[0008] Referring now to FIG. 1, an exemplary embodiment of an apparatus 100 for slide imaging is illustrated. The apparatus 100 can include a computing device. The apparatus 100 can include a processor 104. The processor 104 can include any processor 104 described in the present disclosure, without limitation. The processor 104 may be included in a computing device. The apparatus 100 may include at least one processor 104 and a memory 108 communicatively connected to the at least one processor 104, and instructions 112 for configuring the at least one processor 104 to execute one or more processes described herein are stored in the memory 108. The computing device can include any computing device described in the present disclosure, including but not limited to a microcontroller, a microprocessor, a digital signal processor (DSP), and / or a system-on-chip (SoC) described in the present disclosure. The computing device can include, be included in, and / or communicate with a mobile device such as a mobile phone or a smartphone. The computing device may include a single computing device operating independently, or may include two or more computing devices operating in cooperation, in parallel, sequentially, etc., and the two or more computing devices may be included together in a single computing device or may be included in two or more computing devices. The computing device can interface or communicate with one or more additional devices via a network interface device, as will be described in more detail later. The network interface device can be utilized to connect the computing device to one or more of various networks and one or more devices. Examples of network interface devices include, without limitation, network interface cards (e.g., mobile network interface cards, LAN cards), modems, and any combination thereof.Examples of networks include, but are not limited to, wide area networks (e.g., the Internet, corporate networks), local area networks (e.g., networks associated with offices, buildings, campuses, or other relatively small geographical spaces), telephone networks, data networks associated with telephone / voice providers (e.g., data and / or voice networks of mobile communication providers), direct connections between two computing devices, and any combination thereof. The network can employ wired and / or wireless communication modes. Generally, any network topology can be used. Information (e.g., data, software, etc.) can be communicated to and / or from computers and / or computing devices. Computing devices can include, but are not limited to, for example, a computing device or a cluster of computing devices at a first location, and a second computing device or a cluster of computing devices at a second location. Computing devices can include one or more computing devices specialized for data storage, security, traffic distribution for load balancing, etc. Computing devices can operate one or more computing tasks as described below in parallel, serially, redundantly, or in any other way used for task or memory distribution between computing devices, and can be distributed among multiple computing devices of the computing device. Computing devices can be implemented, as a non-limiting example, using a "shared nothing" architecture.
[0009] Continuing to refer to FIG. 1, a computing device can be designed and / or configured to execute any method, method step, or sequence of method steps described in any embodiment of the present disclosure in any order and to any extent. For example, the computing device may be configured to repeatedly execute a single step or sequence until a desired or commanded result is achieved. The repetition of a step or sequence of steps can be executed iteratively and / or recursively using the output of a previous iteration as input to a subsequent iteration, the aggregation of iteration inputs and / or outputs to produce an aggregated result, the decrement or decrementation of one or more variables such as global variables, and / or the division of a large processing task into a set of smaller processing tasks that are repeatedly addressed. The computing device can execute any step or sequence of steps described in the present disclosure in parallel, such as by using two or more parallel threads, processor cores, etc., to execute a step two or more times simultaneously and / or substantially simultaneously, and the division of tasks between parallel threads and / or processes can be executed according to any protocol suitable for dividing tasks between iterations. Those skilled in the art will recognize, upon considering the entirety of the present disclosure, various ways in which iterations, recursion, and / or parallel processing can be used to subdivide, share, or otherwise process steps, sequences of steps, processing tasks, and / or data.
[0010] Referring still to FIG. 1, "communicatively connected" as used in the present disclosure means being connected by a connection, attachment, or coupling that enables reception and / or transmission of information between two or more related entities. For example, but not limited to, this connection may be wired or wireless, direct or indirect, and between two or more components, circuits, devices, systems, etc., and enables reception and / or transmission of data and / or signals. The data and / or signals therebetween can include, but are not limited to, in particular, electrical, electromagnetic, magnetic, video, audio, wireless, and microwave data and / or signals, combinations thereof, etc. The communicative connection can be achieved, for example, but not limited to, by wired or wireless electronic, digital, or analog communication, directly or via one or more intervening devices or components. Further, the communicative connection can include electrically coupling or connecting at least one output of a device, component, or circuit to at least one input of another device, component, or circuit. For example, but not limited to, it may be via a bus or other facilities for mutual communication between computing device elements. The communicative connection can include, for example, but not limited to, indirect connections via wireless connections, wireless communications, low-power wide-area networks, optical communications, magnetic coupling, capacitive coupling, optical coupling, etc. In some cases, the term "communicatively coupled" may be used in the present disclosure instead of "communicatively connected".
[0011] Referring still to FIG. 1, in some embodiments, the apparatus 100 can be used to generate an image of the slide 116 and / or a sample on the slide 116. As used herein, a "slide" is a container or surface that holds a sample of interest. In some embodiments, the slide 116 may include a glass slide. In some embodiments, the slide 116 may include a formalin-fixed paraffin-embedded slide. In some embodiments, the sample on the slide 116 may be stained. In some embodiments, the slide 116 may be substantially transparent. In some embodiments, the slide 116 may include a thin, flat, and substantially transparent glass slide. In some embodiments, a transparent cover may be applied to the slide 116 such that the sample is present between the slide 116 and the cover. The sample may include, by way of non-limiting example, a blood smear, a cervical smear, a body fluid, and a non-biological sample. In some embodiments, the sample on the slide 116 may include tissue. In some embodiments, the sample on the slide 116 may be frozen.
[0012] Referring still to FIG. 1, in some embodiments, the slide 116 and / or the sample on the slide 116 may be illuminated. In some embodiments, the apparatus 100 may include a light source. As used herein, a "light source" is any device configured to emit electromagnetic radiation. In some embodiments, the light source may emit light having substantially one wavelength. In some embodiments, the light source may emit light having a range of wavelengths. The light source may emit, without limitation, ultraviolet light, visible light, and / or infrared light. By way of non-limiting example, the light source may include a light-emitting diode (LED), an organic LED (OLED), and / or other light emitters. Such a light source may be configured to illuminate the slide 116 and / or the sample on the slide 116. By way of non-limiting example, the light source may illuminate the slide 116 and / or the sample on the slide 116 from below.
[0013] Referring still to FIG. 1, in some embodiments, apparatus 100 may include at least one optical system 120. As used herein, an "optical system" is an arrangement of one or more components that act on or employ electromagnetic radiation. In non-limiting examples, it may include light such as electromagnetic radiation, visible light, infrared light, ultraviolet light, etc. The optical system may include one or more optical elements including, but not limited to, lenses, mirrors, windows, filters, etc. The optical system can form an optical image corresponding to an optical object. For example, the optical system can form an optical image in or on an optical sensor, and the optical sensor can capture (e.g., digitize) the optical image. In some cases, the optical system can have at least one magnification. For example, the optical system may include an objective lens (e.g., a microscope objective lens) and one or more re-imaging optical elements that together form an optical magnification. In some cases, optical zoom may also be referred to as zoom. As used herein, an "optical sensor" is a device that measures light and converts the measured light into one or more signals, which may include, but are not limited to, one or more electrical signals. In some embodiments, optical sensor 120 may include at least one photodetector. As used herein, a "photodetector" is a device that is sensitive to light and can thereby detect light. In some embodiments, the photodetector may include a photodiode, a photore resistor, a photosensor, a photovoltaic chip, etc. In some embodiments, optical sensor 120 may include a plurality of photodetectors. Optical sensor 120 may include, but is not limited to, a camera. Optical sensor 120 can communicate electronically with at least one processor 104 of apparatus 100. As used herein, "electronic communication" is a shared data connection between two or more devices. In some embodiments, apparatus 100 may include two or more optical sensors 120.
[0014] Referring still to FIG. 1, as used herein, "image data" is information representing at least one physical scene, space, and / or object. The image data may include, for example, information representing a sample, slide 116, or a region of the sample or slide. In some cases, the image data may be generated by a camera. "Image data" may be used interchangeably with "image" throughout the present disclosure, and the image is used as a noun. The image may be optical, such as those in which at least one optical system is used to generate an image of an object, but is not limited thereto. The image may be digital, such as when represented as a bitmap, but is not limited thereto. Alternatively, the image may include any medium capable of representing a physical scene, space, and / or object. Alternatively, when "image" is used as a verb in the present disclosure, it refers to the generation and / or formation of an image.
[0015] Referring still to FIG. 1, in some embodiments, the apparatus 100 may include a slide port 140. In some embodiments, the slide port 140 may be configured to hold the slide 116. In some embodiments, the slide port 140 may include one or more alignment features. As used herein, an "alignment feature" is a physical characteristic that serves to fix a slide in a predetermined position and / or align the slide with other components of the apparatus. In some embodiments, the alignment feature may include components for fixing the slide 116, such as a clamp, latch, clip, recess, or other fastener. In some embodiments, the slide port 140 may facilitate the removal and insertion of the slide 116. In some embodiments, the slide port 140 may include a transparent surface through which light can pass. In some embodiments, the slide 116 may be placed on such a transparent surface and / or illuminated by light passing through such a transparent surface. In some embodiments, the slide port 140 may be mechanically connected to the actuator mechanism 124, as described below.
[0016] Referring still to FIG. 1, in some embodiments, the apparatus 100 may include an actuator mechanism 124. As used herein, an "actuator mechanism" is a mechanical component configured to change the relative position between the slide and the optical system. In some embodiments, the actuator mechanism 124 may be mechanically connected to the slide 116, such as the slide 116 within the slide port 140. In some embodiments, the actuator mechanism 124 may be mechanically connected to the slide port 140. For example, the actuator mechanism 124 may move the slide port 140 to move the slide 116. In some embodiments, the actuator mechanism 124 may be mechanically connected to at least one optical system 120. In some embodiments, the actuator mechanism 124 may be mechanically connected to a movable element. As used herein, a "movable element" refers to any movable or portable object, component, and device within the apparatus 100, including but not limited to slides, slide ports, optical systems, etc. In some embodiments, the movable element may move such that the optical system 120 is properly positioned relative to the slide 116 so that the optical system 120 can capture an image of the slide 116 according to a parameter set. In some embodiments, the actuator mechanism 124 may be mechanically connected to an item selected from the list consisting of the slide port 140, the slide 116, and at least one optical system 120. In some embodiments, the actuator mechanism 124 may be configured to change the relative position between the slide 116 and the optical system 120 by moving the slide port 140, the slide 116, and / or the optical system 120.
[0017] Referring still to FIG. 1, the actuator mechanism 124 may include components of a machine that serves to move and / or control a mechanism or system. The actuator mechanism 124 may require a control signal and / or an energy source or power in some embodiments. In some cases, the control signal may be relatively low energy. Exemplary forms of control signals include electric potential or current, pneumatic or flow rate, or hydraulic fluid pressure or flow rate, mechanical force / torque or speed, or even human power. In some cases, the actuator may have an energy source or power source other than the control signal. This may include a primary energy source, which can include, for example, electric power, hydraulic power, pneumatic power, mechanical power, etc. In some embodiments, upon receiving a control signal, the actuator mechanism 124 responds by converting the source power into mechanical motion. In some cases, the actuator mechanism 124 may be understood as a form of automation or automatic control.
[0018] Referring still to FIG. 1, in some embodiments, the actuator mechanism 124 may include a hydraulic actuator. The hydraulic actuator may be composed of a cylinder or a fluid motor that uses hydraulic pressure to facilitate mechanical movement. The output of the hydraulic actuator mechanism 124 may include mechanical movement such as, but not limited to, linear movement, rotational movement, or oscillatory movement. In some embodiments, the hydraulic actuator may employ a hydraulic fluid. Since the liquid is incompressible in some cases, the hydraulic actuator can apply a large force. Further, since the force is equal to the pressure multiplied by the area, the hydraulic actuator can function as a force transducer that changes with the change in area (e.g., the cross-sectional area of the cylinder and / or piston). An exemplary hydraulic cylinder may be composed of a hollow cylindrical tube in which a piston can slide. In some cases, the hydraulic cylinder may be regarded as a single-acting type. The "single-acting type" can be used when the fluid pressure is applied substantially only to one side of the piston. Therefore, the single-acting piston can move only in one direction. In some cases, a spring may be used to give a return stroke to the single-acting piston. In some cases, the hydraulic cylinder may be a double-acting type. The "double-acting type" can be used when the pressure is applied substantially to both sides of the piston. The piston moves due to the difference in force generated between both sides of the piston.
[0019] Referring still to FIG. 1, in some embodiments, the actuator mechanism 124 may include a pneumatic actuator mechanism 124. In some cases, the pneumatic actuator can generate a large force from a relatively small change in gas pressure. In some cases, the pneumatic actuator can respond faster than other types of actuators such as, for example, hydraulic actuators. The pneumatic actuator can use a compressible fluid (e.g., air). In some cases, the pneumatic actuator can operate with compressed air. The operation of the hydraulic and / or pneumatic actuator includes the control of one or more valves, circuits, fluid pumps, and / or fluid manifolds.
[0020] Referring still to FIG. 1, in some cases, the actuator mechanism 124 may include an electric actuator. The electric actuator mechanism 124 may include any of an electromechanical actuator, a linear motor, etc. In some cases, the actuator mechanism 124 may include an electromechanical actuator. The electromechanical actuator can convert the rotational force of an electric rotary motor into linear motion and generate linear motion through a mechanism. Exemplary mechanisms include, but are not limited to, a belt, a screw, a crank, a cam, a linkage, a Scotch yoke, etc., and include a converter from rotational motion to translational motion. In some cases, the control of the electromechanical actuator may include the control of an electric motor. For example, the control signal may control one or more electric motor parameters to control the electromechanical actuator. Exemplary non-limiting electric motor parameters include rotational position, input torque, speed, current, and potential. The electric actuator mechanism 124 may include a linear motor. Since the power from the linear motor is directly output as translational motion rather than being output as rotational motion and then converted into translational motion, the linear motor may be different from the electromechanical actuator. In some cases, the linear motor may have less frictional loss than other devices. The linear motor may be designated into at least three different categories, such as a flat linear motor, a U-channel linear motor, a tubular linear motor, etc. The linear motor may be directly controlled by a control signal that controls one or more linear motor parameters. Exemplary linear motor parameters include, but are not limited to, position, force, speed, potential, and current.
[0021] Referring still to FIG. 1, in some embodiments, the actuator mechanism 124 may include a mechanical actuator mechanism 124. In some cases, the mechanical actuator mechanism 124 may function to perform motion by converting one type of motion, such as rotational motion, into another type of motion, such as linear motion. Exemplary mechanical actuators include rack and pinion. In some cases, a mechanical power source, such as a power takeoff, may function as the power source for the mechanical actuator. The mechanical actuator may employ any number of mechanisms, including, but not limited to, gears, rails, pulleys, cables, linkages, and the like.
[0022] Referring still to FIG. 1, in some embodiments, the actuator mechanism 124 can communicate electronically with an actuator control. As used herein, "actuator control" is a system configured to operate the actuator mechanism so that the slide and the optical system are in a desired relative position. In some embodiments, the actuator control can operate the actuator mechanism 124 based on an input received from the user interface 136. In some embodiments, the actuator control can be configured to operate the actuator mechanism 124 so that the optical system 120 is in a position to capture an image of the entire sample. In some embodiments, the actuator control can be configured to operate the actuator mechanism 124 so that the optical system 120 is in a position to capture an image of a region of interest, a particular horizontal row, a particular point, a particular depth of focus, and the like. The electronic communication between the actuator mechanism 124 and the actuator control may include the transmission of signals. For example, the actuator control can generate a physical movement of the actuator mechanism in response to an input signal. In some embodiments, the input signal can be received by the actuator control from the processor 104 or the input interface 128.
[0023] Referring still to FIG. 1, the "signal" used in the present disclosure is, for example, any understandable representation of data from one device to another device. The signal may include, for example, an optical signal, a hydraulic signal, a pneumatic signal, a mechanical signal, an electrical signal, a digital signal, an analog signal, etc. In some cases, the signal may be used to communicate with a computing device, for example, via one or more ports. In some cases, the signal may be transmitted and / or received by a computing device, for example, via an input / output port. The analog signal may be digitized, for example, by an analog-to-digital converter. In some cases, the analog signal may be processed, for example, by the analog signal processing steps described in the present disclosure, before being digitized. In some cases, the digital signal may be used for communication between two or more devices, including but not limited to a computing device. In some cases, the digital signal may be communicated by one or more communication protocols, including but not limited to the Internet Protocol (IP), the Controller Area Network (CAN) protocol, the serial communication protocol (e.g., Universal Asynchronous Receiver Transmitter [UART]), the parallel communication protocol (e.g., IEEE128 [printer port]), etc.
[0024] Referring still to FIG. 1, in some embodiments, the apparatus 100 can perform one or more signal processing steps on a signal. For example, the apparatus 100 can analyze, modify, and / or synthesize a signal representing data to improve the signal, e.g., by improving transmission, storage efficiency, or signal-to-noise ratio. Exemplary methods of signal processing can include analog, continuous-time, discrete, digital, non-linear, statistical, etc. Analog signal processing can be performed on non-digitized signals or analog signals. Exemplary analog processing can include passive filters, active filters, adder mixers, integrators, delay lines, companders, multipliers, voltage-controlled filters, voltage-controlled oscillators, phase-locked loops, etc. Continuous-time signal processing can, in some cases, be used to process signals that vary continuously within a region, e.g., the time domain. Exemplary non-limiting continuous-time processes can include time-domain processing, frequency-domain processing (Fourier transform), complex frequency-domain processing. Discrete-time signal processing can be used when the signal is sampled at non-continuous or discrete time intervals (i.e., time-quantized). Analog discrete-time signal processing can process signals using the following exemplary circuits sample-and-hold circuits, analog time-division multiplexers, analog delay lines, analog feedback shift registers. Digital signal processing can be used to process digitized discrete-time sampled signals. Generally, digital signal processing can be performed by computing devices, or other specialized digital circuits such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or special digital signal processors (DSPs), but not limited to these. Digital signal processing can be used to perform any combination of typical arithmetic operations such as fixed-point, floating-point, real-valued, complex-valued, multiplication, addition, etc. Digital signal processing can further manipulate circular buffers and look-up tables.Further non-limiting examples of algorithms that can be executed in accordance with digital signal processing techniques include fast Fourier transform (FFT), finite impulse response (FIR) filters, infinite impulse response (IIR) filters, and adaptive filters such as Wiener filters and Kalman filters. Statistical signal processing can be used to process signals as random functions (i.e., stochastic processes) by utilizing statistical characteristics. For example, in some embodiments, a signal may be modeled with a probability distribution indicative of noise, which can be used to reduce the noise of the processed signal.
[0025] Still referring to FIG. 1, in some embodiments, apparatus 100 may include a user interface 136. The user interface 136 may include an output interface 132 and an input interface 128.
[0026] Still referring to FIG. 1, in some embodiments, the output interface 132 may include one or more elements through which apparatus 100 can communicate information to the user. In non-limiting examples, the output interface 132 may include a display. The display may include a high-resolution display. The display can output images, videos, etc. to the user. In another non-limiting example, the output interface 132 may include a speaker. The speaker can output sound to the user. In another non-limiting example, the output interface 132 may include a tactile device. The speaker can output tactile feedback to the user.
[0027] Referring still to FIG. 1, in some embodiments, the optical system 120 may include a camera. In some cases, the camera may include one or more optical systems. Exemplary non-limiting optical systems include spherical lenses, aspherical lenses, mirrors, polarizers, filters, windows, aperture stops, and the like. In some embodiments, one or more of the optical systems associated with the camera may be adjusted to change, by way of non-limiting example, the zoom, depth of field, and / or focal length of the camera. In some embodiments, one or more of such settings may be configured to detect features of the sample on the slide 116. In some embodiments, one or more of such settings may be configured based on a parameter set, as described below. In some embodiments, the camera can capture an image with a shallow depth of field. By way of non-limiting example, the camera can capture an image that is in focus at a first depth of the sample and out of focus at a second depth of the sample. In some embodiments, an autofocus mechanism may be used to determine the focal length. In some embodiments, the focal length may be set by a parameter set. In some embodiments, the camera may be configured to capture multiple images at different focal lengths. By way of non-limiting example, the camera can capture multiple images at different focal lengths such that the image is captured where each focal depth of the sample is in focus in at least one image. In some embodiments, at least one camera may include an image sensor. Exemplary non-limiting image sensors include digital image sensors such as, but not limited to, charge-coupled device (CCD) sensors and complementary metal-oxide-semiconductor (CMOS) sensors. In some embodiments, the camera can be sensitive within the non-visible range of electromagnetic radiation, such as, but not limited to, infrared light.
[0028] Referring still to FIG. 1, in some embodiments, the input interface 128 may include controls for operating the device 100. Such controls may be operated by a user. The input interface 128 may include, by way of non-limiting example, a camera, a microphone, a keyboard, a touch screen, a mouse, a joystick, a foot pedal, buttons, a dial, and the like. The input interface 128 may be capable of receiving, by way of non-limiting example, mechanical input, voice input, visual input, text input, and the like. In some embodiments, voice input to the input interface 128 may be interpreted using an automatic speech recognition function, enabling the user to control the device 100 via voice. In some embodiments, the input interface 128 may approximate the control of a microscope.
[0029] Referring still to FIG. 1, in some embodiments, voice input may be processed using automatic speech recognition. In some embodiments, automatic speech recognition may require training (i.e., enrollment). In some cases, training an automatic speech recognition model may require an individual speaker to read text or isolated vocabulary words. In some cases, the voice training data may include audio components with audible language content that is known a priori by the computing device. Thus, the computing device may train an automatic speech recognition model according to training data that includes audible language content correlated with known content. In this way, the computing device may analyze a person's specific voice and train an automatic speech recognition model for that person's voice, resulting in improved accuracy. Alternatively or additionally, in some cases, the computing device may include a speaker-independent automatic speech recognition model. As used in this disclosure, a "speaker-independent" automatic speech recognition process does not require training for an individual speaker. Conversely, as used in this disclosure, an automatic speech recognition process that employs training specific to an individual speaker is "speaker-dependent".
[0030] Referring still to FIG. 1, in some embodiments, the automatic speech recognition process may perform speech recognition or speaker identification. As used in the present disclosure, "speech recognition" refers to identifying a speaker from audio content rather than the content spoken by the speaker. In some cases, the computing device may first recognize the speaker of the spoken language audio content and then automatically recognize the speaker's voice, for example, by a speaker-dependent automatic speech recognition model or process. In some embodiments, the automatic speech recognition process may be used to authenticate or verify the identity of the speaker. In some cases, the speaker may or may not include a subject. For example, a subject may speak during voice input, but other people may also speak.
[0031] Referring still to FIG. 1, in some embodiments, the automatic speech recognition process may include one or all of acoustic modeling, language modeling, and a statistic-based speech recognition algorithm. In some cases, the automatic speech recognition process may employ a hidden Markov model (HMM). As will be described in detail later, language modeling, such as that employed in natural language processing applications such as document classification and statistical machine translation, may also be employed in the automatic speech recognition process.
[0032] Referring still to FIG. 1, exemplary algorithms employed for automatic speech recognition may include or be based on a hidden Markov model. A hidden Markov model (HMM) can include a statistical model that outputs a sequence of symbols or quantities. Since an audio signal can be considered a piecewise stationary signal or a short-time stationary signal, an HMM can be used for speech recognition. For example, on a short time scale (e.g., 10 milliseconds), speech can be approximated as a stationary process. Speech (i.e., audible language content) can be understood as a Markov model for many probabilistic purposes.
[0033] Referring still to FIG. 1, in some embodiments, the HMM can be automatically trained, is relatively easy to use, and can be computationally feasible. In an exemplary automatic speech recognition process, the hidden Markov model can output a sequence of n-dimensional real-valued vectors (n being a small integer such as 10) at a rate of 1 vector every approximately 10 milliseconds. The vectors may be composed of cepstral coefficients. Cepstral coefficients require the use of the spectral domain. Cepstral coefficients are obtained by performing a Fourier transform on a short time window of the speech that generates the spectrum, decorrelating the spectrum using a cosine transform, and taking the first (i.e., most important) coefficient. In some cases, the HMM may have a statistical distribution in each state that is a mixture of diagonal covariance Gaussian distributions, generating a likelihood for each observation vector. In some cases, each word or phoneme may have a different output distribution. The HMM for a sequence of words or phonemes can be created by concatenating the HMMs for the individual words or phonemes.
[0034] Referring still to FIG. 1, in some embodiments, the automatic speech recognition process can use various combinations of several techniques to improve the results. In some cases, the large vocabulary automatic speech recognition process may include phoneme context dependence. For example, in some cases, phonemes with different left and right contexts may have different recognitions as HMM states. In some cases, the automatic speech recognition process may use cepstrum normalization to normalize for different speakers and recording conditions. In some cases, the automatic speech recognition process can use vocal tract length normalization (VTLN) for male and female normalization and maximum likelihood linear regression (MLLR) for more general speaker adaptation. In some cases, the automatic speech recognition process can determine so-called delta coefficients and delta-delta coefficients to capture speech dynamics and can also use heteroscedastic linear discriminant analysis (HLDA). In some cases, the automatic speech recognition process can use projection based on splicing and linear discriminant analysis (LDA), which can include heteroscedastic linear discriminant analysis or global semi-tied covariance transformation (also called maximum likelihood linear transformation [MLLT]). In some cases, the automatic speech recognition process can omit a purely statistical approach to HMM parameter estimation and instead use discriminative training techniques that optimize some classification-related measure of the training data. Examples of this include maximum mutual information (MMI), minimum classification error (MCE), and minimum phoneme error (MPE).
[0035] Referring still to FIG. 1, in some embodiments, the automatic speech recognition process may be said to decrypt speech (i.e., audible language content). Decryption of speech occurs when the automatic speech recognition system is presented with a new utterance and must compute the most likely sentence. In some cases, speech decryption can include the Viterbi algorithm. The Viterbi algorithm can include a dynamic programming algorithm for obtaining the maximum a posteriori probability estimate of the most likely sequence of hidden states (i.e., the Viterbi path) that results in the observed sequence of events. The Viterbi algorithm can be employed in the context of Markov sources and hidden Markov models. The Viterbi algorithm can be used, for example, to find the best path using a statically created combined hidden Markov model (e.g., the finite state transducer [FST] approach), using a dynamically created combined hidden Markov model that has information from both an acoustic model and a language model.
[0036] Referring still to FIG. 1, in some embodiments, decoding of speech (i.e., audible language content) can include considering not only the best candidates but also good candidates when a new utterance is presented. In some cases, a better scoring function (i.e., re-scoring) can be used to evaluate each of a set of good candidates and select the best candidate according to this refined score. In some cases, the set of candidates can be maintained either as a list (i.e., N-best list approach) or as a subset of models (i.e., lattice). In some cases, re-scoring can be performed by optimizing the Bayesian risk (or an approximation thereof). In some cases, re-scoring can include optimizing a sentence (including keywords) that minimizes the expected value of a given loss function for all possible transcriptions. For example, re-scoring enables the selection of a sentence that minimizes the average distance to other possible sentences weighted by the estimated probabilities. In some cases, different distance calculations can be performed for a particular task, for example, and the loss function employed may include the Levenshtein distance. In some cases, the set of candidates may be truncated to maintain tractability.
[0037] Referring still to FIG. 1, in some embodiments, the automatic speech recognition process can employ a dynamic time warping (DTW)-based approach. The dynamic time warping method can include an algorithm that measures the similarity between two sequences with different times or speeds. For example, even if a person walks slowly in one video and walks fast in another video, and there are accelerations and decelerations during the same observation, the similarity of the walking patterns can be detected. DTW has been applied to video, audio, and graphics, and in fact, all data that can be converted into a linear representation can be analyzed by DTW. In some cases, DTW can be used by the automatic speech recognition process to account for different speech (i.e., audible language content) speeds. In some cases, DTW can enable a computing device to find an optimal match between two given sequences (e.g., time series) under certain limitations. That is, in some cases, the sequences can be non-linearly "warped" to match each other. In some cases, a DTW-based sequence alignment method can be used in the context of a hidden Markov model.
[0038] Referring still to FIG. 1, in some embodiments, the automatic speech recognition process may include a neural network. The neural network can include any neural network, such as those disclosed with reference to FIGS. 2-4, for example. In some cases, the neural network can be used for automatic speech recognition, including phoneme classification, phoneme classification by multi-objective evolutionary algorithms, isolated word recognition, audiovisual speech recognition, audiovisual speaker recognition, and speaker adaptation. In some cases, the neural network employed in automatic speech recognition has fewer explicit assumptions regarding feature statistical characteristics than HMM, and thus may have some characteristics that make it an attractive recognition model for speech recognition. When used to estimate the probability of speech feature segments, the neural network can perform discriminative training in a natural and efficient manner. In some cases, the neural network can be used to effectively classify short-time interval audible language content, such as individual phonemes and isolated words, for example. In some embodiments, the neural network can be employed by the automatic speech recognition process for preprocessing, feature transformation, and / or dimensionality reduction, for example, before HMM-based recognition. In some embodiments, long short-term memory (LSTM) and related recurrent neural networks (RNN) as well as time-delay neural networks (TDNN) can be used for automatic speech recognition over longer time intervals, such as for continuous speech recognition, for example.
[0039] Referring still to FIG. 1, in some embodiments, device 100 captures a first image of slide 116 at a first position. In some embodiments, the first image can be captured using at least one optical system 120.
[0040] Referring still to FIG. 1, in some embodiments, capturing the first image of slide 116 at the first position can include using actuator mechanism 124 and / or actuator control to move the optical system 120 and / or the slide 116 to a desired position. In some embodiments, the first image can include an image of the entire sample and / or the entire slide 116. In some embodiments, the first image can include an image of a region of the sample. In some embodiments, the first image includes an image with a wider angle than a second image (described below). In some embodiments, the first image may include an image with a lower resolution than the second image.
[0041] Referring still to FIG. 1, in some embodiments, the apparatus 100 can identify at least one region of interest in the first image. In some embodiments, machine vision may be used to identify the at least one region of interest. As used herein, a "region of interest" is a specific region within a slide or a digital image of a slide where features are detected. Features can include, by way of non-limiting example, a sample, dust, writing on the slide, cracks in the slide, air bubbles, and the like.
[0042] Referring still to FIG. 1, in some embodiments, apparatus 100 may include a machine learning module 144. Machine learning will be described with reference to FIG. 2. In some embodiments, apparatus 100 can use an ROI identification machine learning model 148 to identify at least one region of interest. In some embodiments, ROI identification machine learning model 148 can be trained using supervised learning. In some embodiments, ROI identification machine learning model 148 can include a classifier. ROI identification machine learning model 148 can be trained on a dataset that includes exemplary images of slides, associated with exemplary regions of the image where features are present. Such a training dataset can be collected, for example, by collecting data from a slide imaging device regarding which regions of an image of a slide an expert focuses on or zooms in on. Once trained, ROI identification machine learning model 148 can accept an image of a slide as input and output data regarding the location of the regions of interest present. In some embodiments, a neural network, such as a convolutional neural network, can be used to identify at least one region of interest. For example, a convolutional neural network can be used to detect edges in an image of a slide, and at least one region of interest can be identified based on the presence of the edges. In some embodiments, at least one region of interest can be identified as a function of differences in brightness and / or color in comparison to the brightness and / or color of the background. In some embodiments, apparatus 100 can use a classifier to identify regions of interest. In some embodiments, a segment of an image is input to the classifier, and the classifier can categorize the segment of the image based on whether a region of interest is present. In some embodiments, the classifier can output a score indicating the degree to which a region of interest is detected and / or a confidence level that a region of interest is present. In some embodiments, features can be detected using a neural network or other machine learning model trained to detect features and / or objects.For example, edges, corners, blobs, or ridges can be detected, and whether a location is determined to be within the region of interest can depend on the detection of such elements. In some embodiments, a machine learning model such as support vector machine technology can be used to determine features based on the detection of elements such as edges, corners, blobs, or ridges.
[0043] Still referring to FIG. 1, in some embodiments, the apparatus 100 can identify at least one region of interest as a function of user input. For example, the user can modify the setting of the degree of sensitivity for detecting the region of interest. In this example, if the user input indicates low sensitivity, the apparatus 100 can detect only large regions of interest. This can include, for example, ignoring potential regions of interest that are below a certain size. In another example, this can include applying a machine learning model, such as a classifier, to an image (or a segment of an image) and identifying the image (or the segment of the image) as a region of interest only if the machine learning model outputs a score higher than a threshold. The score indicates the degree to which a region of interest is detected and / or the confidence that a region of interest exists. In some embodiments, the image is divided into smaller segments, and the segments can be analyzed to determine whether at least one region of interest is present.
[0044] Referring still to FIG. 1, in some embodiments, the apparatus 100 may receive at least one region of interest. In some embodiments, the apparatus 100 may receive at least one region of interest without first capturing a first image. In a non-limiting example, a user may input a region of interest. In some embodiments, the apparatus 100 may capture a first image and apply the received region of interest to the first image. This may be done, for example, when the region of interest is received before the first image is captured. In some embodiments, the apparatus 100 may capture the first image as a function of the region of interest. In a non-limiting example, the apparatus 100 may receive a region of interest from a user through a user input and capture the first image at a first position within the region of interest.
[0045] Referring still to FIG. 1, the apparatus 100 may identify a focus pattern. In some embodiments, the apparatus 100 may identify a focus pattern in at least one region of interest, such as each region of interest detected as described herein. In some embodiments, identifying a focus pattern may include identifying a row, identifying a point within the row, determining an optimal focus at the point, and / or identifying a plane.
[0046] Referring still to FIG. 1, as used herein, a "row" of a digital image of a slide is a segment of the digital image of the slide between two parallel lines. A row can include, for example, a row of pixels having a width of one pixel. In another example, a row may have a width of multiple pixels. A row may or may not cross the pixel grid diagonally. As used herein, a "point" on a digital image of a slide refers to a specific position within the digital image of the slide. For example, in a digital image composed of a pixel grid, a point can have a specific (x,y) position. As used herein, an "optimal focus" is the focal distance that focuses on the object to be focused on.
[0047] Referring still to FIG. 1, in some embodiments, the apparatus 100 may identify rows within the region of interest. The rows may be identified based on a first row sample presence score. The first row sample presence score may be identified using machine vision. The first row sample presence score may be identified based on the output of a row identification machine learning model 152. The first row sample presence score may be identified by determining one or more row sample presence scores of adjacent rows. For example, the first row sample presence score may be determined as a function of a second row sample presence score and a third row sample presence score, and the second and third row sample presence scores are based on rows adjacent to the row of the first row sample presence score. As used herein, a "sample presence score" is a value that represents or estimates the likelihood that a sample exists at a location. As used herein, a "row sample presence score" is a sample presence score where the location is a row. The sample presence score need not represent the likelihood of sample presence as a percentage from 0 to 100%. For example, if the row sample presence score for the first row is 2000 and the row sample presence score for the second row is 3000, it is shown that the second row is more likely to contain samples than the first row. The first row sample presence score may be identified, in non-limiting examples, based on the sum of the row sample presence scores of adjacent rows, and / or the weighted sum of the row sample presence scores of adjacent rows, and / or the minimum value of the row sample presence scores of adjacent rows. In some embodiments, a first set of row sample presence scores is identified using a machine learning model such as the row identification machine learning model 152, and a second set of row sample presence scores is identified based on the row sample presence scores of rows adjacent to the first set of row sample presence scores. In some embodiments, such a second set of row sample presence scores may be used to identify the best rows. The row sample presence score may be identified by determining the sample presence scores of rows that are not directly adjacent to the row in question.For example, the row sample presence score may be determined for the row in question, the adjacent rows, and the rows one row away from the row in question, and each of these row sample presence scores may be a factor (e.g., using a weighted sum of the sample presence scores) for identifying the row. The sample presence score may be determined using machine vision. In some embodiments, the row sample presence score may be determined based on a section of the row that does not completely traverse the image and / or slide. For example, the row sample presence score may be determined relative to the width of the region of interest. In another example, the row sample presence score may be determined for a section of the row having a certain pixel width.
[0048] Referring still to FIG. 1, in some embodiments, identifying the focus pattern may include identifying the row that includes a specific (X, Y) position, such as the position where the optimal focus distance is calculated. In some embodiments, identifying the focus pattern may include capturing a plurality of images at such a position, each of the plurality of images having a different focus distance. Such images can represent a Z-stack, as further described below. Identifying the focus pattern can further include determining the optimally focused image from among the plurality of images. The focus distance of such an image may be determined to be the optimal focus distance for its (X, Y) position. In this way, the optimal focus distance may be determined for a plurality of points on the row. The focus pattern may be identified using the focus distances of the plurality of optimally focused images at the plurality of points on such a row.
[0049] Referring still to FIG. 1, in some embodiments, a row may further include a second position. The apparatus 100 can capture a plurality of second images at such a second position, and each of the plurality of second images has a different focal length. The apparatus 100 may determine the second image that is optimally focused among the plurality of second images of the optimal focus. The apparatus 100 may identify a focus pattern using the focal lengths of the optimally focused first image and the optimally focused second image. For example, the focus pattern may be determined to include a line connecting the above two points. The apparatus 100 can extrapolate a third focal length for a third position as a function of the focus pattern. In some embodiments, extrapolating the third focal length can include using the first or second focal length as the third focal length. In some embodiments, extrapolation can include linear extrapolation, polynomial extrapolation, conical extrapolation, geometric extrapolation, etc. In some embodiments, such a third position may be located outside the row including the first position and / or the second position. In some embodiments, such a third position may be located within a region of interest different from the first position and / or the second position.
[0050] Referring still to FIG. 1, in some embodiments, the sample presence score can be determined using the row identification machine learning model 152. In some embodiments, the row identification machine learning model 152 can be trained using supervised learning. The row identification machine learning model can be trained on a dataset including examples of rows related to whether a sample is present or not from an image of a slide. Such a dataset can be collected, for example, by capturing an image of a slide, manually identifying which rows contain samples, and extracting the rows from the large image. Once trained, the row identification machine learning model 152 can accept a row from an image of a slide, such as a row from a region of interest, as an input, and output a determination of whether a sample is present and / or a sample presence score such as a row sample presence score.
[0051] Referring still to FIG. 1, in some embodiments, the apparatus 100 can identify points within a row, such as the rows identified as described above. In some embodiments, the points may be identified using machine vision. In some embodiments, the points can be identified based on points within a particular row that have a maximum point sample presence score. As used herein, a "point sample presence score" is a sample presence score where the location is a point. The point sample presence score can be determined, in non-limiting examples, based on the color of the point and / or surrounding pixels, or whether the point is inside or outside a potential specimen boundary (which can be determined, for example, by identifying edges in the image and regions of the image surrounded by the edges). In another non-limiting example, the point sample presence score can be determined based on the distance between the point and the edge of the slide.
[0052] Referring still to FIG. 1, in some embodiments, the points may be identified using a point identification machine learning model 156. In some embodiments, the point identification machine learning model 156 can be trained using supervised learning. The point identification machine learning model can be trained on a dataset that includes examples of points from an image of the slide that are relevant to whether a sample is present or not. Such a dataset can be collected, for example, by capturing an image of the slide, manually identifying which points contain samples, and extracting the points from the large image. Once trained, the point identification machine learning model 156 can accept points from an image of the slide, such as points from a region of interest, as input and output a determination of whether a sample is present and / or a point sample presence score. In some embodiments, the apparatus 100 can identify points from a row based on which points have the highest point sample presence score.
[0053] Referring still to FIG. 1, in some embodiments, the apparatus 100 can determine an optimal focus at a point, such as the points identified as described above. In some embodiments, the optimal focus may be determined using an autofocus mechanism. In some embodiments, the optimal focus may be determined using a rangefinder. In some embodiments, the actuator mechanism 124 can move the optical sensor 120 and / or the slide port 140 so that the autofocus mechanism focuses on a desired location. In some embodiments, the autofocus mechanism may be able to focus on multiple points within the frame and can select which point to focus on based on the points identified as described above. In some embodiments, one or more camera parameters other than focus can be adjusted to improve focus and / or improve the image. For example, the aperture can be adjusted to change the degree of focus on a point. For example, the aperture can be adjusted to increase the depth of field to focus on a point. The optimal focus can be expressed, in non-limiting examples, as the focal length or the depth of focus.
[0054] Referring still to FIG. 1, in some embodiments, the apparatus 100 may identify a focus pattern based on an optimal focus and / or point. The optimal focus and point may be represented as positions in three-dimensional space. For example, the X and Y coordinates (horizontal axes) may be determined based on the location of a point on the slide and / or the position of a point in the image. The Z coordinate may be determined based on the optimal focus. For example, the focal length of the optimal focus may be used as the Z coordinate. As used herein, a "focus pattern" is a pattern that approximates the optimal focus level at a plurality of points, including at least one point where the optimal focus has not been measured. One or more (X, Y, Z) coordinates may be used to determine the focus pattern. One or more default parameters may be used to determine the focus pattern (e.g., defaulting to horizontal if insufficient data is available to determine otherwise). For example, a single (X, Y, Z) coordinate may be determined, and the focus pattern may be determined as a plane that extends horizontally in the X and Y directions with a constant Z level. In some embodiments, the focus pattern varies in the vertical (Z) direction. For example, two (X, Y, Z) coordinates may be determined, and a focus pattern may be determined that includes a straight line connecting the two (X, Y, Z) positions and extends horizontally when horizontally translated perpendicular to that line (forming a plane that includes that line). In another example, three (X, Y, Z) coordinates may be determined, and the focus pattern may be determined as a plane that includes all three positions. In another example, several (X, Y, Z) coordinates may be determined, and a regression algorithm may be used to determine the plane that best fits the coordinates. For example, least squares regression may be used. In some embodiments, the focus pattern is not planar. The focus pattern may include a surface, such as the surface of a three-dimensional space. The focus pattern may include one or more curves, bumps, edges, etc. For example, the focus pattern may include a first plane in a first region, a second plane in a second region, and an edge where the planes intersect. In another example, the focus pattern may include a curved and / or bumpy surface where the Z level at each position on the focus pattern surface is determined based on the neighboring (X, Y, Z) coordinates and the distance to their Z levels.In another example, the focus pattern may include a plurality of shapes having (X, Y, Z) coordinates as vertices and lines between the (X, Y, Z) coordinates as boundaries. In some embodiments, the focus pattern can have only one Z value for each (X, Y) coordinate. In some embodiments, which points are evaluated for their optimal focus can be determined, in non-limiting examples, as a function of the desired density of points within the region of interest, the likelihood that a sample exists (such as the output from the ML model described above), and / or user input. For example, the user can manually select points and / or input the desired point density. In some embodiments, the focus pattern may be updated as additional points are scanned. For example, the focus pattern may be in the shape of a plane based on 10 (X, Y, Z) coordinates, and when the 11th (X, Y, Z) coordinate is scanned, the focus pattern may be recalculated and / or updated taking into account the new coordinate. In another example, the focus pattern may be started as a plane based on a single (X, Y, Z) coordinate and updated as additional (X, Y, Z) coordinates are identified. In some embodiments, the focus pattern may be updated for each additional (X, Y, Z) coordinate. In some embodiments, the focus pattern may be updated at a rate less than the rate of (X, Y, Z) coordinate identification. In non-limiting examples, the focus pattern may be updated for every 2, 3, 4, 5, or more (X, Y, Z) coordinates.
[0055] Still referring to FIG. 1, in some embodiments, the data used to identify the focus pattern may be filtered. In some embodiments, one or more outliers may be removed. For example, if almost all (X, Y, Z) points suggest a focus pattern in the shape of a plane and a single (X, Y, Z) point has a Z value that is significantly different from what is estimated from the plane, that (X, Y, Z) point may be deleted. In another example, (X, Y, Z) points identified as being in focus on features other than the sample may be deleted. For example, (X, Y, Z) points identified as being in focus on an annotation may be deleted.
[0056] Referring still to FIG. 1, in some embodiments, the process described above can be used to determine a focus pattern for each region of interest. In a region of interest that includes a sample, this results in determining a focus pattern based on one or more points that include the sample. This can help to efficiently identify the focus pattern so that a follow-up image is captured at the correct focal distance. In some regions of interest, such as a region of interest that does not include a sample but instead includes features such as annotations, there may be a risk of focusing on non-sample features. This may be desirable, for example, because capturing a focused image of a feature such as an annotation can help an expert to read the annotation and / or assist an optical character recognition process when transcribing the text. In some regions of interest, both samples and non-sample features may be present. In this case, it is desirable to focus on the sample, and this can be achieved by the process described herein. In some embodiments, the process described herein can provide a more efficient method of identifying a focus pattern than alternatives. For example, the process described herein may require focusing on fewer points to determine the focus pattern.
[0057] Referring still to FIG. 1, in some embodiments, a focus pattern, such as a plane, can be identified as a function of a first image and a first position. The (X, Y) position of a point can be determined from the first position. The optimal focus value can be determined from the first image. Together, these can be used to identify the (X, Y, Z) coordinates that can be used to identify a focus pattern as described herein.
[0058] Referring still to FIG. 1, in some embodiments, a focus pattern, such as the Z-level of the optimal focus, can be used to scan the remainder of the row. For example, it can be used to scan the remaining rows including the point where the optimal focus is identified. In some embodiments, the focus pattern can be used to scan additional rows. For example, a focus pattern determined as a function of the Z-level of a point can be used to scan the rows adjacent to the row containing the point. In some embodiments, the focus pattern can be used to scan other rows within the same region of interest. In some embodiments, the focus pattern can be used to scan rows within other regions of interest, such as neighboring regions of interest.
[0059] Referring still to FIG. 1, in some embodiments, the identification of the region of interest, the identification of the point, the identification of the point, and / or the determination of the focus pattern may be performed locally. For example, device 100 may include a machine learning model that has already been trained and may apply that model to the image. In some embodiments, the identification of the region of interest, the identification of the point, the identification of the point, and / or the determination of the focus pattern may be performed externally. For example, device 100 may transmit the image data to another computing device and may receive the output described herein. In some embodiments, the region of interest may be identified, the rows may be identified, the points may be identified, and / or the focus pattern may be determined in real time.
[0060] Referring still to FIG. 1, in some embodiments, device 100 can determine a scanning pattern. The scanning pattern may be based, for example, on the form of the sample. The scanning pattern may include, in non-limiting examples, zigzag, snake line, and spiral.
[0061] Referring still to FIG. 1, in some embodiments, a snake pattern may be used to scan the slide. The snake pattern may proceed in any horizontal direction. In a non-limiting example, the snake pattern may proceed across the length or width of the slide. In some embodiments, the direction in which the snake pattern proceeds may be selected to minimize the number of turns required. For example, the snake pattern may proceed along the shortest dimension of the slide. In another example, the shape of the sample may be identified using, for example, machine vision, and the snake pattern may proceed in a direction according to the shortest dimension of the sample. In some embodiments, the snake pattern may be selected when high-speed scanning is desired. In some embodiments, the snake pattern can minimize the movement of the camera and / or slide required to scan the slide. In some embodiments, a zigzag pattern may be used to scan the slide. Similar to that described for snake pattern scanning, zigzag pattern scanning can be performed in any horizontal direction, and the direction may be selected to minimize the number of turns and / or rows used to scan features such as the slide and / or sample. In some embodiments, a spiral pattern may be used to scan the slide, features, regions of interest, etc. In some embodiments, the snake pattern and / or zigzag pattern may be optimal for the Z-direction movement from one row to the next. In some embodiments, the spiral pattern may be optimal for dynamic grid region of interest expansion.
[0062] Referring still to FIG. 1, in some embodiments, the apparatus 100 can extrapolate the focal distance at a second position as a function of the focus pattern. In some embodiments, the focus pattern can be determined as a function of one or more (X, Y, Z) points local to the region of interest and / or sub-regions of the region of interest. In some embodiments, a focus pattern, such as a plane, can be used to approximate the optimal focus level using extrapolation (as opposed to interpolation). For example, the focus pattern can be used to approximate the optimal focus level at (X, Y) points outside the range of already scanned (X, Y) points, such as outside the range of X values, outside the range of Y values, or outside the range of a shape that includes already scanned (X, Y) points. In another example, the focus pattern may be used to approximate the optimal focus level, and the optimal focus level can include a Z value outside the range of Z values used to determine the focus pattern. In another example, the (X, Y, Z) points can be extrapolated to the focus pattern across a row. In some embodiments, one or more additional (X, Y, Z) points may be used to update the focus pattern. In some embodiments, the focus pattern identified for one row can be extrapolated to another row, such as an adjacent row. In another example, a local focus pattern, such as a plane, can be extrapolated to identify the optimal focus level outside the local region. In another example, the focus pattern in the first region of interest can be extrapolated to generate the focus pattern in the second region of interest and / or to identify the optimal focus level at the (X, Y) points in the second region of interest.
[0063] Referring still to FIG. 1, in some embodiments, the apparatus 100 can capture a second image of the slide at a second position with a focal distance based on a focus pattern. For example, if the focus pattern is a plane and a focused image of a particular (X, Y) point is desired, the apparatus 100 can capture an image using a focal distance based on the Z coordinate of the plane at those (X, Y) coordinates. In some embodiments, the first position (such as the position where the optimal focus is measured) and the second position can be set to image positions within the same region of interest. In some embodiments, the actuator mechanism may be mechanically connected to the movable element. Also, the actuator mechanism may move the movable element to the second position. In some embodiments, obtaining the second image may include capturing a plurality of images taken at a focal distance based on a focus pattern and constructing the second image from the plurality of images. In some embodiments, the second image may include an image taken as part of a Z-stack. The Z-stack will be described later.
[0064] Referring still to FIG. 1, in some embodiments, the apparatus 100 may determine which regions of interest contain a sample. In some embodiments, this may be applied to a second image, such as a second image taken using a focus distance based on a focus pattern. In some embodiments, a sample identification machine learning model 160 may be used to determine which regions of interest contain a sample. In some embodiments, the sample identification machine learning model 160 may include a classifier. In some embodiments, the sample identification machine learning model 160 may be trained using supervised learning. The sample identification machine learning model 160 may be trained on a dataset that includes exemplary images of slides and / or segments of images of slides that are relevant to whether a sample is present. Such a dataset may be collected, for example, by capturing images of slides and manually identifying those that contain a sample. In some embodiments, multiple machine learning models may be trained to identify different types of samples. Once trained, the sample identification machine learning model 160 may receive an image of a region of interest as input and output a determination as to whether a sample is present. In some embodiments, the sample identification machine learning model 160 may be improved through, for example, the use of reinforcement learning. The feedback used to determine the cost function of the reinforcement learning model may include, for example, user input or annotations on the slide. For example, if an annotation transcribed using optical character recognition indicates a particular type of sample and the sample identification machine learning model 160 indicates that no sample is included in the region of interest, a cost function indicating that the output is false may be determined. In another example, the user may input a label associated with the region of interest. If the label indicates a particular type of sample and the output of the sample identification machine learning model 160 indicates that a sample is present in the region of interest, a cost function indicating that the output is correct may be determined.
[0065] Referring still to FIG. 1, in some embodiments, whether a sample is present may be determined locally. For example, the apparatus 100 may include a machine learning model 160 for sample identification that has already been trained, and the model can be applied to an image or a segment of an image. In some embodiments, whether a sample is present may be determined externally. For example, the apparatus 100 may transmit the image data to another computing device and receive a determination as to whether a sample is present. In some embodiments, whether a sample is present may be determined in real time.
[0066] Referring still to FIG. 1, in some embodiments, a machine vision system and / or an optical character recognition system may be used to determine one or more features of the sample and / or the slide 116. In a non-limiting example, an optical character recognition system may be used to identify writing on the slide 116, which may be used to annotate an image of the slide 116.
[0067] Referring still to FIG. 1, in some embodiments, the apparatus 100 can capture multiple images at different sample depths of focus. As used herein, "sample depth of focus" is the depth within the sample at which the optical system is in focus. As used herein, "focal length" is the focal length on the object side. In some embodiments, the first image and the second image may have different focal lengths and / or sample depths of focus.
[0068] Referring still to FIG. 1, in some embodiments, the apparatus 100 can include a machine vision system. In some embodiments, the machine vision system may include at least one camera. The machine vision system can use images, such as images from at least one camera, to make determinations regarding a scene, space, and / or object. For example, in some cases, the machine vision system can be used for world modeling and alignment of objects in space. In some cases, the alignment may include image processing, such as, but not limited to, object recognition, feature detection, edge / corner detection, etc. Non-limiting examples of feature detection include Scale-Invariant Feature Transform (SIFT), Canny edge detection, Shi Tomasi corner detection, etc. In some cases, the alignment may include one or more transformations that orient a camera frame (or image or video stream) with respect to a three-dimensional coordinate system. Exemplary transformations include, but are not limited to, homography transformation and affine transformation. In one embodiment, the alignment of the first frame with respect to the coordinate system can be verified and / or corrected using object identification and / or computer vision as described above. For example, but not limited to, the initial alignment to two dimensions, represented as alignment to the x and y coordinates, may however be performed using the two-dimensional projection of three-dimensional points onto the first frame. The third dimension of the alignment, representing depth and / or the z-axis, can be detected by comparing the two frames. For example, if the first frame includes a pair of frames captured using a pair of cameras (also referred to as a stereo camera in the present disclosure), the stereo pair of the image of the object can be detected using image recognition and / or edge detection software. The two stereo pairs are compared to derive the z-axis value of points on the object, enabling the derivation of additional z-axis points within and / or around the object, for example, using interpolation. This may be repeated for multiple objects within the field of view, including, but not limited to, environmental features of interest identified by an object classifier and / or indicated by an operator.In one embodiment, the x-axis and the y-axis may be selected to span a plane common to two cameras used to capture a stereoscopic image and / or the xy-plane of the first frame, such that the translational components of x and y and φ may be pre-input into the translation matrix and the rotation matrix, as described above, for the affine transformation of the object coordinates. As described above, the initial x and y coordinates and / or the estimation of the transformation matrix may alternatively or additionally be performed between the first frame and the second frame. As described above, for each point of the object and / or edge and / or a plurality of points on the edge of the object, the x and y coordinates of the first stereoscopic frame may be input, along with an initial estimated value of the z coordinate, based on an assumption about the object, such as the assumption that the ground is substantially parallel to the xy-plane, as selected above. As described above, the Z coordinate and / or the x, y, and z coordinates aligned using the image capture and / or object identification process may then be compared with the coordinates predicted using the initial estimate of the transformation matrix, and an error function may be used and calculated by comparing the two sets of points with the new x, y, and / or z coordinates, and may be repeatedly estimated and compared until the error function falls below a threshold level. In some cases, the machine vision system may be able to use a classifier, such as any of the classifiers described throughout the present disclosure.
[0069] Referring still to FIG. 1, in some embodiments, the image data may be processed using optical character recognition. In some embodiments, optical character recognition or an optical character reader (OCR) includes automatically converting an image of written (e.g., typed, handwritten, or printed) text into machine-encoded text. In some cases, recognizing at least one keyword from the image data can include one or more processes including, but not limited to, optical character recognition (OCR), optical word recognition, intelligent character recognition, intelligent word recognition, etc. In some cases, the OCR may recognize the written text one glyph or one character at a time. In some cases, optical word recognition can recognize the written text one word at a time, for example, in a language that uses a space as a word delimiter. In some cases, intelligent character recognition (ICR) can recognize the written text one glyph or one character at a time, for example, by employing a machine learning process. In some cases, intelligent word recognition (IWR) can recognize the written text one word at a time, for example, by employing a machine learning process.
[0070] Referring still to FIG. 1, in some cases, the OCR may be an "offline" process that analyzes static documents or image frames. In some cases, handwriting motion analysis can be used as input for handwriting recognition. For example, this technique can capture not only the shape of the glyphs or words, but also the order in which segments are drawn, the direction, and the pattern of placing and lifting the pen. This additional information can potentially make the handwriting recognition more accurate. In some cases, this technique is also called "online" character recognition, dynamic character recognition, real-time character recognition, intelligent character recognition.
[0071] Referring still to FIG. 1, in some cases, the OCR process can employ preprocessing of the image data. The preprocessing process can include, but is not limited to, skew correction, despeckling, binarization, line removal, layout analysis or "zoning", line and word detection, script recognition, character separation or "segmentation", and normalization. In some cases, the skew correction process can include applying a transformation (e.g., a homography or an affine transformation) to the image data to align the text. In some cases, the despeckling process can include removing positive and negative spots and / or smoothing the edges. In some cases, the binarization process can include converting the image from color or grayscale to black and white (i.e., a binary image). Binarization can be performed as a simple way to separate text (or any other desired image component) from the background of the image data. In some cases, binarization may be required, for example, if the OCR algorithm employed only supports binary images. In some cases, the line removal process can include removing images other than glyphs and characters (e.g., boxes and lines). In some cases, the layout analysis or "zoning" process can identify columns, paragraphs, captions, etc. as separate blocks. In some cases, the line and word detection process can establish reference values for the shapes of words and characters, and words can be separated as needed. In some cases, the script recognition process can identify the script, for example, in a multilingual document, and enable the selection of an appropriate OCR algorithm. In some cases, the character separation or "segmentation" process can separate signal characters, for example, with a character-based OCR algorithm. In some cases, the normalization process can normalize the aspect ratio and / or scale of the image data.
[0072] Referring still to FIG. 1, in some embodiments, the OCR process can include an OCR algorithm. Exemplary OCR algorithms can include matrix matching processing and / or feature extraction processing. Matrix matching can include comparing an image with stored glyphs on a pixel-by-pixel basis. In some cases, matrix matching is also known as "pattern matching", "pattern recognition", and / or "image correlation". Matrix matching may depend on whether the input glyph is correctly separated from the rest of the image data. Matrix matching may also depend on the stored glyph being the same font and the same scale as the input glyph. Matrix matching may work best with typed text.
[0073] Referring still to FIG. 1, in some embodiments, the OCR process can include a feature extraction process. In some cases, feature extraction can break glyphs into at least one feature. Exemplary non-limiting features can include corners, edges, lines, closed loops, line directions, line intersections, and the like. In some cases, feature extraction can reduce the dimensionality of the representation and make the recognition process more efficient computationally. In some cases, the extracted features are compared to an abstract vector-like representation of the characters and reduced to one or more glyph prototypes. General techniques for feature detection in computer vision are applicable to this type of OCR. In some embodiments, a machine learning process such as a nearest neighbor classifier (e.g., k-nearest neighbor algorithm) can be used to compare the image features to stored glyph features and select the closest match. The OCR can employ any machine learning process described in the present disclosure, such as the machine learning process described with reference to FIGS. 2-4, for example. Exemplary non-limiting OCR software includes Cuneiform and Tesseract. Cuneiform is a multi-language, open-source optical character recognition system originally developed by Cognitive Technologies of Moscow, Russia. Tesseract is free OCR software originally developed by Hewlett-Packard of Palo Alto, California, USA.
[0074] Referring still to FIG. 1, in some cases, OCR may employ a two-pass approach for character recognition. In the first pass, an attempt may be made to recognize characters. Each well-recognized character is passed to an adaptive classifier as training data. The adaptive classifier obtains a chance to more accurately recognize characters by further analyzing the image data. Since the adaptive classifier may have learned something useful at a timing that is a bit too late to recognize characters in the first pass, a second pass is performed on the image data. The second pass may include adaptive recognition, and characters recognized with high confidence in the first pass can be used to better recognize the remaining characters in the second pass. In some cases, the two-pass approach may be advantageous for special fonts or low-quality image data. Another exemplary OCR software tool is OCRopus. The development of OCRopus is led by the German Research Center for Artificial Intelligence (DFKI) in Kaiserslautern, Germany. In some cases, OCR software may also employ a neural network.
[0075] Referring still to FIG. 1, in some cases, OCR can include post - processing. For example, the accuracy of OCR can be improved in cases where the output is constrained by a lexicon. The lexicon may include a list or set of words that are allowed to appear in the document. In some cases, the lexicon may include, for example, all words in English or a more specialized lexicon for a particular field. In some cases, the output stream may be a plain - text stream or a file of characters. In some cases, the OCR process can retain the original layout of the image data. In some cases, near - neighbor analysis can use co - occurrence frequencies and focus on the fact that certain words are frequently seen together to correct errors. For example, "Washington, D.C." is much more common in English than "Washington DOC." In some cases, the OCR process can utilize prior knowledge about the grammar of the recognized language. For example, grammar rules may be used to determine whether a word is a verb or a noun. For recognition and classification, the concept of distance can be adopted. For example, the Levenshtein distance algorithm can be used in the post - processing of OCR to further optimize the results.
[0076] Referring still to FIG. 1, in some embodiments, the apparatus 100 can remove artifacts from the image. As used herein, an "artifact" is a visual inaccuracy, an image element that distracts from the element of interest, an image element that obscures the element of interest, or other undesirable elements of the image.
[0077] Referring still to FIG. 1, apparatus 100 may include an image processing module. As used in this disclosure, an "image processing module" is a component designed to process digital images. In one embodiment, the image processing module may include a plurality of software algorithms, such as, but not limited to, a plurality of image processing techniques as described hereinafter, that can analyze, manipulate, or otherwise improve an image. In another embodiment, the image processing module may include hardware components, such as, but not limited to, one or more graphics processing units (GPUs) that can speed up the processing of a large number of images. In some cases, the image processing module may be implemented using one or more image processing libraries, such as, but not limited to, OpenCV, PIL / Pillow, ImageMagick.
[0078] Referring still to FIG. 1, the image processing module may be configured to receive images from the optical sensor 120. One or more images may be transmitted from the optical sensor 120 to the image processing module via any suitable electronic communication protocol, including but not limited to packet-based protocols such as Transmission Control Protocol / Internet Protocol (TCP-IP), File Transfer Protocol (FTP). Receiving an image may also include obtaining the image from a data store that includes the image, as described hereinafter. For example, but not limited to, an image may be obtained using a query that specifies a timestamp for which the image is required to match.
[0079] Referring still to FIG. 1, the image processing module may be configured to process images. In one embodiment, the image processing module may be configured to compress and / or encode an image to reduce file size and storage requirements while maintaining the essential visual information necessary for further processing steps as described hereinafter. In one embodiment, the compression and / or encoding of the image may facilitate the fast transmission of the image. In some cases, the image processing module may be configured to perform lossless compression on the image, and lossless compression can maintain the original image quality. In a non-limiting example, the image processing module may utilize one or more lossless compression algorithms such as, but not limited to, Huffman coding, Lempel-Ziv-Welch (LZW), Run-Length Encoding (RLE), etc. to identify and remove the redundancy of the image without losing information. In such an embodiment, compressing and / or encoding each image of the image may include converting the file format of each image to PNG, GIF, lossless JPEG2000, etc. In one embodiment, an image compressed via lossless compression can be completely reconstructed into its original form (e.g., the resolution, dimensions, color representation, format, etc. of the original image). In another case, the image processing module may be configured to perform lossy compression on the image, and lossy compression may sacrifice some image quality to achieve a higher compression ratio. In a non-limiting example, the image processing module may utilize one or more lossy compression algorithms such as, but not limited to, the discrete cosine transform (DCT) of JPEG or the wavelet transform of JPEG2000 to discard less important information in the image, and as a result, the file size is reduced, but the degradation of the image quality is slight. In such an embodiment, the compression and / or encoding of the image may include converting the file format of each image to JPEG, WebP, lossy JPEG2000, etc.
[0080] Referring still to FIG. 1, in one embodiment, processing an image can include determining a degree of depiction quality of a region of interest of the image. In one embodiment, the image processing module can determine the blurriness of the image. In a non-limiting example, the image processing module can perform blur detection by taking an approximation such as a Fourier transform of the image, or a fast Fourier transform (FFT), and analyzing the distribution of low and high frequencies in the depiction of the resulting frequency domain of the image. For example, without limitation, the number of high frequency values below a threshold level may indicate blurriness. In another non-limiting example, blur detection may be performed by convolving the image, an image channel, etc. with a Laplacian kernel, which may generate, for example without limitation, a numerical score that reflects the number of sharp changes in intensity shown in each image such that a high score indicates sharpness and a low score indicates blurriness. In some cases, blur detection can be performed using a gradient-based operator that measures an operator based on the gradient or first derivative of the image, based on the hypothesis that sharp changes indicate sharp edges in the image and thus a low degree of blurriness. In some cases, blur detection may be performed using a wavelet-based operator that utilizes the ability of the coefficients of the discrete wavelet transform to describe the frequency and spatial content of the image. In some cases, blur detection may be performed using a statistic-based operator that utilizes some image statistics as texture descriptors to calculate a focus level. In another case, blur detection may be performed by using discrete cosine transform (DCT) coefficients to calculate the focus level of the image from the frequency content. Additionally or alternatively, the image processing module can be configured to rank the images according to the degree of depiction quality of the region of interest and select the highest-ranked image from a plurality of images.
[0081] Referring still to FIG. 1, processing the image can include enhancing the image or at least one region of interest by a plurality of image processing techniques to improve the quality of the image (or the degree of the quality of the depiction) for better processing and analysis, as further described in the present disclosure. In one embodiment, the image processing module may be configured to perform a noise reduction operation on the image, and the noise reduction operation can remove or minimize noise (resulting from various causes such as sensor limitations, poor lighting conditions, image compression, etc.), and as a result, a cleaner and visually consistent image can be obtained. In some cases, the noise reduction operation may be performed using one or more image filters. For example, but not limited to, the noise reduction operation can include Gaussian filtering, median filtering, bilateral filtering, etc. The noise reduction process can be performed by the image processing module by averaging or filtering the pixel values in the vicinity of each pixel of the image to reduce random variations.
[0082] Referring still to FIG. 1, in another embodiment, the image processing module may be configured to perform a contrast enhancement operation on the image. In some cases, the image may exhibit low contrast, for example, the features may be difficult to distinguish from the background. By performing a contrast enhancement operation, the contrast of the image can be improved by expanding the intensity range of the image and / or redistributing the intensity values (i.e., the degree of light and dark of the pixels within the image). In a non-limiting example, the intensity value represents the gray level or color of each pixel and can scale from 0 to 255 for an 8-bit image and from 0 to 16,777,215 for a 24-bit color image. In some cases, the contrast enhancement operation can include, but is not limited to, histogram equalization, adaptive histogram equalization (CLAHE), contrast stretching, etc. The image processing module can be configured to adjust the light and dark levels within the image to make the features more distinguishable (i.e., to enhance the degree of depiction quality). Additionally or alternatively, the image processing module may be configured to perform a brightness normalization operation to correct for variations in lighting conditions (i.e., non-uniform brightness levels). In some cases, the image can include a consistent brightness level across the entire region after the brightness normalization operation performed by the image processing module. In a non-limiting example, the image processing module may perform global normalization or local mean normalization, and the average intensity value of the entire image or a region of the image may be calculated and used to adjust the brightness level.
[0083] Referring still to FIG. 1, in other embodiments, the image processing module may be configured to perform a color space conversion operation to enhance the degree of rendering quality. In a non-limiting example, in the case of a color image (i.e., an RGB image), the image processing module may be configured to convert the RGB image to a grayscale or HSV color space. Such a conversion may emphasize the difference in intensity values between the region or feature of interest and the background. The image processing module may be further configured to perform image sharpening operations such as, but not limited to, unsharp masking, Laplacian sharpening, high-pass filtering, etc. The image processing module can use image sharpening operations to emphasize edges and fine details associated with regions or features of interest within the image by emphasizing high-frequency components within the image.
[0084] Referring still to FIG. 1, processing an image may include separating an area of interest or features from the rest of the image as a function of multiple image processing techniques. The image may include the highest ranked image selected by the image processing module as described above. In one embodiment, the multiple image processing techniques may include one or more morphological operations, which are techniques developed based on random functions used to process geometric structures using set theory, lattice theory, topology, and structuring elements. For the purposes of the present disclosure, a "structuring element" is a small matrix or kernel that defines the shape and size of a morphological operation. In some cases, the structuring element may be placed at the center of each pixel of the image and used to determine the output pixel value at that location. In a non-limiting example, separating an area of interest or features from an image may include applying a dilation operation, which is a basic morphological operation configured to expand or grow the boundaries of objects (e.g., cells, dust particles, etc.) within the image. In another non-limiting example, separating an area of interest or features from an image may include applying an erosion operation, which is a basic morphological operation configured to shrink or contract the boundaries of objects within the image. In another non-limiting example, separating an area of interest or features from an image may include applying an opening operation, which is a basic morphological operation configured to remove small objects or thin structures from the image while retaining larger structures. In a further non-limiting example, separating an area of interest or features from an image may include applying a closing operation, which is a basic morphological operation configured to fill small gaps or holes in objects within the image while retaining the overall shape and size of the objects. These morphological operations can be performed by the image processing module to enhance object edges, remove noise, or fill gaps in areas of interest or features prior to further processing.
[0085] Referring still to FIG. 1, in one embodiment, separating a region of interest or features from an image can include using an edge detection technique that can detect one or more shapes defined by edges. The "edge detection technique" used in the present disclosure includes a mathematical method for identifying points in a digital image where the brightness of the image changes abruptly and / or becomes discontinuous. In one embodiment, such points can be organized into line segments of straight lines and / or curves called "edges". The edge detection technique can be performed by an image processing module using any suitable edge detection algorithm, including but not limited to Canny edge detection, Sobel operator edge detection, Prewitt operator edge detection, Laplacian operator edge detection, and / or differential edge detection. The edge detection technique may include edge detection based on phase congruency, which finds all positions in the image where all sine waves in the frequency domain generated using, for example, Fourier decomposition may have a matching phase indicating the position of the edge. The edge detection technique can be used to detect the shape of features of interest, such as cells, that indicate cell membranes or cell walls. In one embodiment, the edge detection technique can be used to find closed figures formed by edges.
[0086] Referring still to FIG. 1, in a non-limiting example, separating the feature of interest from the image can include determining the feature of interest by edge detection techniques. The feature of interest can include a specific region within the digital image that contains information related to further processing as described later. In a non-limiting example, the image data located outside the feature of interest can include irrelevant or redundant information. The image portion containing irrelevant or redundant information can be ignored by the image processing module, thereby concentrating resources on the feature of interest. In some cases, the feature of interest may have different sizes, shapes, and / or positions within the image. In a non-limiting example, the feature of interest may be shown as a circle surrounding the cell nucleus. In some cases, the feature of interest can specify one or more coordinates and distances, such as the center and radius of the circle surrounding the cell nucleus in the image. Then, the image processing module can be configured to separate the feature of interest from the image based on the feature of interest. In a non-limiting example, the image processing module can crop the image according to a bounding box surrounding the feature of interest.
[0087] Referring still to FIG. 1, the image processing module may be configured to perform connected component analysis (CCA) on an image for separating the features of interest. "Connected component analysis (CCA)" used in the present disclosure, also known as connected component labeling, is an image processing technique used to identify and label connected regions within a binary image (i.e., an image in which each pixel has only two possible values: 0 or 1, black or white, or foreground or background). The "connected region" described herein is a group of adjacent pixels that share the same value and are connected based on a predefined neighborhood system such as a 4-connected neighborhood or an 8-connected neighborhood, although not limited thereto. In some cases, the image processing module can convert the image into a binary image by thresholding, which can include setting a threshold for separating the pixels of the image corresponding to the features of interest (foreground) from the pixels corresponding to the background. Pixels having an intensity value above the threshold can be set to 1 (white), and pixels below the threshold can be set to 0 (black). In one embodiment, CCA can be employed to detect and extract the features of interest by identifying a plurality of connected regions that exhibit specific properties or features of the features of interest. The image processing module can then filter the plurality of connected regions by analyzing the properties of the plurality of connected regions, such as, but not limited to, area, aspect ratio, height, width, perimeter, etc. In a non-limiting example, the connected components that closely resemble the dimensions and aspect ratio of the features of interest may be retained as the features of interest by the image processing module, and other components may be discarded. The image processing module can be further configured to extract the features of interest from the image for further processing as described later. Referring still to FIG. 1, in one embodiment, separating the features of interest from the image can include dividing the region depicting the features of interest into a plurality of sub-regions.
[0088] Dividing the region into sub-regions can include dividing the region as a function of the feature of interest and / or CCA via an image segmentation process. As used in this disclosure, "image segmentation process" is a process of dividing a digital image into one or more segments, where each segment represents a different part of the image. The image segmentation process can change the representation of the image. The image segmentation process can be executed by an image processing module. In a non-limiting example, the image processing module can perform region-based segmentation, which includes growing regions from one or more seed points or pixels on the image based on similarity criteria. The similarity criteria can include, but are not limited to, color, intensity, texture, etc. In a non-limiting example, region-based segmentation can include region growing, region merging, watershed algorithms, etc.
[0089] Still referring to FIG. 1, in some embodiments, the apparatus 100 can remove artifacts identified by the machine vision system or the optical character recognition system described above. Non-limiting examples of artifacts that can be removed include dust particles, air bubbles, cracks in the slide 116, writings on the slide 116, shadows, visual noise such as a grainy image, etc. In some embodiments, the artifacts can be partially removed and / or the visibility can be reduced.
[0090] Referring still to FIG. 1, in some embodiments, artifacts can be removed using an artifact removal machine learning model. In some embodiments, the artifact removal machine learning model can be trained on a dataset including images associated with images without artifacts. In some embodiments, the artifact removal machine learning model can accept an image including artifacts as input and output an image without artifacts. For example, the artifact removal machine learning model can accept an image including bubbles in a slide as input and output an image without bubbles. In some embodiments, the artifact removal machine learning model can include a generative machine learning model such as a diffusion model. The diffusion model can learn the structure of the dataset by modeling how data points diffuse in the latent space. In some embodiments, artifact removal can be performed locally. For example, device 100 can include a pre-trained artifact removal machine learning model and apply the model to an image. In some embodiments, artifact removal can be performed externally. For example, device 100 can transmit image data to another computing device and receive an image with artifacts removed. In some embodiments, artifacts can be removed in real time. In some embodiments, artifacts can be removed based on identification by a user. For example, the user can use a mouse cursor to drag a box around an artifact, and device 100 can remove the artifact within the box.
[0091] Referring still to FIG. 1, in some embodiments, the apparatus 100 can display an image to the user. In some embodiments, the first image can be displayed to the user in real time. In some embodiments, the image can be displayed to the user using the output interface 132. For example, the first image can be displayed on a display such as a screen. In some embodiments, the first image can be displayed to the user in the context of a graphical user interface (GUI). For example, the GUI can include controls for navigating the image, such as zoom in / out controls and controls for changing the location being displayed. The GUI can include a touch screen.
[0092] Referring still to FIG. 1, in some embodiments, the apparatus 100 can receive a parameter set from a user. As used herein, a "parameter set" is a set of values that identify how an image is to be captured. The parameter set can be implemented as a data structure as described hereinafter. In some embodiments, the apparatus 100 can receive a parameter set from the user using the input interface 128. The parameter set can include X and Y coordinates indicating where the user wants to look. The parameter set can include a desired magnification level. As used herein, a "magnification level" is a data item indicating how much an image is to be magnified / reduced for capture. The magnification level can take into account optical zoom and / or digital zoom. As a non-limiting example, the magnification level can be "8x". The parameter set may include the depth of focus and / or focal length of the desired sample. In a non-limiting example, the user can operate the input interface 128 such that the parameter set includes the X and Y coordinates and a magnification level corresponding to a more magnified view of a particular region of the sample. In some embodiments, the parameter set corresponds to a more magnified view of a particular region of the sample included in the first image. This can be done, for example, to obtain a more detailed view of a small object. As used herein, unless otherwise specified, "X coordinate" and "Y coordinate" refer to coordinates along the vertical axis, and the plane defined by these axes is parallel to the plane of the surface of the slide 116. In some cases, setting the magnification may include changing one or more optical elements within the optical system. For example, setting the magnification may include replacing the first objective lens with a second objective lens having a different magnification. Further, by replacing one or more optical components that "down beam" from the objective lens, the overall magnification of the optical system can be changed and the magnification can be set. In some cases, setting the magnification may include changing the digital magnification. Digital zoom includes outputting an image at a different resolution, i.e., after resizing the image, using the output interface.In some embodiments, the apparatus 100 can capture an image having a specific field of view. The field of view can be determined, for example, based on the level of magnification or how wide an angle the camera uses to take the image.
[0093] Referring still to FIG. 1, in some embodiments, the apparatus 100 can move one or more of the slide port 140, the slide 116, and the at least one optical system 120 to a second position. In some embodiments, the location of the second position may be based on a parameter set. In some embodiments, the location of the second position may be based on the identification of a region of interest, a row, or a point, as described above. The second position can be determined, for example, by changing the position of the optical system 120 relative to the slide 116 based on a parameter set. For example, the parameter set can indicate that the second position is achieved by changing the X coordinate 5 mm in a particular direction. In this example, the second position can be obtained by changing the original position of the optical system 5 mm in that direction. In some embodiments, such movement can be performed using an actuator mechanism 124. In some embodiments, the actuator mechanism 124 can move the slide port 140 so that the slide 116 is in a position relative to the at least one optical system 120 such that the optical sensor 120 can capture an image as directed by the parameter set. For example, the slide 116 can rest on the slide port 140, and the movement of the slide port 140 can move the slide 116. In some embodiments, the actuator mechanism 124 can move the slide 116 so that the slide 116 is in a position relative to the at least one optical system 120 such that the optical sensor 120 can capture an image as directed by the parameter set. For example, the slide 116 can be connected to the actuator mechanism 124 such that the actuator mechanism 124 can move the slide 116 relative to the at least one optical system 120. In some embodiments, the actuator mechanism 124 can move the at least one optical system 120 so that the slide 116 is in a position relative to the slide 116 such that the optical system 120 can capture an image as directed by the parameter set.For example, slide 116 may be stationary, and actuator mechanism 124 can move at least one optical system 120 to a position relative to slide 116. In some embodiments, actuator mechanism 124 can move a plurality of slide port 140, slide 116, and at least one optical system 120 so that they are in the correct relative positions. In some embodiments, actuator mechanism 124 can move slide port 140, slide 116, and / or at least one optical system 120 in real time. For example, user input of a parameter set can cause substantially immediate movement of items by actuator mechanism 124.
[0094] Still referring to FIG. 1, in some embodiments, apparatus 100 may capture a second image of slide 116 at a second position. In some embodiments, apparatus 100 can capture the second image using at least one optical system 120. In some embodiments, the second image can include an image of a region of the sample. In some embodiments, the second image can include an image of a region photographed in the first image. For example, the second image can include an image with a higher resolution per unit area that magnifies the region within the first image. In another example, the second image can be captured using a focal length based on a focus pattern. Thereby, the second image is displayed so that the user can detect smaller details within the imaged region.
[0095] Still referring to FIG. 1, in some embodiments, the second image includes a shift in X and Y coordinates relative to the first image. For example, the second image can partially overlap the first image.
[0096] Referring still to FIG. 1, in some embodiments, the apparatus 100 may capture a second image in real time. For example, the user may operate the input interface 128 to create a parameter set, and the actuator mechanism 124 may initiate movement of the slide 116 relative to the optical system 120 substantially immediately after the input interface 128 is operated, and the optical system 120 may capture a second image substantially immediately after the actuator mechanism 124 completes its movement. In some embodiments, artifacts may also be removed in real time. In some embodiments, the image may also be annotated in real time. In some embodiments, the focus pattern may also be determined in real time, and the image taken after the first image may be taken using a focal distance according to the focus pattern.
[0097] Referring still to FIG. 1, in some embodiments, the apparatus 100 can display a second image to the user. In some embodiments, the second image can be displayed to the user using the output device 132. The second image can be displayed as described above with respect to the output device and the display of the first image. In some embodiments, displaying the second image to the user can include replacing an area of the first image with the second image to generate a hybrid image and displaying the hybrid image to the user. As used herein, a "hybrid image" is an image constructed by combining a first image and a second image. In some embodiments, by generating a hybrid image in this way, the second image can be retained. For example, if the second image has a higher resolution per unit area and covers a smaller area than the first image, the second image can replace segments of the first area corresponding to the area covered by the second image that have a lower resolution per unit area. In some embodiments, at the boundary between the first image and the second image in the hybrid image, image adjustment can be performed to cancel out the visual difference between the first image and the second image. In a non-limiting example, the color can be adjusted so that the background color of the image is constant along the boundary of the image. As another non-limiting example, the brightness of the image can be adjusted so that there is no significant difference between the brightnesses of the images. In some embodiments, as described above, artifacts can be removed from the second image and / or the hybrid image. In some embodiments, the second image can be displayed to the user in real time. For example, the adjustment (such as annotation and / or artifact removal) can start substantially immediately after the second image is captured, and the adjusted version of the second image can be displayed to the user substantially immediately after the adjustment is performed. In some embodiments, an unadjusted version of the second image can be displayed to the user while the adjustment is being performed. In some embodiments, if there are multiple images covering a particular area, when the user zooms out using the user interface 136, a low-resolution image of the area may be displayed.
[0098] Referring still to FIG. 1, in some embodiments, the apparatus 100 may transmit a data structure including a first image, a second image, a hybrid image, and / or a plurality of images to an external device. Such an external device can include, in non-limiting examples, a phone, a tablet, or a computer. In some embodiments, such transmission may configure the external device to display the image.
[0099] Referring still to FIG. 1, in some embodiments, the apparatus 100 can annotate an image. In some embodiments, the apparatus 100 can annotate the first image. In some embodiments, the apparatus 100 can annotate the second image. In some embodiments, the apparatus 100 can annotate the hybrid image. For example, when creating the hybrid image, the apparatus 100 can recognize cells depicted in the hybrid image as a particular type of cell and attach an annotation indicating the type of cell to the hybrid image. In non-limiting examples, the apparatus 100 can associate text with a particular location within the image, and the text describes a feature present at that location within the image. In some embodiments, the apparatus 100 can annotate an image selected from a list consisting of the first image, the second image, and the hybrid image.
[0100] Referring still to FIG. 1, in some embodiments, the annotation can be performed as a function of user input of an annotation command to the input interface 128. As used herein, an "annotation command" is data generated based on user input indicating whether to create an annotation or user input describing the annotation to be created. For example, the user can select an option to control whether the apparatus 100 annotates an image. In another example, the user can manually annotate an image. In some embodiments, the annotation can be performed automatically. In some embodiments, the annotation can be performed using an annotation machine learning model. In some embodiments, the annotation machine learning model can include an optical character recognition model, as described above. In some embodiments, the annotation machine learning model can be trained using a data set that includes image data associated with text depicted in an image. In some embodiments, the annotation machine learning model can receive input image data and output annotated image data and / or annotations to be applied to the image data. In some embodiments, the annotation machine learning model can be used to convert text written on the slide 116 into an annotation on the image. In some embodiments, the annotation machine learning model can include a machine vision model, as described above. In some embodiments, an annotation machine learning model that includes a machine vision model can be trained on a data set that includes image data associated with annotations indicating features of the image data. In some embodiments, the annotation machine learning model can receive input image data and output annotated image data and / or annotations to be applied to the image data. Non-limiting examples of features that the annotation machine learning model can be trained to recognize include cell types, cell features, and objects within the slide 116 such as bubbles. In some embodiments, the image can be annotated in real time. For example, the annotation can begin immediately after the image is captured and / or immediately after a command to annotate the image is received, and the annotated image can be displayed to the user immediately after the annotation is complete.
[0101] Referring still to FIG. 1, in some embodiments, the apparatus 100 can determine a visual element data structure. In some embodiments, the apparatus 100 can display visual elements to the user as a function of the visual element data structure. As used herein, a "visual element data structure" is a data structure that describes visual elements. By way of non-limiting example, visual elements can include a first image, a second image, a hybrid image, and GUI elements.
[0102] Referring still to FIG. 1, in some embodiments, the visual element data structure can include visual elements. As used herein, a "visual element" is data that is visually displayed to the user. In some embodiments, the visual element data structure can include rules for displaying visual elements. In some embodiments, the visual element data structure can be determined as a function of a first image, a second image, and / or a hybrid image. In some embodiments, the visual element data structure can be determined as a function of items from a list consisting of a first image, a second image, a hybrid image, GUI elements, and annotations. By way of non-limiting example, the visual element data structure can be generated such that visual elements that describe features of the first image, such as annotations, are displayed to the user.
[0103] Referring still to FIG. 1, in some embodiments, visual elements can include one or more elements such as text, images, shapes, charts, particle effects, interactive features, and the like. By way of non-limiting example, visual elements can include a touch screen button for setting a magnification level.
[0104] Referring still to FIG. 1, the visual element data structure can include rules for managing whether or not, or when, visual elements are to be displayed. In a non-limiting example, the visual element data structure can include a rule to display visual elements including annotations that describe a first image, a second image, and / or a hybrid image when a user selects a particular region of the first image, the second image, and / or the hybrid image using a GUI.
[0105] Referring still to FIG. 1, the visual element data structure can include rules for presenting a plurality of visual elements, or a plurality of visual elements at a time. In one embodiment, about 1, 2, 3, 4, 5, 10, 20, or 50 visual elements are displayed simultaneously. For example, multiple annotations can be displayed simultaneously.
[0106] Referring still to FIG. 1, in some embodiments, the apparatus 100 can transmit visual elements to a display such as the output interface 132. The display can communicate visual elements to the user. The display can include, for example, a smartphone screen, a computer screen, a tablet screen, etc. The display can be configured to provide a visual interface. The visual interface can include one or more virtual interactive elements such as, but not limited to, buttons, menus, etc. The display can include one or more physical interactive elements such as buttons, a computer mouse, or a touch screen that enable a user to input data to the display. The interactive elements can be configured to enable interaction between the user and the computing device. In some embodiments, the visual element data structure is determined as a function of data input by the user to the display.
[0107] Referring still to FIG. 1, the variables and / or data described herein can be represented as a data structure. In some embodiments, the data structure can include one or more functions and / or variables, such as a class in object-oriented programming. In some embodiments, the data structure can include data in the form of boolean values, integers, floating-point numbers, strings, dates, etc. By way of non-limiting example, an annotation data structure can include a string value representing the text of the annotation. In some embodiments, the data of the data structure can be organized in a linked list, tree, array, matrix, tensor, etc. By way of non-limiting example, an annotation data structure can be organized in an array. In some embodiments, the data structure may include one or more elements of metadata or may be associated with one or more elements of metadata. The data structure can include one or more self-referential data elements that can be used by the processor 104 when interpreting the data structure. By way of non-limiting example, the data structure can include a " <date>」 and 「< / date> " tag indicating that the content between the tags is a date.
[0108] Referring still to FIG. 1, the data structure can be stored, for example, in the memory 108 or a database. The database can be implemented as, but is not limited to, a relational database, a key-value type database such as a NOSQL database, or any other format or structure used as a database that those skilled in the art would recognize as appropriate upon considering the entire disclosure. The database can alternatively or additionally be implemented using a distributed data storage protocol and / or data structure such as a distributed hash table. The database can include, as described above, a plurality of data entries and / or records. The data entries in the database can be flagged or linked to one or more additional information elements and can be reflected in linked tables such as tables associated by one or more indexes within the data entry cells and / or within a relational database. Those skilled in the art, upon considering the entire disclosure, will recognize the various ways in which the data entries in the database can store, retrieve, organize, and / or reflect the data and / or records used herein, as well as the categories and / or populations of data, without conflicting with the present disclosure.
[0109] Referring still to FIG. 1, in some embodiments, the data structure can be read and / or operated on by the processor 104. In a non-limiting example, an image data structure can be read and displayed to a user. In another non-limiting example, as described above, the image data structure can be modified to remove artifacts.
[0110] Referring still to FIG. 1, in some embodiments, the data structure can be calibrated. In some embodiments, the data structure can be trained using a machine learning algorithm. In a non-limiting example, the data structure can include an array of data representing the bias of the connections of a neural network. In this example, the neural network is trained on a set of training data, and the backpropagation algorithm can be used to modify the array of data. Machine learning models and neural networks are further described herein.
[0111] One or more features of apparatus 100 may match one or more of the features disclosed below. (A) U.S. Patent Application No. 18 / 217,378, filed July 25, 2023, titled "APPARATUS AND A METHOD FOR DETECTING ASSOCIATIONS AMONG DATASETS OF DIFFERENT TYPES" (incorporated herein by reference in its entirety) (B) U.S. Patent Application No. 18 / 226,017, filed July 25, 2023, titled "APPARATUS AND A METHOD FOR GENERATING A CONFIDENCE SCORE ASSOCIATED WITH A SCANNED LABEL" (incorporated herein by reference in its entirety) (C) U.S. Patent Application No. 18 / 226,058, filed July 25, 2023, titled "IMAGING DEVICE AND A METHOD FOR IMAGE GENERATION OF A SPECIMEN" (incorporated herein by reference in its entirety) (D) U.S. Patent Application No. 18 / 226,100, filed July 25, 2023, titled "APPARATUS AND METHODS FOR REAL-TIME IMAGE GENERATION" (incorporated herein by reference in its entirety).
[0112] Referring now to FIG. 2, an exemplary embodiment of a machine learning module 200 that can execute one or more machine learning processes as described in the present disclosure is illustrated. The machine learning module can use machine learning processes to perform decisions, classifications, and / or analysis steps, methods, processes, etc. as described in the present disclosure. As used in the present disclosure, a "machine learning process" is a process that automatically uses training data 204 to generate hardware or software logic, data structures, and / or algorithms instantiated in functions executed by a computing device / module to generate an output 208 when data is provided as an input 212, which is contrasted with non-machine learning software programs where the commands to be executed are pre-determined by a user and described in a programming language.
[0113] Referring still to FIG. 2, "training data" as used herein is data that includes correlations that can be used by a machine learning process to model relationships between two or more categories of data elements. For example, without limitation, training data 204 may include a plurality of data entries, also known as "training examples", each entry representing a set of data elements that were recorded, received, and / or generated together, and the data elements can be correlated by, for example, their co-presence in a given data entry, their proximity in a given data entry, etc. The plurality of data entries of training data 204 can exhibit one or more trends in the correlations between categories of data elements. For example, without limitation, a high value of a first data element belonging to a first category of data elements tends to correlate with a high value of a second data element belonging to a second category of data elements, indicating the possibility of a proportional or other mathematical relationship linking the values belonging to the two categories. The plurality of categories of data elements can be associated in training data 204 according to various correlations. A correlation can be one that indicates a causal relationship and / or a predictive link between categories of data elements and can be modeled as a relationship, such as a mathematical relationship, by a machine learning process, as will be described in more detail below. Training data 204 can be formatted and / or organized by categories of data elements, for example, by associating data elements with one or more descriptors corresponding to the categories of data elements. As a non-limiting example, training data 204 can include data that has been input into a standardized form by a person or a process such that the input of a given data element in a given field of the form is mapped to one or more descriptors of a category. Elements of training data 204 may be linked to descriptors of a category by tags, tokens, or other data elements.For example, but not limited to, the training data 204 may be provided in a format that links the positions of data, such as a fixed-length format, a comma-separated values (CSV) format, and / or a self-describing format such as Extensible Markup Language (XML), JavaScript Object Notation (JSON), enabling a process or device to detect the categories of the data.
[0114] Alternatively or additionally, still referring to FIG. 2, the training data 204 may include one or more uncategorized elements. That is, the training data 204 may not be formatted or may not include descriptors for some elements of the data. Machine learning algorithms and / or other processes may sort the training data 204 according to one or more categorizations, for example, using natural language processing algorithms, tokenization, detection of correlation values in raw data, etc. Categories may be generated using correlation and / or other processing algorithms. As a non-limiting example, in a corpus of text, phrases that constitute "n" compound words such as a noun modified by another noun may be identified according to the statistically significant prevalence of n-grams that contain such words in a particular order. Such n-grams may be categorized as elements of a language, such as "words", that are tracked like single words, and new categories may be generated as a result of statistical analysis. Similarly, in a data entry that includes text data, a person's name may be identified by referring to a list, dictionary, or other glossary, enabling ad hoc categorization by a machine learning algorithm and / or automatic association of the data in the data entry with a descriptor or a given format. The ability to automatically categorize data entries allows the same training data 204 to be applied to two or more different machine learning algorithms, as will be described in more detail below. The training data 204 used by the machine learning module 200 may correlate any input data as described in the present disclosure with any output data as described in the present disclosure. As a non-limiting illustrative example, the input may include an image of an area of interest, and the output may include a determination of whether a sample is present or not.
[0115] Referring further to FIG. 2, the training data may be filtered, sorted, and / or selected using one or more supervised and / or unsupervised machine learning processes and / or models, as will be described in further detail below, such models may include, but are not limited to, a training data classifier 216. The training data classifier 216 may include a "classifier", and the "classifier" used in the present disclosure is a machine learning model defined below, for example, a mathematical model, neural network, or data structure that represents and / or uses a program generated by a machine learning algorithm known as a "classification algorithm", as will be described in further detail below, that sorts inputs into categories or bins of data and outputs the categories or bins of data and / or associated labels. The classifier may be configured to output at least one data that labels or otherwise identifies a set of data that has been clustered and found to be proximal under a distance metric as described below. The distance metric can include any norm, such as, but not limited to, the Pythagorean norm. The machine learning module 200 may generate a classifier using a classification algorithm defined as a process by which a computing device and / or any module and / or component operating on the computing device derives a classifier from the training data 204. Classification may be performed using, but not limited to, linear classifiers such as logistic regression and / or naive Bayes classifiers, nearest neighbor classifiers such as k-nearest neighbor classifiers, support vector machines, least squares support vector machines, Fisher's linear discriminant, quadratic classifiers, decision trees, boosting trees, random forest classifiers, learning vector quantization, and / or neural network-based classifiers. By way of non-limiting example, the training data classifier 216 may classify elements of the training data as to whether a sample is present or not.
[0116] Referring further to FIG. 2, the training examples used as training data can be selected from a population of potential examples according to cohorts related to the analysis problem to be solved, classification tasks, and the like. Alternatively or additionally, the training data may be selected to span situations or sets of inputs that the machine learning model and / or process are likely to encounter when deployed. For example, without limitation, for each category of input data to a machine learning process or model that may exist within a range of values in a population of phenomena such as images, user data, processing data, physical data, etc., the computing device, processor, and / or machine learning model can select training examples representing each possible value on such a range, and / or representative samples of values on such a range. The selection of representative samples can include, for example, selecting training examples at a rate that matches the distribution of such values statistically determined and / or predicted according to relative frequencies such that values that occur more frequently in the population of analyzed data are represented by more training examples than values that occur less frequently. Alternatively or additionally, the set of training examples can be compared to a collection of representative values in a database and / or presented to a user such that the process can detect one or more values not included in the set of training examples, either automatically or via user input. The computing device, processor, and / or module may automatically generate missing training examples. This may be done by receiving and / or obtaining the missing input values and / or output values and associating the missing input values and / or output values with corresponding output values and / or input values coexisting in the data record and provided by a user and / or other devices, etc.
[0117] Referring still to FIG. 2, a computer, processor, and / or module may be configured to sanitize training data. As used herein, "sanitizing" training data is a process by which training examples that impede convergence to a useful result of a machine learning model and / or process are removed. For example, without limitation, training examples can include input and / or output values that are outliers from typical values that a machine learning algorithm using the training examples would likely encounter, such that the machine learning algorithm adapts to amounts that are unlikely to be inputs and / or outputs. For example, values that are greater than a threshold number of standard deviations away from an average, mean, or expected value may be removed. Alternatively or additionally, one or more training examples may be identified as having low-quality data. Here, "low-quality" is defined as having a signal-to-noise ratio below a threshold.
[0118] As a non-limiting example, further referring to FIG. 2, an image classifier or other machine learning model, and / or an image used for training a process that takes an image as input or generates an image as output may be rejected if the image quality is below a threshold. For example, without limitation, a computing device, a processor, and / or a module may perform blur detection and be able to eliminate one or more blurs. Blur detection may be performed, as a non-limiting example, by taking an approximation such as a Fourier transform of an image, or a fast Fourier transform (FFT), and analyzing the distribution of low and high frequencies in the frequency domain depiction of the resulting image. The number of high frequency values below a threshold level may indicate a possible blur. In a further non-limiting example, blur detection may be performed by convolving an image, an image channel, etc. with a Laplacian kernel, which may generate a numerical score that reflects the number of sharp changes in intensity shown in the image, where a high score indicates sharpness and a low score indicates blur. Blur detection can be performed using a gradient-based operator that measures an operator based on the gradient or first derivative of an image, based on the hypothesis that sharp changes indicate sharp edges in the image and thus a low degree of blur. Blur detection may be performed using a wavelet-based operator that utilizes the ability of the coefficients of a discrete wavelet transform to describe the frequency and spatial content of an image. Blur detection may be performed using a statistic-based operator that utilizes several image statistics as texture descriptors to calculate a focus level. Blur detection may be performed by using discrete cosine transform (DCT) coefficients to calculate the focus level of an image from its frequency content.
[0119] Referring still to FIG. 2, a computing device, processor, and / or module may be configured to assume one or more training examples. For example, but not limited to, if a machine learning model and / or process has one or more inputs and / or outputs that require, transmit, or receive a certain number of bits, samples, or other units of data, elements of one or more training examples used as or compared to the inputs and / or outputs may be modified to have such a unit number of data. For example, a computing device, processor, and / or module can convert a smaller number of units, such as a low-pixel image, to a desired number of units, such as by upsampling or interpolation. As a non-limiting example, a low-pixel image may have 100 pixels, but the desired number of pixels may be 128 pixels. The processor can interpolate the low-pixel image to convert 100 pixels to 128 pixels. Also, those skilled in the art should note that upon reading this disclosure, they will know various methods of interpolating a smaller number of data units, such as samples, pixels, bits, etc., to a desired number of such units. In some examples, a set of interpolation rules may be trained by a highly detailed input and / or output, a corresponding set of inputs and / or outputs downsampled to a smaller number of units, and a neural network or other machine learning model trained to predict interpolated pixel values using the training data. As a non-limiting example, sample inputs and / or outputs, such as a sample image having sample-expanded data units (e.g., pixels added between original pixels), can be input into a neural network or machine learning model and can output a pseudo-replica sample image with dummy values assigned to pixels between the original pixels based on the set of interpolation rules.As a non-limiting example, in the context of an image classifier, a machine learning model may have a set of high-definition images, a set of interpolation rules trained by a set of images downsampled to fewer pixels, and a neural network or other machine learning model trained using those examples to predict interpolated pixel values in the context of face images. As a result, an input having sample-expanded data units (with dummy values added between the original data units) is run through the trained neural network and / or model, and values can be filled in to replace the dummy values. Alternatively or additionally, a processor, computing device, and / or module may utilize a sample expander method, a low-pass filter, or both. As used in this disclosure, a "low-pass filter" is a filter that passes signals of frequencies lower than a selected cutoff frequency and attenuates signals of frequencies higher than the cutoff frequency. The exact frequency response of the filter depends on the filter design. A computing device, processor, and / or module may use averaging, such as luma averaging or chroma averaging in an image, to fill in data units between the original data units.
[0120] In some embodiments, continuing to refer to FIG. 2, a computing device, processor, and / or module can downsample the elements of a training example to a desired fewer number of data elements. As a non-limiting example, an image with a high number of pixels may have 256 pixels, but the desired number of pixels may be 128 pixels. The processor can downsample the high-pixel-count image to convert 256 pixels to 128 pixels. In some embodiments, the processor may be configured to perform downsampling on the data. Downsampling, also known as decimation, may include removing every Nth entry in a sequence of samples, every entry except every Nth entry, etc., which is a process known as "compression" and can be performed, for example, by an N-sample compressor implemented using hardware or software. Anti-aliasing and / or anti-imaging filters, and / or low-pass filters can be used to clean up the side effects of the compression.
[0121] Referring still to FIG. 2, the machine learning module 200 may be configured to execute a delayed learning process 220 and / or protocol. This may alternatively be referred to as a “lazy loading” or “call-when-needed” process and / or protocol, and machine learning is performed by deriving an algorithm that combines the input and the training set upon receipt of the input to be converted to an output, for on-demand generation of the output. For example, an initial set of simulations may be executed to cover initial heuristics and / or “first guesses” in the output and / or relationships. As a non-limiting example, the initial heuristics can include ranking the relevance between the input and elements of the training data 204. The heuristics may include selecting some of the highest-ranked relevant and / or elements of the training data 204. Delayed learning can implement any suitable delayed learning algorithm, including but not limited to, the k-nearest neighbor algorithm, the lazy naive bayes algorithm, etc., and those skilled in the art, upon considering the entire disclosure, will recognize various delayed learning algorithms that can be applied to generate an output as described in the present disclosure, including but not limited to, delayed learning applications of machine learning algorithms as will be described in more detail below.
[0122] Alternatively or additionally, still referring to FIG. 2, a machine learning process as described in the present disclosure may be used to generate the machine learning model 224. As used in the present disclosure, a "machine learning model" is a mathematical and / or algorithmic representation of the relationship between inputs and outputs that is generated using any machine learning process, including but not limited to any of the processes described above, stored in memory, and / or instantiated as a data structure. Once created, the inputs are presented to the machine learning model 224, which generates an output based on the derived relationship. For example, without limitation, a linear regression model generated using a linear regression algorithm can calculate a linear combination of input data using the coefficients derived during the machine learning process and calculate the output data. As a further non-limiting example, the machine learning model 224 may be generated by creating an artificial neural network, such as a convolutional neural network, that includes an input layer of nodes, one or more intermediate layers, and an output layer of nodes. The connections between the nodes may be created through a process of "training" the network, in which elements from a set of training data 204 are applied to the input nodes and an appropriate training algorithm (such as Levenberg-Marquardt, conjugate gradient, simulated annealing, or other algorithms) is then used to adjust the connections and weights between the nodes in adjacent layers of the neural network to generate the desired values at the output nodes. This process is sometimes referred to as deep learning.
[0123] Referring still to FIG. 2, the machine learning algorithm can include at least one supervised machine learning process 228. The at least one supervised machine learning process 228, as defined herein, receives a training set that associates a number of inputs with a number of outputs and attempts to generate one or more data structures that represent and / or instantiate one or more mathematical relationships that associate the inputs with the outputs, each of the one or more mathematical relationships being optimal according to some criterion specified to the algorithm using some scoring function. For example, the supervised learning algorithm can include an image of the region of interest as described above as an input, a determination of whether a sample exists as an output, and a scoring function that represents the desired form of the relationship to be detected between the input and the output. The scoring function can, for example, attempt to maximize the probability that a given input and / or combination of elements of the input is associated with a given output and minimize the probability that a given input is not associated with a given output. The scoring function may be expressed as a risk function that represents the "expected loss" of the algorithm that associates the input with the output, where the loss is calculated as an error function that represents the degree to which the prediction generated by the relationship is inaccurate when compared to a given input-output pair provided in the training data 204. Those skilled in the art will recognize the various possible variations of the at least one supervised machine learning process 228 that can be used to determine the relationship between the input and the output when considering the entire disclosure. The supervised machine learning process can include the classification algorithm defined above.
[0124] Referring further to FIG. 2, the training of the supervised machine learning process can include, but is not limited to, iteratively updating coefficients, biases, and weights based on an error function, an expected loss, and / or a risk function. For example, the output generated by a supervised machine learning model using an input example of a training example may be compared to an output example from the training example, and the error function may be generated based on the comparison, which may include any error function suitable for use with any machine learning algorithm described in the present disclosure, such as the square of the difference between one or more sets of the compared values. Such an error function may be sequentially used to update one or more weights, biases, coefficients, or other parameters of the machine learning model through any suitable process, including but not limited to a gradient descent process, a least squares process, and / or other processes described in the present disclosure. This may be done iteratively and / or recursively to gradually adjust the weights, biases, coefficients, or other parameters. The update may be performed using one or more backpropagation algorithms in a neural network. The iterative and / or recursive update of the weights, biases, coefficients, or other parameters as described above may be performed until the currently available training data is depleted and / or until a convergence determination is passed, where "convergence determination" is a determination against a condition selected to indicate that the model and / or its weights, biases, coefficients, or other parameters have reached a certain level of accuracy. The convergence determination can, for example, compare the difference between two or more consecutive errors or error function values, and a difference below a threshold can be considered to indicate convergence. Alternatively or additionally, one or more errors and / or error function values evaluated in a training iteration may be compared to a threshold.
[0125] Referring still to FIG. 2, a computing device, processor, and / or module may be configured to execute the methods, method steps, sequences of method steps, and / or algorithms described with reference to this figure in any order and to any extent of repetition. For example, a computing device, processor, and / or module may be configured to repeatedly execute a single step, sequence, and / or algorithm until a desired or commanded result is achieved. The repetition of a step or sequence of steps may be executed iteratively and / or recursively using the output of a previous iteration as input to a subsequent iteration, the aggregation of inputs and / or outputs of iterations to produce an aggregated result, the decrement or decrementation of one or more variables such as global variables, and / or the division of a large processing task into a set of smaller processing tasks that are repeatedly addressed. A computing device, processor, and / or module may execute any step, sequence of steps, or algorithm in parallel, such as by using two or more parallel threads, processor cores, etc., to execute the steps simultaneously and / or more than twice substantially simultaneously, and the division of tasks among the parallel threads and / or processes may be executed according to any protocol suitable for the division of tasks among iterations. Those skilled in the art will recognize, upon consideration of the present disclosure in its entirety, the various ways in which iterations, recursion, and / or parallel processing may be used to subdivide, share, or otherwise process steps, sequences of steps, processing tasks, and / or data.
[0126] Referring further to FIG. 2, the machine learning process can include at least one unsupervised machine learning process 232. An unsupervised machine learning process, as used herein, is a process that derives inferences in a dataset without regard to labels, and as a result, an unsupervised machine learning process can freely discover any structure, relationship, and / or correlation provided in the data. The unsupervised machine learning process 232 may not require a response variable, and the unsupervised machine learning process 232 can be used to find interesting patterns and / or inferences between variables, to determine the degree of correlation between two or more variables, etc.
[0127] Referring still to FIG. 2, the machine learning module 200 can be designed and configured to create a machine learning model 224 using techniques for the development of a linear regression model. The linear regression model may include ordinary least squares regression, which aims to minimize the sum of the squares of the differences between the predicted and actual results according to an appropriate norm (e.g., vector space distance norm) for measuring such differences. The coefficients of the resulting linear equation can be modified to improve the minimization. The linear regression model may include ridge regression, and the function to be minimized may include a term that multiplies the square of each coefficient by a scalar amount to penalize large coefficients in addition to the least squares function. The linear regression model may include a least absolute shrinkage and selection operator (lasso) model, in which ridge regression is combined with multiplying the coefficient 1 divided by twice the number of samples by the least squares term. The linear regression model may include a multi-task lasso model, and the norm applied to the least squares term of the lasso model is the Frobenius norm corresponding to the square root of the sum of the squares of all terms. The linear regression model may include an elastic net model, a multi-task elastic net model, a least angle regression model, a LARS lasso model, an orthogonal matching pursuit model, a Bayesian regression model, a logistic regression model, a stochastic gradient descent model, a perceptron model, a passive aggressive algorithm, a robustness regression model, a Huber regression model, or other suitable models that one of ordinary skill in the art may envision upon consideration of the entire disclosure. In an embodiment, the linear regression model may be generalized to a polynomial regression model, whereby a polynomial (e.g., quadratic, cubic, or higher order) that provides the best fit of the predicted output / actual output is sought. As will be apparent to one of ordinary skill in the art upon consideration of the entire disclosure, methods similar to those described above can be applied to minimize the error function.
[0128] Referring still to FIG. 2, the machine learning algorithm may include, but is not limited to, linear discriminant analysis. The machine learning algorithm may include quadratic discriminant analysis. The machine learning algorithm may include kernel ridge regression. The machine learning algorithm may include a support vector machine including, but not limited to, regression processing based on support vector classification. The machine learning algorithm may include a stochastic gradient descent algorithm including classification and regression algorithms based on stochastic gradient descent. The machine learning algorithm may include a nearest neighbor algorithm. The machine learning algorithm may include various forms of latent space regularization such as variational regularization. The machine learning algorithm may include a Gaussian process such as Gaussian process regression. The machine learning algorithm may include a cross decomposition algorithm including partial least squares method and / or canonical correlation analysis. The machine learning algorithm may include a naive Bayes method. The machine learning algorithm may include a decision tree-based algorithm such as decision tree classification or regression algorithm. The machine learning algorithm may include an ensemble method such as a bagging meta-estimator, a random tree forest, AdaBoost, gradient tree boosting, and / or a voting classifier method. The machine learning algorithm may include a neural network algorithm including convolutional neural network processing.
[0129] Referring still to FIG. 2, the machine learning model and / or process may be deployed or instantiated by incorporation into a program, apparatus, system, and / or module. For example, but not limited to, the machine learning model, neural network, and / or some or all of its parameters may be stored and / or deployed in any memory or circuit. Parameters such as coefficients, weights, and / or biases may be stored as circuit-based constants such as wires set to logic “1” and “0” voltage levels and / or an array of binary inputs and / or outputs in a logic circuit to represent numbers according to any suitable encoding system including two's complement, or may be stored in any volatile memory and / or non-volatile memory. Similarly, mathematical operations and the input and / or output of data to and from models, neural network layers, etc. may be instantiated in hardware circuits and / or in the form of instructions in firmware, machine code such as binary arithmetic code instructions, assembly language, or any high-level programming language. Any technique for the hardware and / or software instantiation of memory, instructions, data structures, and / or algorithms may be used to instantiate the machine learning process and / or model, which includes, but is not limited to, the manufacture and / or configuration of non-reconfigurable hardware elements, circuits, and / or modules such as ASICs but not limited thereto, the manufacture and / or configuration of reconfigurable hardware elements, circuits, and / or modules such as FPGAs but not limited thereto, the manufacture and / or configuration of non-rewritable memory elements, circuits, and / or modules such as non-rewritable ROMs but not limited thereto, the manufacture and / or configuration of reconfigurable and / or rewritable memory elements, circuits, and / or modules such as rewritable ROMs or other memory technologies described in the present disclosure but not limited thereto, and / or any combination of the manufacture and / or configuration of any computing device and / or its components as described in the present disclosure.Such deployed and / or instantiated machine learning models and / or algorithms can receive inputs from any other processes, modules, and / or components described in this disclosure and generate outputs for any other processes, modules, and / or components described in this disclosure.
[0130] Still referring to FIG. 2, for purposes of modifying, improving, and / or enhancing a machine learning model and / or algorithm, any process of training, retraining, deploying, and / or instantiating the machine learning model and / or algorithm may be performed and / or repeated after an initial deployment and / or instantiation. Such retraining, deployment, and / or instantiation may be performed as a periodic or regular process, for example, after a measure of quantity such as the number of bytes of processed data or other metric, the number of times of use or execution of the processes described in this disclosure, and / or according to a software, firmware, or other update schedule, as retraining, deployment, and / or instantiation at regular elapsed time intervals. Alternatively or additionally, retraining, deployment, and / or instantiation may be event-based and, without limitation, may be triggered by user input indicating sub-optimal or otherwise problematic performance and / or by an automated field test and / or auditing process, which may compare the machine learning model and / or algorithm and / or the output of an error and / or its error function to any threshold, convergence determination, etc., and / or may compare the output of the processes described herein to similar thresholds, convergence determinations, etc. Event-based retraining, deployment, and / or instantiation may alternatively or additionally be triggered by the receipt and / or generation of one or more new training examples, and the number of new training examples may be compared to a pre-configured threshold, and if the pre-configured threshold is exceeded, retraining, deployment, and / or instantiation may be triggered.
[0131] Referring still to FIG. 2, re-training and / or additional training can be performed using any version of a machine learning model and / or algorithm that is currently or previously deployed as a starting point, using any process for the training described above. Training data for re-training can be collected, pre-processed, sorted, classified, sanitized, or otherwise processed according to any process described in this disclosure. The training data can include training examples that are used, received, and / or generated from any version of any system, module, machine learning model or algorithm, device, and / or method described in this disclosure, including inputs and associated outputs, and such examples can be modified and / or labeled according to user feedback or other processes to indicate a desired result, and / or have an actual or measured result from a process modeled and / or predicted by a system, module, machine learning model or algorithm, device, and / or method as the "desired" result to be compared with the output for the training process as described above.
[0132] Redployment may be performed using any reconfiguration and / or rewriting of reconfigurable and / or rewritable circuits and / or memory elements, or alternatively, redployment may be performed by manufacturing new hardware and / or software components, circuits, instructions, etc., which may be added to existing hardware and / or software components, circuits, instructions, etc., and / or replace existing hardware and / or software components, circuits, instructions, etc.
[0133] Referring further to FIG. 2, one or more of the above-described processes or algorithms may be executed by at least one dedicated hardware unit 236. For the purposes of this figure, a "dedicated hardware unit" is a hardware component, circuit, etc. that is specifically designated or selected to perform one or more specific tasks and / or processes described with reference to this figure, such as, but not limited to, preconditioning and / or sanitizing training data, and / or training machine learning algorithms and / or models, other than the main control circuit and / or processor that executes the method steps described in this disclosure. The dedicated hardware unit 236 may include a hardware unit capable of efficiently performing iterative or large-scale calculations, such as matrix-based calculations for updating or adjusting the parameters, weights, coefficients, and / or biases of a machine learning model and / or neural network, using pipeline processing, parallel processing, etc. Such a hardware unit may include, for example, dedicated circuits for matrix operations and / or signal processing operations including a plurality of arithmetic circuit units and / or logic circuit units, such as multipliers and / or adders that can operate, for example, simultaneously and / or in parallel, and may be optimized for such processing. Such a dedicated hardware unit 236 may include, but is not limited to, a graphics processing unit (GPU), a dedicated signal processing module, an FPGA, or other reconfigurable hardware configured to instantiate parallel processing units for one or more specific tasks. A computing device, processor, apparatus, or module may be configured to direct one or more dedicated hardware units 236 to perform one or more operations described herein, such as evaluating model and / or algorithm outputs, one-time or iterative updating of parameters, coefficients, weights, and / or biases, and / or any other operations such as vector and / or matrix operations described in this disclosure.
[0134] Next, referring to FIG. 3, an exemplary embodiment of a neural network 300 is illustrated. A neural network 300, also known as an artificial neural network, is a network of "nodes" or data structures that have one or more inputs, one or more outputs, and a function that determines an output based on the inputs. Such nodes can be organized into a network, such as, but not limited to, a convolutional neural network that includes an input layer of nodes 304, one or more intermediate layers 308, and an output layer of nodes 312. Connections between nodes may be created via a process of "training" the network, in which elements from a set of training data are applied to input nodes and an appropriate training algorithm (such as Levenberg-Marquardt, conjugate gradient, simulated annealing, or other algorithms) is then used to adjust the connections and weights between nodes in adjacent layers of the neural network to produce a desired value at the output nodes. This process is sometimes referred to as deep learning. Connections are made only from input nodes towards output nodes in a "feedforward" network, and in a "recurrent network" the output of one layer can be fed back as input to the same or a different layer. As a further non-limiting example, a neural network can include a convolutional neural network that includes an input layer of nodes, one or more intermediate layers, and an output layer of nodes. As used in this disclosure, a "convolutional neural network" is a neural network in which at least one hidden layer is a convolutional layer that convolves an input to the layer with a subset of the input known as a "kernel" along with one or more additional layers such as a pooling layer, a fully connected layer, etc.
[0135] Next, referring to FIG. 4, an exemplary embodiment of a node 400 of a neural network is illustrated. The node can receive a plurality of inputs x from, among other things, inputs to the neural network that includes the node and / or from other nodes, and can receive numerical values from these inputs. ican include. When one or more inputs are provided, the node can execute one or more activation functions to generate its output. The activation functions include, but are not limited to, a binary step function that compares the input with a threshold and outputs a logical 1 or logical 0 output or the equivalent, a linear activation function where the output is proportional to the input, and / or a non-linear activation function where the output is not proportional to the input. The non-linear activation functions include, but are not limited to, a sigmoid function of the form [Number] of the form, [Number] a tanh (hyperbolic tangent) function of the form, f(x) = tanh 2 (x), a tanh derivative function such as f(x) = max(0, x), a rectified linear unit function such as f(x) = max(ax, x) for some a, a "leaky" and / or "parametric" rectified linear unit function for some value of α, [Number] an exponential linear unit function such as (in some embodiments, this function may be replaced and / or weighted by its own derivative), where the input to the instant layer is x i is, [Number] a softmax function such as, a swish function such as f(x) = x * sigmoid(x), for some values of a, b, r [Number] a Gaussian error linear unit function such as, and / or [Number] can include scaled exponential linear unit functions such as etc. Basically, the input x that can be used as an activation function i There is no limitation on the nature of the function. As a non-limiting and exemplary example, each node multiplies the input x i by the weight w i to perform a weighted sum of the inputs. Additionally or alternatively, a bias b may be added to the weighted sum of the inputs such that an offset independent of the input to the layer is added to each unit of the neural network layer. The weighted sum may then be input to a function φ to generate one or more outputs y. The weight w i applied to the input x i can indicate whether its input is "excitatory" such that it has a strong influence on one or more outputs y, for example by having a large corresponding weight value, and / or "inhibitory" such that it has a weak influence on one further output y, for example by having a small corresponding weight value. The value of the weight w i may be determined by training the neural network using training data, and this training may be performed using any suitable process as described above.
[0136] Still referring to FIG. 4, the "convolutional neural network" used in the present disclosure is a neural network in which at least one hidden layer is a convolutional layer that convolves the input to the layer with a subset of the input known as a "kernel" together with one or more additional layers such as a pooling layer, a fully connected layer, etc. A CNN can include, but is not limited to, an extension of a deep neural network (DNN), which is defined as a neural network having two or more hidden layers.
[0137] Referring still to FIG. 4, in some embodiments, the convolutional neural network can learn from an image. In a non-limiting example, the convolutional neural network may perform tasks such as image classification, detection of objects depicted in the image, image segmentation, and / or image processing. In some embodiments, the convolutional neural network may operate such that each node in the input layer is connected only to a region of nodes in the hidden layer. In some embodiments, the regions may collectively create a map of features from the input layer to the hidden layer. In some embodiments, the convolutional neural network may include a layer in which the weights and biases of all nodes are the same. In some embodiments, this allows the convolutional neural network to detect features such as edges at different locations within the image.
[0138] Referring now to FIG. 5, an exemplary embodiment of a slide imaging method 500 is illustrated. This can help reduce errors due to out-of-focus. At step 505, the entire slide image may be captured at a low resolution (e.g., 1x). At step 510, image segmentation may be performed to detect regions where the substance is present. This may delimit all regions, including areas of dust on the slide, pen marks, printed text (such as annotations), etc., with bounding boxes. At step 515, for each bounding box encompassing the region of interest, a sample presence probability score may be calculated. This can be performed by various algorithms, such as K-means, NLMD (non-local mean denoising), features such as the hue color space of each row, and segmentation models such as U-Net. This can be used to find the best row 520. The best row may include rows sandwiched between the upper and lower rows, and the sum of the weighted scores of those rows is the highest score regarding the presence of the sample. Once the best row is determined, the best (x,y) position can be determined. The best (x,y) position may include the (x,y) point that maximizes the probability of sample presence at that point, based on the boundaries, colors, etc. of the specimen. The optimal focus of that bounding box may be determined by collecting a Z-stack at the best (x,y) point 525. As used herein, a "Z-stack" is a plurality of images with varying focal distances at a specific (x,y) position. In some embodiments, the Z-stack may have a height smaller than the height of the slide. In some embodiments, the Z-stack may include 2, 3, 4, 5, 6, 7, 8, 9, 10, or more images captured at various focal distances. In some embodiments, the images captured as part of the Z-stack may be at intervals of less than 1 millimeter. For example, reducing the number of images captured can increase the speed, thus potentially improving efficiency. Capturing images closely can potentially determine the optimal focal distance with high precision. In some embodiments, due to these factors, it may be desirable to capture a relatively small number of images over a relatively small height.This can increase the importance of the selection of the focal length to capture a Z stack such that the object to be focused is located within the height of the Z stack. The focus pattern can be used to estimate the focal length at the (x, y) position where the Z stack is captured. By using such an estimate, the certainty that the object to be focused is within the height of the Z stack can be improved. In some embodiments, the distance of the Z stack can be identified and / or captured using the planarity of the camera field of view and / or at least two points along its row. Such planarity can include a focus pattern as described herein. In steps 530 and 535, that Z level may be used to scan the row and identify a focus pattern such as a plane. A focus pattern such as a plane may be identified using a set of points along the row as a function of the optimal focus of the points of that row. Such a plane may be used to estimate the focal length at other positions, such as the positions of adjacent rows. Such a plane may be recalculated based on additional data when new data is obtained. For example, the optimal focus is identified at additional points and the plane may be updated based on this additional data. Such a plane may also be used to estimate the focal length in another region of interest 540. This procedure can be repeated for all regions of interest on slide 545.
[0139] Referring now to FIGS. 6A-6C, the progression of a slide through the various steps described herein is illustrated. Slide 604 may include annotations 608A, 608B, dust 612A, 612B, and samples 616A, 616B. The steps described herein can address the problem of selecting the correct focus for the regions where the samples are present, despite the presence of dust and annotations. By dividing slide 604 into separate regions of interest 620A-620F, each segment can be scanned at a different focus that is optimal for that region. This makes the focus determination independent of the spatial distribution of the samples, dust, and annotations. In some embodiments, it is best to perform image classification to detect dust and annotations to avoid scanning downstream rather than during the scan. In some embodiments, this can address the problem of running an advanced model live on the scanning device during the scan and the risk of false positives when classifying a region as dust or annotation and skipping the scan of the associated region. Also, annotations that may be useful for downstream tasks may need to be scanned for use in downstream multimodal learning. In some embodiments, performing optimal focus determination following segmentation can be a beneficial approach for detecting all regions of interest at the optimal focus obtained by best row estimation.
[0140] Next, referring to FIG. 7, an exemplary embodiment of a method 700 for slide imaging is illustrated. One or more steps of method 700 may be implemented as described herein with reference to other figures, without limitation. One or more steps of method 700 may be implemented using at least one processor, without limitation.
[0141] Still referring to FIG. 7, in some embodiments, method 700 may include receiving at least one region of interest 705.
[0142] Referring still to FIG. 7, in some embodiments, method 700 may include capturing a first image of the slide at a first position within at least one region of interest 710 using at least one optical system.
[0143] Referring still to FIG. 7, in some embodiments, method 700 may include identifying a focus pattern as a function of a first image and a first position 715. In some embodiments, identifying the focus pattern includes identifying a row that includes the first position, capturing a plurality of first images at the first position, wherein each of the plurality of first images has a different focal distance, determining the first image that is optimally focused among the plurality of first images of optimal focus, and identifying the focus pattern using the focal distances of the plurality of optimally focused images in a set of points along the row. In some embodiments, the row further includes a second position, and the method includes capturing a plurality of second images at the second position using at least one processor and at least one optical system, wherein each of the plurality of second images has a different focal distance, determining the second image that is optimally focused among the plurality of second images of optimal focus using at least one processor, identifying the focus pattern using the focal distances of the optimally focused first image and the optimally focused second image using at least one processor, and further including extrapolating a third focal distance at a third position as a function of the focus pattern using at least one processor. In some embodiments, the third position is located outside the row. In some embodiments, the third position is located within a region of interest different from the first position. In some embodiments, identifying the row includes identifying the row based on a first row sample presence score from a first set of row sample presence scores. In some embodiments, identifying the row based on the first row sample presence score includes determining, from a second set of sample presence scores, the row having the highest sample presence score for which adjacent rows have the highest sample presence scores, wherein the second set of sample presence scores is determined using machine vision. In some embodiments, identifying the plane includes identifying a plurality of points, identifying a plurality of optimal foci at the plurality of points, identifying a subset of the plurality of points, and generating a plane as a function of the corresponding subset of optimal foci at those points.In some embodiments, identifying the focus pattern further includes updating the focus pattern, and updating the focus pattern includes identifying additional points and optimal foci at the additional points, and updating the focus pattern as a function of the additional points and the optimal foci at the additional points.
[0144] Still referring to FIG. 7, in some embodiments, method 700 may include extrapolating the focal length at a second position as a function of focus pattern 720.
[0145] Still referring to FIG. 7, in some embodiments, method 700 may include capturing a second image of the slide at a second position and a focal length 725 using at least one optical system. In some embodiments, capturing the second image includes capturing a plurality of images taken at a focal length based on the focus pattern, and constructing the second image from the plurality of images.
[0146] Still referring to FIG. 7, in some embodiments, method 700 may further include moving a movable element to a second position using an actuator mechanism.
[0147] Still referring to FIG. 7, in some embodiments, method 700 may further include identifying points within a row having a maximum point sample presence score using machine vision.
[0148] Still referring to FIG. 7, in some embodiments, method 700 includes capturing a low-magnification image of the slide using at least one processor and an optical system, the low-magnification image having a magnification lower than that of the first image, and identifying at least one region of interest within the low-magnification image using at least one processor and machine vision, and determining whether a sample is included in either the first image or the second image using at least one processor.
[0149] As will be apparent to one or more of ordinary skill in the art in the computer technology field, any one or more of the aspects and embodiments described in this specification can be appropriately implemented using one or more machines (e.g., one or more computing devices utilized as user computing devices for electronic documents, one or more server devices such as document servers) programmed in accordance with the teachings of this specification. Appropriate software coding can be readily made by a skilled programmer based on the teachings of this disclosure, as will be apparent to one of ordinary skill in the software technology field. The above-described aspects and implementations employing software and / or software modules can also include appropriate hardware for supporting the implementation of machine-executable instructions of the software and / or software modules.
[0150] Such software may be a computer program product employing a machine-readable storage medium. A machine-readable storage medium can store and / or encode a sequence of instructions for execution by a machine (e.g., a computing device) and can be any medium that causes a machine to execute any one of the methodologies and / or embodiments described herein. Examples of machine-readable storage media include, but are not limited to, magnetic disks, optical disks (e.g., CD, CD-R, DVD, DVD-R, etc.), magneto-optical disks, read-only memory "ROM" devices, random access memory "RAM" devices, magnetic cards, optical cards, solid state memory devices, EPROM, EEPROM, and any combination thereof. As used herein, a machine-readable medium is not limited to a single medium and is intended to include, for example, a collection of compact disks, a collection of physically separated media such as one or more hard disk drives combined with computer memory. A machine-readable storage medium as used herein does not include signal transmissions in transient form.
[0151] Such software can also include information (e.g., data) carried as a data signal on a data carrier such as a carrier wave. For example, machine-executable information can be a data carrier signal embodied in a data carrier that encodes a sequence of instructions for execution by a machine (e.g., a computing device) or a portion thereof, and any associated information (e.g., data structures and data) that causes the machine to execute any one of the methodologies and / or embodiments described herein.
[0152] Examples of computing devices include, but are not limited to, e - book reading devices, computer workstations, terminal computers, server computers, handheld devices (e.g., tablet computers, smartphones, etc.), web appliances, network routers, network switches, network bridges, any machine capable of executing a sequence of instructions specifying actions to be performed by that machine, and any combination thereof. In one example, a computing device may include and / or be included in a kiosk.
[0153] FIG. 8 shows a schematic representation of one embodiment of a computing device in an exemplary form of a computer system 800 in which a set of instructions for causing any one or more of the aspects and / or methodologies of the present disclosure to be executed by a control system can be executed. It is also contemplated that multiple computing devices can be utilized to implement a set of instructions specially configured to cause any one or more of the aspects and / or methodologies of the present disclosure to be executed on one or more of the devices. The computer system 800 includes a processor 804 and a memory 808 that communicate with each other and with other components via a bus 812. The bus 812 can include any of several types of bus structures, including, but not limited to, a memory bus, a memory controller, a peripheral bus, a local bus, and any combination thereof, using any of various bus architectures.
[0154] Processor 804 can include any suitable processor, such as, but not limited to, a processor incorporating a logic circuit that performs arithmetic and logical operations, such as an arithmetic logic unit (ALU), which is controlled by a state machine and can be directed by operation inputs from memory and / or sensors. Processor 804 may be configured according to, by way of non-limiting example, a von Neumann architecture and / or a Harvard architecture. Processor 804 may include, without limitation, a microcontroller, a microprocessor, a digital signal processor (DSP), a field programmable gate array (FPGA), a complex programmable logic device (CPLD), a graphics processing unit (GPU), a general-purpose GPU, a tensor processing unit (TPU), an analog or mixed-signal processor, a trusted platform module (TPM), a floating-point unit (FPU), and / or a system-on-chip (SoC), may incorporate these, and / or may be incorporated into these.
[0155] Memory 808 may include various components (e.g., machine-readable media), including, but not limited to, random access memory components, read-only components, and any combination thereof. In one example, a basic input / output system 816 (BIOS) including basic routines that help to transfer information between elements within computer system 800, such as during startup, may be stored in memory 808. Memory 808 can also include instructions (e.g., software) 820 that embody any one or more of the aspects and / or methodologies of the present disclosure (e.g., stored on one or more machine-readable media). In another example, memory 808 can further include any number of program modules, including, but not limited to, an operating system, one or more application programs, other program modules, program data, and any combination thereof.
[0156] The computer system 800 can also include a storage device 824. Examples of storage devices (e.g., storage device 824) include, but are not limited to, hard disk drives, magnetic disk drives, optical disk drives in combination with optical media, solid state memory devices, and any combination thereof. The storage device 824 can be connected to the bus 812 by an appropriate interface (not shown). Examples of interfaces include, but are not limited to, SCSI, Advanced Technology Attachment (ATA), Serial ATA, Universal Serial Bus (USB), IEEE 1394 (FIREWIRE (registered trademark)), and any combination thereof. In one example, the storage device 824 (or one or more of its components) can be removably interfaced with the computer system 800 (e.g., via an external port connector (not shown)). In particular, the storage device 824 and associated machine-readable medium 828 can provide non-volatile and / or volatile storage of machine-readable instructions, data structures, program modules, and / or other data for the computer system 800. In one example, the software 820 can reside entirely or partially within the machine-readable medium 828. In another example, the software 820 can reside entirely or partially within the processor 804.
[0157] The computer system 800 can also include an input device 832. In one example, a user of the computer system 800 can input commands and / or other information into the computer system 800 via the input device 832. Examples of the input device 832 include, but are not limited to, alphanumeric input devices (e.g., keyboards), pointing devices, joysticks, game pads, voice input devices (e.g., microphones, voice response systems, etc.), cursor control devices (e.g., mice), touch pads, optical scanners, video capture devices (e.g., still cameras, video cameras), touch screens, and any combination thereof. The input device 832 can be interface-connected to the bus 812 via any of a variety of interfaces (not shown) including, but not limited to, serial interfaces, parallel interfaces, game ports, USB interfaces, FIREWIRE interfaces, direct interfaces to the bus 812, and any combination thereof. The input device 832 can include a touch screen interface that can be part of or separate from a display 836 described below. The input device 832 can be used as a user selection device to select one or more graphical representations in the graphical interface as described above.
[0158] The user can also input commands and / or other information into the computer system 800 via a storage device 824 (e.g., a removable disk drive, a flash drive, etc.) and / or a network interface device 840. A network interface device such as the network interface device 840 can be utilized to connect the computer system 800 to one or more of various networks such as the network 844 and one or more remote devices 848 connected thereto. Examples of network interface devices include, but are not limited to, network interface cards (e.g., mobile network interface cards, LAN cards), modems, and any combination thereof. Examples of networks include, but are not limited to, wide area networks (e.g., the Internet, corporate networks), local area networks (e.g., networks associated with offices, buildings, campuses, or other relatively small geographical spaces), telephone networks, data networks associated with telephone / voice providers (e.g., data and / or voice networks of mobile communication providers), direct connections between two computing devices, and any combination thereof. A network such as the network 844 can employ wired and / or wireless communication modes. Generally, any network topology can be used. Information (e.g., data, software 820, etc.) can be communicated to and / or from the computer system 800 via the network interface device 840.
[0159] Computer system 800 may further include a video display adapter 852 that communicates a displayable image to a display device such as display device 836. Examples of display devices include, but are not limited to, liquid crystal displays (LCDs), cathode ray tubes (CRTs), plasma displays, light emitting diode (LED) displays, and any combination thereof. Display adapter 852 and display device 836 can be utilized in combination with processor 804 to provide a graphical representation of aspects of the present disclosure. In addition to the display device, computer system 800 may include one or more other peripheral output devices including, but not limited to, audio speakers, printers, and any combination thereof. Such peripheral output devices can be connected to bus 812 via peripheral interface 856. Examples of peripheral interfaces include, but are not limited to, serial ports, USB connections, FIREWIRE connections, parallel connections, and any combination thereof.
[0160] The foregoing has described in detail exemplary embodiments of the present invention. Various modifications and additions can be made without departing from the spirit and scope of the present invention. The features of each of the various embodiments described above can be combined as appropriate with the features of other described embodiments to provide various combinations of features in related new embodiments. Further, although a number of separate embodiments have been described above, what is described herein is merely illustrative of the application of the principles of the present invention. Further, specific methods herein may be illustrated and / or described as being performed in a particular order, but the order can be significantly varied within the scope of ordinary technology in order to implement the methods, systems, and software according to the present disclosure. Accordingly, this specification is intended to be construed as illustrative only and not to limit the scope of the present invention in any other way.
[0161] Exemplary embodiments are disclosed above and illustrated in the accompanying drawings. It will be understood by those skilled in the art that various changes, omissions, and additions can be made to what is specifically disclosed herein without departing from the spirit and scope of the present invention.
Claims
1. 1. An apparatus for imaging a slide, comprising: At least one optical system including an optical sensor; a slide port configured to hold the slide; At least one processor; a memory communicatively coupled to the at least one processor, the at least one processor comprising: Receiving at least one region of interest; capturing, using the at least one optical system, a plurality of first images of the slide at a first location within the at least one region of interest, each first image taken at a different focal length relative to a focal plane; identifying pixel arrays in a fixed direction in the plurality of first images as rows; identifying a set including a plurality of points along the row, and for each of the points determining an optimally focused image among the plurality of first images, thereby Identifying a focus pattern using a focal length of the optimally focused image; extrapolating a focal length of a second position based on the focal pattern; Using the at least one optical system, capture a second high resolution image of the slide at the second position and at the focal length of an area included in the area captured in the first image. and a memory storing instructions for configuring the An apparatus comprising:
2. 10. The apparatus of claim 1, further comprising an actuator mechanism mechanically connected to the at least one optical system, the actuator mechanism configured to move the at least one optical system to the second position.
3. The line further includes a second location, and the instructions cause the processor to: using the optical system to capture a plurality of second images of the slide at the second position, each image taken at a different focal length relative to a different focal plane; determining an optimally focused second image from the plurality of second images; identifying the focus pattern using the focal length of the first optimally focused image and the focal length of the second optimally focused image; Extrapolating a third focal length at a third position based on the focal pattern. The apparatus of claim 1 further configured to:
4. The apparatus of claim 3 , wherein the third location is located outside a unidirectional array of pixels identified as a row in the plurality of first images.
5. The apparatus of claim 3 , wherein the third location is located in a different region of interest than the first location.
6. 2. The apparatus of claim 1, wherein identifying the row comprises being based on a first row sample presence score that is a first set of indicators indicating a likelihood that a sample is present in a pixel array of a certain orientation in the plurality of first images.
7. 7. The apparatus of claim 6, wherein identifying the row based on the first row sample presence score includes identifying a row having a highest sample presence score among neighboring rows, the row having the highest sample presence score determined based on a second set of sample presence scores calculated using machine vision.
8. 2. The apparatus of claim 1, wherein the instructions further configure the processor to use machine vision to identify a point in the row having a maximum value among sample presence scores calculated for each point in the row.
9. identifying the focal pattern includes identifying a plane; identifying a plurality of points; and for each of the plurality of points, identifying a focal length of an optimally focused image of the plurality of first images as an optimal focus; generating a plane as a function of a subset of the plurality of points and a corresponding subset of best focus at those points; The apparatus of claim 1 , comprising:
10. Identifying the focus pattern further includes updating the focus pattern, where updating the focus pattern comprises: identifying an additional point; and for the additional point, identifying a focal length of an optimally focused image of the plurality of first images as an optimal focus; updating the focus pattern based on a calculation using the additional point and an optimum focus at the additional point; The apparatus of claim 9 , comprising:
11. Capturing the second image includes: capturing a plurality of images taken at focal lengths based on the focus pattern; constructing the second image from the plurality of images; and The apparatus of claim 1 , comprising:
12. The processor, capturing a low magnification image of the slide using the at least one optical system, the low magnification image having a magnification lower than a magnification of the first image; identifying the at least one region of interest in the low magnification image using machine vision; Determining whether the sample is included in either the first image or the second image The apparatus of claim 1 , further configured to:
13. 1. A method of imaging a slide, comprising: receiving, using at least one processor, at least one region of interest; capturing, using the at least one processor and at least one optical system, a plurality of first images of the slide at a first location within the at least one region of interest, each first image taken at a different focal length relative to a focal plane; Using the at least one processor, identifying pixel arrays in a fixed direction in the plurality of first images as rows; identifying a set of points along the row and for each of the points, determining an optimally focused image among the plurality of first images; and identifying a focus pattern using a focal length of the optimally focused image; extrapolating, using the at least one processor, a focal length at a second position based on the focal pattern; and capturing, using said at least one processor and said at least one optical system, a second image of said slide at said second position and at said focal length, said second image being a high resolution image of an area of said slide that is included in the area captured in said first image; The method includes:
14. The method of claim 13 , further comprising moving the optical system to the second position using an actuator mechanism.
15. the row further comprises a second location, and the method further comprises: capturing a plurality of second images of the slide at the second location using the at least one processor and the at least one optical system, each image taken at a different focal distance relative to a different focal plane; determining, using the at least one processor, an optimally focused second image of the plurality of second images; and identifying, using the at least one processor, the focal length of the first optimally focused image and the focal length of the second optimally focused image; extrapolating, using the at least one processor, a third focal length at a third position based on the focal pattern; The method of claim 13 further comprising:
16. The method of claim 15 , wherein the third location is located outside a unidirectional array of pixels identified as a row in the plurality of first images.
17. The method of claim 15 , wherein the third location is located in a different region of interest than the first location.
18. 14. The method of claim 13, wherein identifying the rows comprises: based on first row sample presence scores that are a first set of indicators indicative of the likelihood that a sample is present in a pixel array of a given orientation in the plurality of first images.
19. 20. The method of claim 18, wherein identifying the row based on the first row sample presence score includes identifying a row having a highest sample presence score among neighboring rows, the row having the highest sample presence score determined based on a second set of sample presence scores calculated using machine vision.
20. The method of claim 13, further comprising using machine vision to identify a point in the row that has a maximum value among the sample presence scores calculated for each point in the row.
21. identifying the focal pattern includes identifying a plane; identifying a plurality of points; and for each of the plurality of points, identifying a focal length of an optimally focused image of the plurality of first images as an optimal focus; generating a plane as a function of a subset of the plurality of points and a corresponding subset of best focus at those points; The method of claim 13, comprising:
22. Identifying the focus pattern further includes updating the focus pattern, where updating the focus pattern comprises: identifying an additional point; and for the additional point, identifying a focal length of an optimally focused image of the plurality of first images as an optimal focus; updating the focus pattern based on a calculation using the additional point and an optimum focus at the additional point; 22. The method of claim 21 , comprising:
23. Capturing the second image includes: capturing a plurality of images taken at focal lengths based on the focus pattern; constructing the second image from the plurality of images; and The method of claim 13, comprising:
24. capturing a low magnification image of the slide using at least one processor and the optical system, the low magnification image having a magnification lower than a magnification of the first image; identifying the at least one region of interest in the low magnification image using the at least one processor and machine vision; determining, using the at least one processor, whether the sample is included in either the first image or the second image; The method of claim 13 further comprising:
Citation Information
Patent Citations
Automated staining method and apparatus for biological materials
JP2010520487A
Image acquisition device and focusing method for the same
JP2014089411A
How to speed up modeling of digital slide scanners
JP2021524049A
System and Method for Enhanced Predictive Autofocusing
US20090195688A1
Optical scanning arrangement and method
US20200379232A1