Apparatus and method for slide imaging

The apparatus and method improve slide imaging by capturing initial images, identifying focal patterns, and adjusting focus to ensure clear imaging of regions of interest, addressing out-of-focus issues caused by obstacles in single-pass digitization.

JP2025128278APending Publication Date: 2025-09-02PRAMANA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025096565
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-28
Filing Date
2025-06-10
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

Single-pass slide digitization approaches face challenges in determining the best-focus Z plane due to obstacles like dust and pen marks, leading to out-of-focus issues in tissue imaging.

Method used

An apparatus and method that includes an optical system, processor, and memory to capture initial images, identify focal patterns, and extrapolate focal lengths for improved focus, allowing multiple image captures to ensure accurate focus on regions of interest.

Benefits of technology

Enhances imaging efficiency by minimizing false positives and ensuring clear focus on relevant areas, reducing errors and improving image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025128278000001_ABST
    Figure 2025128278000001_ABST
Patent Text Reader

Abstract

To provide techniques related to slide imaging.SOLUTION: An exemplary apparatus for real time image generation includes: at least one optical system; a slide port configured to hold a slide; an actuator mechanism mechanically connected to a mobile element; a user interface including an input interface and an output interface; and at least one processor configured to receive at least one region of interest, capture, using the at least one optical system, a first image of the slide at a first position within the at least one region of interest, identify a focus pattern as a function of the first image and the first position, extrapolate a focal distance for a second position as a function of the focus pattern, and capture, using the at least one optical system, a second image of the slide at the second position and at the focal distance.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates generally to the field of medical imaging, and more particularly to an apparatus and method for slide imaging. [Background technology]

[0002] When scanning slides, the cost of errors due to out-of-focus can be significant, especially when decisions are made based on the slide. Single-pass slide digitization approaches necessarily require determining the best-focus Z plane during a single scan, which forces decisions to be made in the presence of tissue, dust, pen marks, and other elements on the slide. For example, if dust is encountered, the scan will select the wrong Z reference plane, and the tissue beyond the dust will not be in focus. Similarly, pen marks will cause the tissue to be completely out-of-focus because the Z plane of the pen mark is above the coverslip. Summary of the Invention [Means for solving the problem]

[0003] In one aspect, an apparatus for imaging a slide may include at least one optical system including an optical sensor, a slide port configured to hold a slide, at least one processor, and a memory communicatively connected to the at least one processor, the memory storing instructions to configure the at least one processor to receive at least one region of interest, capture a first image of the slide at a first position within the at least one region of interest using the at least one optical system, identify a focal pattern as a function of the first image and the first position, extrapolate a focal length of a second position as a function of the focal pattern, and capture a second image of the slide at a second position and the focal length using the at least one optical system.

[0004] In another aspect, a method of imaging a slide may include receiving at least one region of interest using at least one processor; capturing a first image of the slide at a first location within the at least one region of interest using at least one processor and at least one optical system; identifying a focal pattern as a function of the first image and the first location using at least one processor; extrapolating a focal length of a second location as a function of the focal pattern using at least one processor; and capturing a second image of the slide at the second location and the focal length using at least one processor and the at least one optical system.

[0005] These and other aspects and features of non-limiting embodiments of the present invention will become apparent to those skilled in the art upon review of the following description of specific non-limiting embodiments of the present invention in conjunction with the accompanying drawings. [Brief explanation of the drawings]

[0006] For the purpose of illustrating the invention, the drawings show aspects of one or more embodiments of the invention. It is to be understood, however, that the invention is not limited to the precise arrangements and instrumentalities shown in the drawings. [Figure 1] FIG. 1 illustrates an exemplary apparatus for slide imaging. [Figure 2] FIG. 1 illustrates an exemplary machine learning model. [Figure 3] FIG. 1 illustrates an exemplary neural network. [Figure 4] FIG. 1 illustrates an exemplary neural network node. [Figure 5] FIG. 1 illustrates an exemplary method of slide imaging. [Figure 6A] FIG. 1 illustrates a slide including various features, identifying a region of interest, and scanning a line within the region of interest. [Figure 6B]FIG. 1 illustrates a slide including various features, identifying a region of interest, and scanning a line within the region of interest. [Figure 6C] FIG. 1 illustrates a slide including various features, identifying a region of interest, and scanning a line within the region of interest. [Figure 7] FIG. 1 illustrates an exemplary method of slide imaging. [Figure 8]

[0013] Figure 1 is a block diagram of a computing system that may be used to implement any one or more of the methodologies disclosed herein and any one or more portions thereof. The drawings are not necessarily to scale and may be illustrated with phantom lines, schematic representations, or fragmentary views. In certain instances, details that are not necessary for an understanding of the embodiments or that obscure other details may be omitted. DETAILED DESCRIPTION OF THE INVENTION

[0007] At a high level, aspects of the present disclosure relate to devices and methods for slide imaging. The devices described herein can generate images of a slide and / or a sample on the slide. In one embodiment, the device can capture a first image and identify one or more regions of interest. The region of interest may include features such as writing, dust, a sample, or the like. The sample may include tissue. In one embodiment, the device can identify a focus pattern in the region of interest. For example, the device can identify the focus pattern as a plane based on a plurality of points at which best focus is determined. In one embodiment, the device can determine which regions contain samples. In some embodiments, the determination of which regions contain samples may be made after one or more other steps described herein. In some embodiments, delaying the determination of which regions contain samples can make the slide imaging process more efficient. For example, efficiency can be improved because it is difficult to implement sophisticated models on a scanning device during scanning. In another example, efficiency can be improved because classifying a region as dust or an annotation and skipping scanning the region minimizes the risk of false positives. In other instances, it may be useful to scan annotations, so performing the steps in this order may be optimal. Exemplary embodiments illustrating aspects of the present disclosure are described below in the context of several specific examples.

[0008] Referring now to FIG. 1 , an exemplary embodiment of an apparatus 100 for slide imaging is illustrated. The apparatus 100 may include a computing device. The apparatus 100 may include a processor 104. The processor 104 may include, but is not limited to, any processor 104 described in this disclosure. The processor 104 may be included in a computing device. The apparatus 100 may include at least one processor 104 and a memory 108 communicatively coupled to the at least one processor 104, the memory 108 storing instructions 112 that configure the at least one processor 104 to perform one or more processes described herein. The computing device may include any computing device described in this disclosure, including, but not limited to, a microcontroller, a microprocessor, a digital signal processor (DSP), and / or a system-on-chip (SoC) described in this disclosure. The computing device may include, be included in, and / or communicate with a mobile device, such as a mobile phone or smartphone. A computing device may include a single computing device operating independently, or may include two or more computing devices operating in cooperation, parallel, serially, etc., where the two or more computing devices may be included together in a single computing device or may be included in two or more computing devices. A computing device may interface or communicate with one or more additional devices, as described in more detail below, via a network interface device. A network interface device may be utilized to connect a computing device to one or more of various networks and one or more devices. Examples of network interface devices include, but are not limited to, a network interface card (e.g., a mobile network interface card, a LAN card), a modem, and any combination thereof.Examples of networks include, but are not limited to, wide area networks (e.g., the Internet, enterprise networks), local area networks (e.g., networks associated with an office, building, campus, or other relatively small geographic space), telephone networks, data networks associated with telephone / voice providers (e.g., a mobile communications provider's data and / or voice network), direct connections between two computing devices, and any combination thereof. Networks can employ wired and / or wireless communication modes. In general, any network topology can be used. Information (e.g., data, software, etc.) can be communicated to and / or from computers and / or computing devices. Computing devices can include, but are not limited to, a computing device or cluster of computing devices at a first location and a second computing device or cluster of computing devices at a second location. Computing devices can include one or more computing devices specialized for data storage, security, traffic distribution for load balancing, etc. A computing device may distribute one or more computing tasks, as described below, across multiple computing devices that may operate in parallel, serially, redundantly, or any other manner used to distribute tasks or memory among computing devices. As a non-limiting example, a computing device may be implemented using a "shared nothing" architecture.

[0009] With continued reference to FIG. 1 , a computing device may be designed and / or configured to repeatedly perform any method, method step, or sequence of method steps in any embodiment described in this disclosure, in any order, and to any degree. For example, a computing device may be configured to repeatedly perform a single step or sequence until a desired or ordered result is achieved. The repetition of a step or sequence of steps may be performed iteratively and / or recursively using the output of a previous iteration as input for a subsequent iteration, aggregating the inputs and / or outputs of an iteration to produce an aggregate result, decreasing or decrementing one or more variables, such as global variables, and / or dividing a large processing task into a set of smaller processing tasks that are addressed iteratively. A computing device may perform any step or sequence of steps described in this disclosure in parallel, such as performing a step two or more times simultaneously and / or nearly simultaneously, using two or more parallel threads, processor cores, etc., and the division of tasks among parallel threads and / or processes may be performed according to any protocol suitable for dividing tasks among iterations. Those skilled in the art will recognize, upon reviewing this disclosure in its entirety, various ways in which steps, sequences of steps, processing tasks, and / or data may be subdivided, shared, or otherwise processed using iterative, recursive, and / or parallel processing.

[0010] Still referring to FIG. 1 , as used in this disclosure, “communicatively connected” means connected by a connection, attachment, or coupling that allows for the receipt and / or transmission of information between two or more entities. For example, but not limited to, this connection may be wired or wireless, direct or indirect, and between two or more components, circuits, devices, systems, etc., allowing for the reception and / or transmission of data and / or signals. The data and / or signals therebetween may include, but are not limited to, electrical, electromagnetic, magnetic, visual, audio, radio, and microwave data and / or signals, combinations thereof, among others. A communicative connection may be achieved, for example, but not limited to, by wired or wireless electronic, digital, or analog communication, directly or through one or more intervening devices or components. Furthermore, a communicative connection may include electrically coupling or connecting at least one output of one device, component, or circuit to at least one input of another device, component, or circuit. For example, but not limited to, this may be via a bus or other facility for intercommunication between computing device elements. A communicative connection may also include an indirect connection via, for example, but not limited to, a wireless connection, a wireless communication, a low-power wide area network, optical communication, magnetic coupling, capacitive coupling, optical coupling, etc. In some instances, the term "communicatively coupled" may be used in this disclosure instead of communicatively connected.

[0011] Still referring to FIG. 1 , in some embodiments, the device 100 may be used to generate images of the slide 116 and / or a sample on the slide 116. As used herein, a "slide" refers to a container or surface that holds a sample of interest. In some embodiments, the slide 116 may include a glass slide. In some embodiments, the slide 116 may include a formalin-fixed, paraffin-embedded slide. In some embodiments, the sample on the slide 116 may be stained. In some embodiments, the slide 116 may be substantially transparent. In some embodiments, the slide 116 may include a thin, flat, substantially transparent glass slide. In some embodiments, a transparent cover may be applied to the slide 116 such that the sample resides between the slide 116 and the cover. Samples may include, by non-limiting example, blood smears, cervical smears, bodily fluids, and non-biological samples. In some embodiments, the sample on the slide 116 may include tissue. In some embodiments, the sample on the slide 116 may be frozen.

[0012] Still referring to FIG. 1 , in some embodiments, the slide 116 and / or samples on the slide 116 may be illuminated. In some embodiments, the apparatus 100 may include a light source. As used herein, a "light source" is any device configured to emit electromagnetic radiation. In some embodiments, the light source may emit light having substantially one wavelength. In some embodiments, the light source may emit light having a range of wavelengths. The light source may emit, without limitation, ultraviolet, visible, and / or infrared light. In non-limiting examples, the light source may include a light-emitting diode (LED), an organic light-emitting diode (OLED), and / or other light emitter. Such a light source may be configured to illuminate the slide 116 and / or samples on the slide 116. In non-limiting examples, the light source may illuminate the slide 116 and / or samples on the slide 116 from below.

[0013] Still referring to FIG. 1 , in some embodiments, the device 100 may include at least one optical system 120. As used in this disclosure, an "optical system" is an arrangement of one or more components that act on or employ electromagnetic radiation. In non-limiting examples, the optical system may include light, such as electromagnetic radiation, visible light, infrared light, or ultraviolet light. The optical system may include one or more optical elements, including, but not limited to, lenses, mirrors, windows, filters, etc. The optical system may form an optical image corresponding to an optical object. For example, the optical system may form an optical image at or on an optical sensor, which may capture (e.g., digitize) the optical image. In some cases, the optical system may have at least one magnification. For example, the optical system may include an objective lens (e.g., a microscope objective lens) and one or more reimaging optical elements that together form the optical magnification. In some cases, optical magnification may be referred to as zooming. As used herein, an "optical sensor" refers to a device that measures light and converts the measured light into one or more signals, which may include, but are not limited to, one or more electrical signals. In some embodiments, the optical sensor 120 may include at least one photodetector. As used herein, a "photodetector" refers to a device that is sensitive to light and can thereby detect light. In some embodiments, the photodetector may include a photodiode, a photoresistor, a photosensor, a photovoltaic chip, or the like. In some embodiments, the optical sensor 120 may include multiple photodetectors. The optical sensor 120 may include, but is not limited to, a camera. The optical sensor 120 may be in electronic communication with at least one processor 104 of the device 100. As used in this disclosure, "electronic communication" refers to a shared data connection between two or more devices. In some embodiments, the device 100 may include two or more optical sensors 120.

[0014] Still referring to FIG. 1 , as used herein, “image data” refers to information representing at least one physical scene, space, and / or object. Image data may include, for example, information representing a sample, a slide 116, or a region of a sample or slide. In some cases, image data may be generated by a camera. “Image data” may be used interchangeably with “image” throughout this disclosure, and image is used as a noun. An image may be optical, such as, but not limited to, one in which at least one optical system is used to generate an image of an object. An image may be digital, such as, but not limited to, one represented as a bitmap. Alternatively, an image may include any medium capable of representing a physical scene, space, and / or object. Alternatively, when “image” is used as a verb in this disclosure, it refers to the generation and / or formation of an image.

[0015] Still referring to FIG. 1 , in some embodiments, the device 100 may include a slide port 140. In some embodiments, the slide port 140 may be configured to hold the slide 116. In some embodiments, the slide port 140 may include one or more alignment features. As used herein, a “alignment feature” is a physical characteristic that helps secure the slide in place and / or align the slide with other components of the device. In some embodiments, the alignment feature may include a component that secures the slide 116, such as a clamp, latch, clip, recess, or other fastener. In some embodiments, the slide port 140 may facilitate removal or insertion of the slide 116. In some embodiments, the slide port 140 may include a transparent surface through which light can pass. In some embodiments, the slide 116 may be placed on such a transparent surface and / or illuminated by light traveling through such a transparent surface. In some embodiments, the slide port 140 may be mechanically connected to the actuator mechanism 124, as described below.

[0016] Still referring to FIG. 1 , in some embodiments, the apparatus 100 may include an actuator mechanism 124. As used herein, an “actuator mechanism” is a mechanical component configured to change the relative position of the slide and the optical system. In some embodiments, the actuator mechanism 124 may be mechanically connected to the slide 116, such as the slide 116 in the slide port 140. In some embodiments, the actuator mechanism 124 may be mechanically connected to the slide port 140. For example, the actuator mechanism 124 may move the slide port 140 to move the slide 116. In some embodiments, the actuator mechanism 124 may be mechanically connected to at least one optical system 120. In some embodiments, the actuator mechanism 124 may be mechanically connected to a movable element. As used herein, the term “movable element” refers to any movable or portable object, component, or device within the apparatus 100, such as, but not limited to, a slide, a slide port, or an optical system. In some embodiments, the movable element may move to properly position the optical system 120 relative to the slide 116 so that the optical system 120 can capture an image of the slide 116 according to a set of parameters. In some embodiments, the actuator mechanism 124 may be mechanically connected to an item selected from the list consisting of the slide port 140, the slide 116, and at least one optical system 120. In some embodiments, the actuator mechanism 124 may be configured to change the relative position of the slide 116 and the optical system 120 by moving the slide port 140, the slide 116, and / or the optical system 120.

[0017] Still referring to FIG. 1 , the actuator mechanism 124 may include machine components responsible for moving and / or controlling a mechanism or system. The actuator mechanism 124, in some embodiments, may require a control signal and / or an energy source or power. In some cases, the control signal may be relatively low energy. Exemplary forms of the control signal include electrical potential or current, air pressure or flow, hydraulic fluid pressure or flow, mechanical force / torque or velocity, or even human power. In some cases, the actuator may have an energy or power source other than the control signal. This may include a primary energy source, which may include, for example, electrical power, hydraulic pressure, pneumatic pressure, mechanical power, etc. In some embodiments, upon receiving a control signal, the actuator mechanism 124 responds by converting source power into mechanical motion. In some cases, the actuator mechanism 124 may be understood as a form of automation or automatic control.

[0018] Still referring to FIG. 1 , in some embodiments, the actuator mechanism 124 may include a hydraulic actuator. A hydraulic actuator may be comprised of a cylinder or fluid motor that uses hydraulic force to facilitate mechanical movement. The output of the hydraulic actuator mechanism 124 may include mechanical movement, such as, but not limited to, linear, rotary, or oscillatory movement. In some embodiments, the hydraulic actuator may employ hydraulic fluid. Because liquids are potentially incompressible, hydraulic actuators may exert large forces. Furthermore, because force is equal to pressure multiplied by area, hydraulic actuators may function as force transducers with changes in area (e.g., the cross-sectional area of ​​the cylinder and / or piston). An exemplary hydraulic cylinder may be comprised of a hollow cylindrical tube through which a piston can slide. In some cases, the hydraulic cylinder may be considered single-acting. The term "single-acting" may be used when fluid pressure is applied substantially only to one side of the piston. Thus, a single-acting piston can move in only one direction. In some cases, a spring may be used to provide a return stroke for the single-acting piston. In some cases, the hydraulic cylinder may be double-acting. "Double acting" may be used when pressure is applied to substantially both sides of the piston. The force difference between the two sides of the piston causes the piston to move.

[0019] Still referring to FIG. 1 , in some embodiments, the actuator mechanism 124 may include a pneumatic actuator mechanism 124. In some cases, pneumatic actuators can generate large forces from relatively small changes in gas pressure. In some cases, pneumatic actuators can respond more quickly than other types of actuators, such as hydraulic actuators. Pneumatic actuators can use compressible fluids (e.g., air). In some cases, pneumatic actuators can operate with compressed air. Operation of hydraulic and / or pneumatic actuators includes control of one or more valves, circuits, fluid pumps, and / or fluid manifolds.

[0020] Still referring to FIG. 1 , in some cases, the actuator mechanism 124 may include an electric actuator. The electric actuator mechanism 124 may include either an electromechanical actuator, a linear motor, or the like. In some cases, the actuator mechanism 124 may include an electromechanical actuator. An electromechanical actuator can convert the rotational force of an electric rotary motor into linear motion, generating linear motion through a mechanism. Exemplary mechanisms include, but are not limited to, a rotary-to-translational converter, such as a belt, a screw, a crank, a cam, a linkage, or a scotch yoke. In some cases, control of the electromechanical actuator may include control of an electric motor; for example, a control signal may control one or more electric motor parameters to control the electromechanical actuator. Exemplary, non-limiting electric motor parameters include rotational position, input torque, speed, current, and potential. The electric actuator mechanism 124 may include a linear motor. A linear motor may differ from an electromechanical actuator in that power from a linear motor is output directly as translational motion, rather than being output as rotational motion and converted to translational motion. In some cases, a linear motor may incur less frictional loss than other devices. Linear motors may be designated into at least three different categories, such as flat linear motors, U-channel linear motors, and tubular linear motors. Linear motors may be directly controlled by control signals that control one or more linear motor parameters. Exemplary linear motor parameters include, but are not limited to, position, force, velocity, potential, and current.

[0021] Still referring to FIG. 1 , in some embodiments, the actuator mechanism 124 may include a mechanical actuator mechanism 124. In some cases, the mechanical actuator mechanism 124 may function to perform movement by converting one type of movement, such as rotational movement, into another type of movement, such as linear movement. An exemplary mechanical actuator includes a rack and pinion. In some cases, a mechanical power source, such as a power take-off, may serve as a power source for the mechanical actuator. The mechanical actuator may employ any number of mechanisms, including, for example, but not limited to, gears, rails, pulleys, cables, linkages, etc.

[0022] Still referring to FIG. 1 , in some embodiments, the actuator mechanism 124 can be in electronic communication with an actuator control. As used herein, an “actuator control” is a system configured to operate the actuator mechanism to position the slide and the optical system in a desired relative position. In some embodiments, the actuator control can operate the actuator mechanism 124 based on input received from the user interface 136. In some embodiments, the actuator control can be configured to operate the actuator mechanism 124 so that the optical system 120 is positioned to capture an image of the entire sample. In some embodiments, the actuator control can be configured to operate the actuator mechanism 124 so that the optical system 120 is positioned to capture an image of a region of interest, a particular horizontal row, a particular point, a particular depth of focus, etc. Electronic communication between the actuator mechanism 124 and the actuator control can include the transmission of signals. For example, the actuator control can generate physical movement of the actuator mechanism in response to an input signal. In some embodiments, the input signal can be received by the actuator control from the processor 104 or the input interface 128.

[0023] Still referring to FIG. 1 , a “signal” as used in this disclosure is any understandable representation of data, for example, from one device to another. Signals may include optical signals, hydraulic signals, pneumatic signals, mechanical signals, electrical signals, digital signals, analog signals, etc. In some cases, signals may be used to communicate with a computing device, for example, via one or more ports. In some cases, signals may be transmitted and / or received by a computing device, for example, via an input / output port. Analog signals may be digitized, for example, by an analog-to-digital converter. In some cases, analog signals may be processed, for example, by analog signal processing steps described in this disclosure, before being digitized. In some cases, digital signals may be used to communicate between two or more devices, including, but not limited to, computing devices. In some cases, digital signals may be communicated via one or more communication protocols, including, but not limited to, Internet Protocol (IP), Controller Area Network (CAN) protocol, serial communication protocols (e.g., Universal Asynchronous Receiver Transmitter [UART]), parallel communication protocols (e.g., IEEE 128 [printer port]), etc.

[0024] Still referring to FIG. 1 , in some embodiments, the device 100 can perform one or more signal processing steps on a signal. For example, the device 100 can analyze, modify, and / or synthesize a signal representing data to improve the signal, for example, by improving transmission, storage efficiency, or signal-to-noise ratio. Exemplary methods of signal processing can include analog, continuous-time, discrete, digital, nonlinear, statistical, etc. Analog signal processing can be performed on non-digitized or analog signals. Exemplary analog processing can include passive filters, active filters, summing mixers, integrators, delay lines, companders, multipliers, voltage-controlled filters, voltage-controlled oscillators, phase-locked loops, etc. Continuous-time signal processing can sometimes be used to process signals that vary continuously within a domain, e.g., the time domain. Exemplary, non-limiting continuous-time processes can include time-domain processing, frequency-domain processing (Fourier transform), and complex frequency-domain processing. Discrete-time signal processing can be used when a signal is sampled at non-continuous or discrete time intervals (i.e., quantized in time). Analog discrete-time signal processing can process signals using the following exemplary circuits: sample and hold circuits, analog time division multiplexers, analog delay lines, and analog feedback shift registers. Digital signal processing can be used to process digitized discrete-time sampled signals. Generally, digital signal processing can be performed by a computing device or other specialized digital circuitry, such as, but not limited to, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a specialized digital signal processor (DSP). Digital signal processing can be used to perform any combination of typical arithmetic operations, such as fixed-point, floating-point, real-valued, complex-valued, multiplication, addition, etc. Digital signal processing can also operate circular buffers and look-up tables.Further non-limiting examples of algorithms that may be implemented according to digital signal processing techniques include fast Fourier transforms (FFTs), finite impulse response (FIR) filters, infinite impulse response (IIR) filters, and adaptive filters such as Wiener filters and Kalman filters. Statistical signal processing may be used to exploit statistical properties to process signals as random functions (i.e., stochastic processes). For example, in some embodiments, a signal may be modeled with a probability distribution that represents noise, which may be used to reduce the noise in the processed signal.

[0025] 1, in some embodiments, the device 100 may include a user interface 136. The user interface 136 may include the output interface 132 and the input interface 128.

[0026] Still referring to FIG. 1 , in some embodiments, output interface 132 may include one or more elements through which device 100 may communicate information to a user. In a non-limiting example, output interface 132 may include a display. The display may include a high-resolution display. The display may be capable of outputting images, videos, etc. to the user. In another non-limiting example, output interface 132 may include a speaker. The speaker may be capable of outputting audio to the user. In another non-limiting example, output interface 132 may include a haptic device. The speaker may be capable of outputting haptic feedback to the user.

[0027] Still referring to FIG. 1 , in some embodiments, the optical system 120 may include a camera. In some cases, the camera may include one or more optical systems. Exemplary, non-limiting optical systems include spherical lenses, aspherical lenses, reflectors, polarizers, filters, windows, aperture stops, etc. In some embodiments, one or more optical systems associated with the camera may be adjusted to change the camera's zoom, depth of field, and / or focal length, in non-limiting examples. In some embodiments, one or more of such settings may be configured to detect features of a sample on the slide 116. In some embodiments, one or more of such settings may be configured based on a parameter set, as described below. In some embodiments, the camera may capture images with a shallow depth of field. In a non-limiting example, the camera may capture images focused on a first depth of the sample and out of focus on a second depth of the sample. In some embodiments, an autofocus mechanism may be used to determine the focal length. In some embodiments, the focal length may be set by a parameter set. In some embodiments, the camera may be configured to capture multiple images at different focal lengths. In a non-limiting example, the camera can capture multiple images at different focal lengths, such that an image is captured at each focal depth of the sample in at least one image. In some embodiments, the at least one camera can include an image sensor. Exemplary, non-limiting image sensors include digital image sensors, such as, but not limited to, charge-coupled device (CCD) sensors and complementary metal-oxide semiconductor (CMOS) sensors. In some embodiments, the camera can be sensitive in the non-visible range of electromagnetic radiation, such as, but not limited to, infrared.

[0028] Still referring to FIG. 1 , in some embodiments, input interface 128 may include controls for operating device 100. Such controls may be operated by a user. Input interface 128 may include, in non-limiting examples, a camera, a microphone, a keyboard, a touchscreen, a mouse, a joystick, foot pedals, buttons, dials, etc. Input interface 128 may accept, in non-limiting examples, mechanical input, voice input, visual input, text input, etc. In some embodiments, voice input to input interface 128 may be interpreted using automatic speech recognition functionality, allowing a user to control device 100 via voice. In some embodiments, input interface 128 may approximate the controls of a microscope.

[0029] Still referring to FIG. 1 , in some embodiments, speech input may be processed using automatic speech recognition. In some embodiments, automatic speech recognition may require training (i.e., enrollment). In some cases, training an automatic speech recognition model may require an individual speaker to read text or isolated vocabulary. In some cases, speech training data may include speech components having audible linguistic content, which content is known a priori by the computing device. Thus, the computing device can train the automatic speech recognition model according to training data including audible linguistic content that correlates to the known content. In this manner, the computing device can analyze a person's particular speech and train the automatic speech recognition model for that person's speech, resulting in improved accuracy. Alternatively or additionally, in some cases, the computing device may include a speaker-independent automatic speech recognition model. As used in this disclosure, a “speaker-independent” automatic speech recognition process does not require training for an individual speaker. Conversely, as used in this disclosure, an automatic speech recognition process that employs training specific to an individual speaker is “speaker-dependent.”

[0030] Still referring to FIG. 1 , in some embodiments, the automatic speech recognition process may perform voice recognition or speaker identification. As used in this disclosure, “voice recognition” refers to identifying a speaker from audio content, rather than what the speaker says. In some cases, a computing device may first recognize a speaker of linguistic audio content and then automatically recognize the speaker's voice, for example, by a speaker-dependent automatic speech recognition model or process. In some embodiments, the automatic speech recognition process may be used to authenticate or verify the identity of a speaker. In some cases, the speaker may or may not include the subject. For example, the subject may speak during audio input, but other people may speak as well.

[0031] 1 , in some embodiments, the automatic speech recognition process may include one or all of acoustic modeling, language modeling, and statistically-based speech recognition algorithms. In some cases, the automatic speech recognition process may employ hidden Markov models (HMMs). As will be described in more detail below, language modeling, such as that employed in natural language processing applications such as document classification and statistical machine translation, may also be employed in the automatic speech recognition process.

[0032] Still referring to FIG. 1 , an exemplary algorithm employed for automatic speech recognition may include or be based on a hidden Markov model. A hidden Markov model (HMM) may include a statistical model that outputs a sequence of symbols or quantities. HMMs may be used for speech recognition because speech signals may be viewed as piecewise stationary or short-term stationary signals. For example, on short time scales (e.g., 10 milliseconds), speech may be approximated as a stationary process. Speech (i.e., audible speech content) may be understood as a Markov model for many probabilistic purposes.

[0033] Still referring to FIG. 1 , in some embodiments, HMMs can be trained automatically, are relatively simple to use, and are computationally feasible. In an exemplary automatic speech recognition process, a hidden Markov model can output a sequence of n-dimensional real-valued vectors (where n is a small integer, such as 10) at a rate of approximately one vector every 10 milliseconds. The vectors may be composed of cepstral coefficients. Cepstral coefficients must use the spectral domain. Cepstral coefficients are found by Fourier transforming a short time window of speech to generate a spectrum, decorrelating the spectrum using a cosine transform, and taking the first (i.e., most significant) coefficients. In some cases, the HMM may have a statistical distribution at each state that is a mixture of diagonal covariance Gaussian distributions, yielding a likelihood for each observation vector. In some cases, each word or phoneme may have a different output distribution. An HMM for a sequence of words or phonemes can be created by concatenating HMMs for separate words or phonemes.

[0034] Still referring to FIG. 1 , in some embodiments, the automatic speech recognition process may use various combinations of techniques to improve results. In some cases, the large vocabulary automatic speech recognition process may include context-dependent phonemes. For example, in some cases, phonemes with different left and right contexts may have different recognition as states of the HMM. In some cases, the automatic speech recognition process may use cepstrum normalization to normalize for different speakers and recording conditions. In some cases, the automatic speech recognition process may use vocal tract length normalization (VTLN) for gender normalization and maximum likelihood linear regression (MLLR) for more general speaker adaptation. In some cases, the automatic speech recognition process may determine so-called delta and delta-delta coefficients to capture speech dynamics and may use heteroscedastic linear discriminant analysis (HLDA). In some cases, the automatic speech recognition process may use splicing and projections based on linear discriminant analysis (LDA), which may include heteroscedastic linear discriminant analysis or global semi-coupled covariance transform (also known as maximum likelihood linear transform [MLLT]). In some cases, automatic speech recognition processes may forego a purely statistical approach to HMM parameter estimation and instead use discriminative training techniques that optimize some classification-related measure of the training data, examples of which include maximum mutual information (MMI), minimum classification error (MCE), and minimum phoneme error (MPE).

[0035] Still referring to FIG. 1 , in some embodiments, an automatic speech recognition process may be said to decode speech (i.e., audible linguistic content). Speech decoding occurs when an automatic speech recognition system is presented with a new utterance and must calculate the most likely sentence. In some cases, speech decoding may include a Viterbi algorithm. The Viterbi algorithm may include a dynamic programming algorithm to obtain a maximum a posteriori estimate of the most likely sequence of hidden states (i.e., the Viterbi path) that results in an observed sequence of events. The Viterbi algorithm may be employed in the context of Markov sources and hidden Markov models. The Viterbi algorithm may be used, for example, to find the best path using a statically created combinatorial hidden Markov model (e.g., a finite-state transducer [FST] approach) or using a dynamically created combinatorial hidden Markov model that has information from both an acoustic model and a language model.

[0036] Still referring to FIG. 1 , in some embodiments, decoding speech (i.e., audible language content) may include considering not only the best candidate but also good candidates when presented with a new utterance. In some cases, a refined scoring function (i.e., rescoring) may be used to evaluate each of the set of good candidates and select the best candidate according to this refined score. In some cases, the set of candidates may be maintained either as a list (i.e., an N-best list approach) or as a subset of the model (i.e., a lattice). In some cases, rescoring may be performed by optimizing Bayes risk (or an approximation thereof). In some cases, rescoring may include optimizing the sentence (including keywords) that minimizes the expected value of a given loss function with respect to all possible transcriptions. For example, rescoring enables the selection of a sentence that minimizes the average distance to other possible sentences weighted by their estimated probabilities. In some cases, the loss function employed may include the Levenshtein distance, although, for example, different distance calculations may be performed for specific tasks. In some cases, the set of candidates may be pruned to maintain tractability.

[0037] Still referring to FIG. 1 , in some embodiments, an automatic speech recognition process can employ a dynamic time warping (DTW)-based approach. Dynamic time warping can include an algorithm that measures the similarity between two sequences that differ in time or speed. For example, a person may walk slowly in one video and quickly in another, and even if they accelerate and decelerate during the same observation, similarities in their walking patterns can be detected. DTW has been applied to video, audio, and graphics; in fact, any data that can be converted to a linear representation can be analyzed with DTW. In some cases, DTW can be used by an automatic speech recognition process to accommodate different speaking (i.e., audible language content) rates. In some cases, DTW can enable a computing device to find the best match between two given sequences (e.g., time series) under certain constraints. That is, in some cases, sequences can be nonlinearly "warped" to match each other. In some cases, DTW-based sequence alignment methods can be used in the context of hidden Markov models.

[0038] Still referring to FIG. 1 , in some embodiments, an automatic speech recognition process may include a neural network. The neural network may include any neural network, such as those disclosed with reference to FIGS. 2-4 . In some cases, neural networks may be used for automatic speech recognition, including phoneme classification, multi-objective evolutionary phoneme classification, isolated word recognition, audio-visual speech recognition, audio-visual speaker recognition, and speaker adaptation. In some cases, neural networks employed in automatic speech recognition make fewer explicit assumptions about feature statistical properties than HMMs, and therefore may possess several qualities that make them attractive recognition models for speech recognition. When used to estimate the probability of speech feature segments, neural networks can provide discriminative training in a natural and efficient manner. In some cases, neural networks may be used to effectively classify audible linguistic content over short time intervals, such as individual phonemes and isolated words. In some embodiments, neural networks may be employed by automatic speech recognition processes for preprocessing, feature transformation, and / or dimensionality reduction, for example, prior to HMM-based recognition. In some embodiments, long short-term memory (LSTM) and related recurrent neural networks (RNNs) and time-delay neural networks (TDNNs) may be used for automatic speech recognition, e.g., over longer time intervals for continuous speech recognition.

[0039] 1, in some embodiments, the device 100 captures a first image of the slide 116 at a first position. In some embodiments, the first image may be captured using at least one optical system 120.

[0040] 1 , in some embodiments, capturing a first image of slide 116 at a first position may include using actuator mechanism 124 and / or actuator control to move optical system 120 and / or slide 116 to a desired position. In some embodiments, the first image includes an image of the entire sample and / or entire slide 116. In some embodiments, the first image includes an image of a region of the sample. In some embodiments, the first image includes a wider-angle image than a second image (described below). In some embodiments, the first image may include a lower-resolution image than the second image.

[0041] Still referring to FIG. 1 , in some embodiments, the apparatus 100 can identify at least one region of interest in the first image. In some embodiments, machine vision may be used to identify the at least one region of interest. As used herein, a "region of interest" is a particular area within a slide or a digital image of a slide in which features are detected. Features may include, in non-limiting examples, a sample, dust, writing on the slide, cracks in the slide, air bubbles, etc.

[0042] Still referring to FIG. 1 , in some embodiments, the device 100 may include a machine learning module 144. Machine learning will be described with reference to FIG. 2 . In some embodiments, the device 100 may use an ROI identification machine learning model 148 to identify at least one region of interest. In some embodiments, the ROI identification machine learning model 148 may be trained using supervised learning. In some embodiments, the ROI identification machine learning model 148 may include a classifier. The ROI identification machine learning model 148 may be trained on a dataset including example images of slides associated with example regions of the images in which features are present. Such a training dataset may be collected, for example, by collecting slide imaging device data regarding which regions of slide images an expert focuses on or zooms in on. Once trained, the ROI identification machine learning model 148 may accept slide images as input and output data regarding the location of any regions of interest present. In some embodiments, a neural network, such as a convolutional neural network, may be used to identify at least one region of interest. For example, a convolutional neural network may be used to detect edges in slide images, and at least one region of interest may be identified based on the presence of the edges. In some embodiments, at least one region of interest may be identified as a function of brightness and / or color difference compared to the brightness and / or color of a background. In some embodiments, device 100 may identify the region of interest using a classifier. In some embodiments, segments of the image may be input to the classifier, which may categorize the image segments based on whether a region of interest is present. In some embodiments, the classifier may output a score indicating the degree to which a region of interest is detected and / or a confidence level that a region of interest exists. In some embodiments, features may be detected using a neural network or other machine learning model trained to detect features and / or objects.For example, edges, corners, blobs, or ridges may be detected, and whether a location is determined to be within the region of interest may depend on the detection of such features. In some embodiments, machine learning models, such as support vector machine techniques, may be used to determine features based on the detection of features such as edges, corners, blobs, or ridges.

[0043] Still referring to FIG. 1 , in some embodiments, device 100 can identify at least one region of interest as a function of user input. For example, a user can modify a setting for the degree of sensitivity for detecting regions of interest. In this example, if the user input indicates low sensitivity, device 100 can only detect large regions of interest. This can include, for example, ignoring potential regions of interest below a certain size. In another example, this can include applying a machine learning model, such as a classifier, to the image (or a segment of the image) and identifying the image (or a segment of the image) as a region of interest only if the machine learning model outputs a score higher than a threshold. The score indicates the degree to which a region of interest is detected and / or the confidence that a region of interest exists. In some embodiments, the image can be divided into smaller segments, and the segments can be analyzed to determine whether at least one region of interest is present.

[0044] Still referring to FIG. 1 , in some embodiments, device 100 may receive at least one region of interest. In some embodiments, device 100 may receive at least one region of interest without first capturing a first image. In a non-limiting example, a user may input the region of interest. In some embodiments, device 100 may capture a first image and apply the received region of interest to the first image. This may be done, for example, if the region of interest is received before the first image is captured. In some embodiments, device 100 may capture the first image as a function of the region of interest. In a non-limiting example, device 100 may receive a region of interest from a user through user input and may capture the first image at a first location within the region of interest.

[0045] 1 , apparatus 100 may identify a focus pattern. In some embodiments, apparatus 100 may identify a focus pattern in at least one region of interest, such as each region of interest detected as described herein. In some embodiments, identifying a focus pattern may include identifying a row, identifying a point within the row, determining a best focus at the point, and / or identifying a plane.

[0046] Still referring to FIG. 1, as used herein, a "row" of a digital image of a slide is a segment of the digital image of a slide between two parallel lines. A row may include, for example, a row of pixels having a width of one pixel. In another example, a row may have a width of multiple pixels. A row may or may not cross a grid of pixels diagonally. As used herein, a "point" on a digital image of a slide refers to a specific location within the digital image of a slide. For example, in a digital image composed of a grid of pixels, a point may have a specific (x,y) location. As used herein, "best focus" is the focal length at which the object being focused on is in focus.

[0047] Still referring to FIG. 1 , in some embodiments, the device 100 may identify rows within a region of interest. The row may be identified based on a first row sample presence score. The first row sample presence score may be identified using machine vision. The first row sample presence score may be identified based on the output of the row identification machine learning model 152. The first row sample presence score may be identified by determining the row sample presence scores of one or more adjacent rows. For example, the first row sample presence score may be determined as a function of the second row sample presence score and the third row sample presence score, where the second and third row sample presence scores are based on rows adjacent to the row of the first row sample presence score. As used herein, a "sample presence score" is a value that represents or estimates the likelihood of a sample being present at a location. As used herein, a "row sample presence score" is a sample presence score for a location that is a row. The sample presence score does not necessarily represent the likelihood of a sample being present as a percentage between 0 and 100%. For example, if the row sample presence score for the first row is 2000 and the row sample presence score for the second row is 3000, this indicates that the second row is more likely to contain a sample than the first row. The first row sample presence score may be identified based on, in non-limiting examples, the sum of the row sample presence scores of adjacent rows, and / or the weighted sum of the row sample presence scores of adjacent rows, and / or the minimum of the row sample presence scores of adjacent rows. In some embodiments, a first set of row sample presence scores is identified using a machine learning model, such as row identification machine learning model 152, and a second set of row sample presence scores is identified based on the row sample presence scores of rows adjacent to the first set of row sample presence scores. In some embodiments, such second set of row sample presence scores may be used to identify the best row. The row sample presence scores may be identified by determining the sample presence scores of rows that are not directly adjacent to the row in question.For example, line sample presence scores may be determined for the line in question, adjacent lines, and lines one line away from the line in question, and each of these line sample presence scores may be used to identify the line (e.g., using a weighted sum of the sample presence scores). The sample presence scores may be determined using machine vision. In some embodiments, the line sample presence scores may be determined based on a section of the line that does not extend completely across the image and / or slide. For example, the line sample presence score may be determined relative to the width of the region of interest. In another example, the line sample presence score may be determined for a section of the line having a certain pixel width.

[0048] Still referring to FIG. 1 , in some embodiments, identifying a focal pattern may include identifying a row that includes a particular (X,Y) location, such as a location where an optimal focal length is calculated. In some embodiments, identifying a focal pattern may include capturing multiple images at such locations, each of the multiple images having a different focal length. Such images may represent a Z-stack, as described further below. Identifying a focal pattern may further include determining an optimally focused image from among the multiple images. The focal length of such image may be determined to be the optimal focal length for that (X,Y) location. In this manner, an optimal focal length may be determined for multiple points in the row. A focal pattern may be identified using the focal lengths of multiple optimally focused images at multiple points in such row.

[0049] Still referring to FIG. 1 , in some embodiments, the row may further include a second position. The apparatus 100 may capture multiple second images at such second positions, each of the multiple second images having a different focal length. The apparatus 100 may determine an optimally focused second image from the multiple best-focused second images. The apparatus 100 may identify a focus pattern using the focal lengths of the optimally focused first image and the optimally focused second image. For example, the focus pattern may be determined to include a line connecting the two points. The apparatus 100 may extrapolate a third focal length for the third position as a function of the focus pattern. In some embodiments, extrapolating the third focal length may include using the first or second focal length as the third focal length. In some embodiments, the extrapolation may include linear extrapolation, polynomial extrapolation, conic extrapolation, geometric extrapolation, etc. In some embodiments, such a third position may be located outside the row that includes the first position and / or the second position. In some embodiments, such a third location may be located within a different region of interest than the first location and / or the second location.

[0050] Still referring to FIG. 1 , in some embodiments, the sample presence score may be determined using a line identification machine learning model 152. In some embodiments, the line identification machine learning model 152 may be trained using supervised learning. The line identification machine learning model may be trained on a dataset containing examples of lines from images of slides that are associated with whether or not a sample is present. Such a dataset may be collected, for example, by capturing images of slides, manually identifying which lines contain samples, and extracting lines from the larger image. Once trained, the line identification machine learning model 152 can accept as input lines from images of slides, such as lines from a region of interest, and output a determination of whether or not a sample is present and / or a sample presence score, such as a line sample presence score.

[0051] Still referring to FIG. 1 , in some embodiments, device 100 can identify points within a row, such as the row identified above. In some embodiments, points may be identified using machine vision. In some embodiments, points may be identified based on the point within a particular row having the highest point sample presence score. As used herein, a "point sample presence score" is a sample presence score for a point. The point sample presence score may be determined, in non-limiting examples, based on the color of the point and / or surrounding pixels, or whether the point is inside or outside a potential specimen boundary (which may be determined, for example, by identifying edges in the image and regions of the image surrounded by the edges). In another non-limiting example, the point sample presence score may be determined based on the distance between the point and the edge of the slide.

[0052] Still referring to FIG. 1 , in some embodiments, points may be identified using a point identification machine learning model 156. In some embodiments, the point identification machine learning model 156 may be trained using supervised learning. The point identification machine learning model may be trained on a dataset containing examples of points from images of slides associated with whether a sample is present. Such a dataset may be collected, for example, by capturing images of slides, manually identifying which points contain samples, and extracting points from the larger image. Once trained, the point identification machine learning model 156 can accept points from images of slides, such as points from a region of interest, as input and output a determination of whether a sample is present and / or a point sample presence score. In some embodiments, the device 100 can identify points from rows based on which points have the highest point sample presence score.

[0053] Still referring to FIG. 1 , in some embodiments, device 100 can determine best focus at a point, such as the point identified above. In some embodiments, best focus may be determined using an autofocus mechanism. In some embodiments, best focus may be determined using a rangefinder. In some embodiments, actuator mechanism 124 can move optical sensor 120 and / or sliding port 140 to cause the autofocus mechanism to focus on a desired location. In some embodiments, the autofocus mechanism may be capable of focusing on multiple points within a frame and can select which point to focus on based on the point identified above. In some embodiments, one or more camera parameters other than focus can be adjusted to improve focus and / or improve the image. For example, aperture can be adjusted to change the degree to which a point is focused. For example, aperture can be adjusted to increase the depth of field to focus on a point. Best focus can be expressed as focal length or depth of focus, in non-limiting examples.

[0054] Still referring to FIG. 1 , in some embodiments, the device 100 may identify a focus pattern based on the best focus and / or point. The best focus and point may be expressed as positions in three-dimensional space. For example, the X and Y coordinates (horizontal axis) may be determined based on the location of the point on the slide and / or the location of the point on the image. The Z coordinate may be determined based on the best focus. For example, the focal length of the best focus may be used as the Z coordinate. As used herein, a "focus pattern" is a pattern that approximates the best focus level at multiple points, including at least one point where best focus has not been measured. One or more (X, Y, Z) coordinates may be used to determine the focus pattern. One or more default parameters may be used to determine the focus pattern (e.g., horizontal as a default if there is insufficient data to determine otherwise). For example, a single (X, Y, Z) coordinate may be determined, and the focus pattern may be determined as a plane extending horizontally in the X and Y directions, with the Z level held constant. In some embodiments, the focus pattern varies vertically (in the Z direction). For example, two (X, Y, Z) coordinates may be determined, and a focal pattern may be determined that includes a line connecting the two (X, Y, Z) positions and extends horizontally (forming a plane containing the line) when moved horizontally perpendicular to the line. In another example, three (X, Y, Z) coordinates may be determined, and the focal pattern may be determined as a plane that includes all three positions. In another example, several (X, Y, Z) coordinates may be determined, and a regression algorithm may be used to determine a plane that best fits the coordinates. For example, least-squares regression may be used. In some embodiments, the focal pattern is not planar. The focal pattern may include a surface, such as a surface in three-dimensional space. The focal pattern may include one or more curves, bumps, edges, etc. For example, the focal pattern may include a first plane in a first region, a second plane in a second region, and an edge where the planes intersect. In another example, the focus pattern may include a curved and / or bumpy surface where the Z level of each location on the focus pattern surface is determined based on the (X, Y, Z) coordinates of neighboring locations and the distance to those Z levels.In another example, the focus pattern may include multiple shapes with (X, Y, Z) coordinates as vertices and lines between the (X, Y, Z) coordinates as boundaries. In some embodiments, the focus pattern may have only one Z value for each (X, Y) coordinate. In some embodiments, which points are evaluated for their best focus may be determined as a function of, by way of non-limiting example, the desired density of points within the region of interest, the likelihood of a sample being present (such as output from the ML model described above), and / or user input. For example, a user may manually select points and / or input a desired point density. In some embodiments, the focus pattern may be updated as additional points are scanned. For example, the focus pattern may be in the shape of a plane based on 10 (X, Y, Z) coordinates, and an 11th (X, Y, Z) coordinate may be scanned, and the focus pattern may be recalculated and / or updated to account for the new coordinates. In another example, the focus pattern may start as a plane based on a single (X, Y, Z) coordinate and be updated as additional (X, Y, Z) coordinates are identified. In some embodiments, the focus pattern may be updated for every additional (X, Y, Z) coordinate. In some embodiments, the focus pattern may be updated at a rate less than the rate of (X, Y, Z) coordinate identification. In non-limiting examples, the focus pattern may be updated every 2, 3, 4, 5, or more (X, Y, Z) coordinates.

[0055] Still referring to FIG. 1 , in some embodiments, the data used to identify the focus pattern may be filtered. In some embodiments, one or more outliers may be removed. For example, if nearly all (X,Y,Z) points suggest a focus pattern in the shape of a plane, and a single (X,Y,Z) point has a Z value that is significantly different from that estimated from the plane, that (X,Y,Z) point may be removed. In another example, (X,Y,Z) points identified as focusing on features other than the sample may be removed. For example, (X,Y,Z) points identified as focusing on an annotation may be removed.

[0056] Still referring to FIG. 1 , in some embodiments, the process described above can be used to determine a focus pattern for each region of interest. For regions of interest that include a sample, this results in determining a focus pattern based on one or more points that include the sample. This can help efficiently identify the focus pattern so that follow-up images are captured at the correct focus distance. For some regions of interest, such as regions of interest that do not include a sample but instead include features such as annotations, focusing on non-sample features may be desirable because, for example, capturing a focused image of a feature such as an annotation may help an expert read the annotation and / or aid the optical character recognition process in transcribing text. For some regions of interest, both sample and non-sample features may be present. In this case, focusing on the sample is desirable, and this can be achieved by the processes described herein. In some embodiments, the processes described herein may provide a more efficient method of identifying a focus pattern than alternatives. For example, the processes described herein may require focusing on fewer points to determine a focus pattern.

[0057] Still referring to FIG. 1 , in some embodiments, a focus pattern, such as a plane, can be identified as a function of a first image and a first location. An (X,Y) location of a point can be determined from the first location. A best focus value can be determined from the first image. Together, these can be used to identify (X,Y,Z) coordinates that can be used to identify the focus pattern, as described herein.

[0058] Still referring to FIG. 1 , in some embodiments, a focus pattern, such as the Z level of best focus, may be used to scan the remainder of the row. For example, best focus may be used to scan the remaining row including the identified point. In some embodiments, the focus pattern may be used to scan additional rows. For example, a focus pattern determined as a function of the Z level of the point may be used to scan rows adjacent to the row including the point. In some embodiments, the focus pattern may be used to scan other rows within the same region of interest. In some embodiments, the focus pattern may be used to scan rows in other regions of interest, such as neighboring regions of interest.

[0059] Still referring to FIG. 1 , in some embodiments, identifying regions of interest, identifying points, identifying points, and / or determining focus patterns may be performed locally. For example, apparatus 100 may include a pre-trained machine learning model and apply the model to the image. In some embodiments, identifying regions of interest, identifying points, identifying points, and / or determining focus patterns may be performed externally. For example, apparatus 100 may transmit image data to another computing device and may receive outputs described herein. In some embodiments, identifying regions of interest, identifying lines, identifying points, and / or determining focus patterns may be performed in real time.

[0060] 1, in some embodiments, the device 100 can determine a scanning pattern. The scanning pattern may be based on, for example, the morphology of the sample. The scanning pattern may include, in non-limiting examples, a zigzag, a snake line, and a spiral.

[0061] Still referring to FIG. 1 , in some embodiments, a snake pattern may be used to scan a slide. The snake pattern may proceed in any horizontal direction. In a non-limiting example, the snake pattern may proceed across the length or width of the slide. In some embodiments, the direction in which the snake pattern proceeds may be selected to minimize the number of turns required. For example, the snake pattern may proceed along the shortest dimension of the slide. In another example, the shape of the sample may be identified, e.g., using machine vision, and the snake pattern may proceed in a direction corresponding to the shortest dimension of the sample. In some embodiments, a snake pattern may be selected when fast scanning is desired. In some embodiments, a snake pattern may minimize the camera and / or slide movement required to scan the slide. In some embodiments, a zigzag pattern may be used to scan a slide. Similar to what was described with respect to snake pattern scanning, the zigzag pattern scan may be performed in any horizontal direction, and the direction may be selected to minimize the number of turns and / or rows used to scan a feature, such as a slide and / or sample. In some embodiments, a spiral pattern may be used to scan a slide, feature, region of interest, etc. In some embodiments, a snake and / or zigzag pattern may be best suited for moving from one row to the next in the Z direction, while in some embodiments, a spiral pattern may be best suited for dynamic grid region of interest expansion.

[0062] Still referring to FIG. 1 , in some embodiments, the apparatus 100 can extrapolate the focal length of the second position as a function of the focus pattern. In some embodiments, the focus pattern can be determined as a function of one or more (X,Y,Z) points local to the region of interest and / or a subregion of the region of interest. In some embodiments, a focus pattern, such as a plane, can be used to approximate the optimal focus level using extrapolation (rather than interpolation). For example, the focus pattern can be used to approximate the optimal focus level at (X,Y) points outside the range of (X,Y) points already scanned, such as outside the range of X values, outside the range of Y values, or outside the range of shapes encompassing the already scanned (X,Y) points. In another example, the focus pattern can be used to approximate the optimal focus level, where the optimal focus level includes Z values ​​outside the range of Z values ​​used to determine the focus pattern. In another example, the (X,Y,Z) points can be extrapolated to a focus pattern spanning a row. In some embodiments, one or more additional (X,Y,Z) points can be used to update the focus pattern. In some embodiments, a focus pattern identified for one row may be extrapolated to another row, such as an adjacent row. In another example, a local focus pattern, such as a plane, may be extrapolated to identify an optimal focus level outside the local region. In another example, a focus pattern in a first region of interest may be extrapolated to generate a focus pattern in a second region of interest and / or to identify an optimal focus level at an (X,Y) point in the second region of interest.

[0063] Still referring to FIG. 1 , in some embodiments, the apparatus 100 can capture a second image of the slide at a second position with a focal length based on the focal pattern. For example, if the focal pattern is planar and a focused image of a specific (X,Y) point is desired, the apparatus 100 can capture the image using a focal length based on the Z coordinate of the plane at those (X,Y) coordinates. In some embodiments, the first position (e.g., the position where best focus is measured) and the second position can be set to image positions within the same region of interest. In some embodiments, an actuator mechanism can be mechanically connected to the movable element. The actuator mechanism can also move the movable element to the second position. In some embodiments, acquiring the second image can include capturing multiple images taken with focal lengths based on the focal pattern and constructing the second image from the multiple images. In some embodiments, the second image can include images taken as part of a Z-stack. Z-stacks are described below.

[0064] Still referring to FIG. 1 , in some embodiments, the device 100 may determine which regions of interest contain a sample. In some embodiments, this may be applied to a second image, such as a second image captured using a focal distance based on the focus pattern. In some embodiments, a sample identification machine learning model 160 may be used to determine which regions of interest contain a sample. In some embodiments, the sample identification machine learning model 160 may include a classifier. In some embodiments, the sample identification machine learning model 160 may be trained using supervised learning. The sample identification machine learning model 160 may be trained on a dataset including example images of slides and / or segments of images of slides that correlate to whether a sample is present. Such a dataset may be collected, for example, by capturing images of slides and manually identifying those that contain a sample. In some embodiments, multiple machine learning models may be trained to identify different types of samples. Once trained, the sample identification machine learning model 160 may accept an image of a region of interest as input and output a determination of whether a sample is present. In some embodiments, the sample identification machine learning model 160 may be improved, such as through the use of reinforcement learning. The feedback used to determine the cost function of the reinforcement learning model may include, for example, user input or annotations on the slide. For example, if annotations transcribed using optical character recognition indicate a particular type of sample and the sample identification machine learning model 160 indicates that the region of interest does not contain the sample, a cost function may be determined that indicates the output is false. In another example, a user may input a label to be associated with a region of interest. If the label indicates a particular type of sample and the output of the sample identification machine learning model 160 indicates that the sample is present in the region of interest, a cost function may be determined that indicates the output is correct.

[0065] Still referring to FIG. 1 , in some embodiments, whether a sample is present may be determined locally. For example, device 100 may include a pre-trained sample identification machine learning model 160 and apply the model to an image or a segment of an image. In some embodiments, whether a sample is present may be determined externally. For example, device 100 may transmit image data to another computing device and receive a determination of whether a sample is present. In some embodiments, whether a sample is present may be determined in real time.

[0066] 1, in some embodiments, a machine vision system and / or an optical character recognition system may be used to determine one or more characteristics of the sample and / or slide 116. In a non-limiting example, an optical character recognition system may be used to identify writing on the slide 116, which may be used to annotate an image of the slide 116.

[0067] Still referring to FIG. 1 , in some embodiments, the apparatus 100 can capture multiple images at different sample focal depths. As used herein, "sample focal depth" refers to the depth within a sample at which an optical system is focused. As used herein, "focal length" refers to the object-side focal length. In some embodiments, the first and second images may have different focal lengths and / or sample focal depths.

[0068] Still referring to FIG. 1 , in some embodiments, device 100 may include a machine vision system. In some embodiments, the machine vision system may include at least one camera. The machine vision system may use images, such as images from at least one camera, to make decisions about a scene, a space, and / or an object. For example, in some cases, the machine vision system may be used for world modeling or aligning objects in space. In some cases, alignment may include image processing, such as, but not limited to, object recognition, feature detection, edge / corner detection, etc. Non-limiting examples of feature detection include Scale Invariant Feature Transform (SIFT), Canny edge detection, Shi Tomasi corner detection, etc. In some cases, alignment may include one or more transformations that orient the camera frame (or image or video stream) relative to a three-dimensional coordinate system. Exemplary transformations include, but are not limited to, homography transformations and affine transformations. In one embodiment, the alignment of the first frame relative to the coordinate system may be verified and / or corrected using object identification and / or computer vision, as described above. For example, but not by way of limitation, an initial alignment into two dimensions, e.g., represented as an alignment to x and y coordinates, may be performed using a two-dimensional projection of a three-dimensional point onto the first frame. A third dimension of alignment, representing depth and / or the z-axis, may be detected by comparing two frames. For example, if the first frame includes a pair of frames captured using a pair of cameras (e.g., also referred to in this disclosure as a stereo camera), image recognition and / or edge detection software may be used to detect a pair of stereo views of the object's image. The two stereo views are compared to derive z-axis values ​​for points on the object, allowing for the derivation of additional z-axis points within and / or around the object using, for example, interpolation. This may be repeated for multiple objects in the field of view, including, but not limited to, environmental features of interest identified by an object classifier and / or indicated by an operator.In one embodiment, the x and y axes may be selected to span a plane common to the two cameras used to capture the stereoscopic image and / or the x-y plane of the first frame, so that the x and y translation components and φ may be pre-entered into a translation matrix and a rotation matrix for an affine transformation of the object's coordinates, as described above. As described above, estimation of the initial x and y coordinates and / or transformation matrix may alternatively or additionally be performed between the first and second frames. As described above, for each point of the object and / or its edges and / or a plurality of points on the object's edges, the x and y coordinates of the first stereoscopic frame may be entered, along with an initial estimate of the z coordinate, based on assumptions about the object, such as, for example, the assumption that the ground is approximately parallel to the x-y plane, as selected above. The Z coordinates and / or x, y, z coordinates aligned using the image capture and / or object identification process as described above may then be compared to coordinates predicted using an initial guess of the transformation matrix, and an error function may be calculated using comparing the two sets of points and the new x, y, and / or z coordinates, and may be iteratively estimated and compared until the error function falls below a threshold level. In some cases, the machine vision system may use a classifier, such as any of the classifiers described throughout this disclosure.

[0069] Still referring to FIG. 1 , in some embodiments, image data may be processed using optical character recognition. In some embodiments, optical character recognition or optical character reader (OCR) involves automatically converting an image of written (e.g., typed, handwritten, or printed) text into machine-coded text. In some cases, recognizing at least one keyword from the image data may include one or more processes, including, but not limited to, optical character recognition (OCR), optical word recognition, intelligent character recognition, intelligent word recognition, etc. In some cases, OCR may recognize written text one glyph or one character at a time. In some cases, optical word recognition can recognize written text one word at a time, for example, for languages ​​that use spaces as word separators. In some cases, intelligent character recognition (ICR) can recognize written text one glyph or one character at a time, for example, by employing machine learning processes. In some cases, intelligent word recognition (IWR) can recognize written text one word at a time, for example, by employing machine learning processes.

[0070] Still referring to FIG. 1, in some cases, OCR may be an "offline" process that analyzes static documents or image frames. In some cases, handwriting analysis may be used as input for handwriting recognition. For example, rather than simply using glyphs or word shapes, this technique can capture the order in which segments are drawn, the direction, pen placement and lift patterns, and other such movements. This additional information can result in more accurate handwriting recognition. In some cases, this technique is also referred to as "online," dynamic, real-time, or intelligent character recognition.

[0071] Still referring to FIG. 1 , in some cases, OCR processing may employ preprocessing of image data. Preprocessing processes may include, but are not limited to, deskew, despeckle, binarization, line removal, layout analysis or “zoning,” line and word detection, script recognition, character separation or “segmentation,” and normalization. In some cases, deskew processing may include applying a transformation (e.g., a homography or affine transformation) to the image data to align text. In some cases, despeckle processing may include removing positive and negative spots and / or smoothing edges. In some cases, binarization processing may include converting an image from color or grayscale to black and white (i.e., a binary image). Binarization may be performed as a simple method of separating text (or any other desired image components) from the background of the image data. In some cases, binarization may be required, for example, because the OCR algorithm being employed only supports binary images. In some cases, line removal processing may include removing glyphs and non-character images (e.g., boxes and lines). In some cases, a layout analysis or "zoning" process may identify columns, paragraphs, captions, etc. as separate blocks. In some cases, a line and word detection process may establish word and character shape metrics and separate words as needed. In some cases, a script recognition process may identify scripts, for example in multilingual documents, allowing an appropriate OCR algorithm to be selected. In some cases, a character separation or "segmentation" process may separate signal characters, for example in a character-based OCR algorithm. In some cases, a normalization process may normalize the aspect ratio and / or scale of the image data.

[0072] Still referring to FIG. 1 , in some embodiments, the OCR process may include an OCR algorithm. Exemplary OCR algorithms include a matrix matching process and / or a feature extraction process. Matrix matching may involve a pixel-by-pixel comparison of the image to stored glyphs. In some cases, matrix matching is also known as “pattern matching,” “pattern recognition,” and / or “image correlation.” Matrix matching may depend on the input glyph being properly separated from the rest of the image data. Matrix matching may also depend on the stored glyph being in a similar font and scale as the input glyph. Matrix matching may work best with typed text.

[0073] Still referring to FIG. 1 , in some embodiments, the OCR process can include a feature extraction process. In some cases, feature extraction can decompose a glyph into at least one feature. Exemplary, non-limiting features can include corners, edges, lines, closed loops, line direction, line intersections, etc. In some cases, feature extraction can reduce the dimensionality of the representation, making the recognition process more computationally efficient. In some cases, the extracted features are compared to an abstract, vector-like representation of the character and reduced to one or more glyph prototypes. Common techniques for feature detection in computer vision can be applied to this type of OCR. In some embodiments, a machine learning process such as a nearest neighbor classifier (e.g., a k-nearest neighbor algorithm) can be used to compare image features with stored glyph features and select the closest match. The OCR can employ any machine learning process described in this disclosure, such as the machine learning processes described with reference to FIGS. 2-4. Exemplary, non-limiting OCR software includes Cuneiform and Tesseract. Cuneiform is a multilingual, open-source optical character recognition system originally developed by Cognitive Technologies of Moscow, Russia. Tesseract is free OCR software originally developed by Hewlett-Packard of Palo Alto, California, USA.

[0074] Still referring to FIG. 1 , in some cases, OCR may employ a two-pass approach to character recognition. The first pass may attempt to recognize characters. Each successful character is passed through an adaptive classifier as training data. The adaptive classifier further analyzes the image data to get a chance to recognize the character more accurately. Because the adaptive classifier may have learned something useful in the first pass that was too late to recognize the character, a second pass is performed on the image data. The second pass may include adaptive recognition, where characters recognized with high confidence in the first pass can be used to better recognize the remaining characters in the second pass. In some cases, a two-pass approach may be advantageous for specialized fonts or low-quality image data. Another exemplary OCR software tool is OCRopus. Development of OCRopus is led by the German Research Center for Artificial Intelligence (DFKI) in Kaiserslautern, Germany. In some cases, OCR software may employ neural networks.

[0075] Still referring to Figure 1, in some cases, OCR can include post-processing. For example, OCR accuracy can be improved if the output is constrained by a lexicon. The lexicon may include a list or set of words that are allowed to appear in a document. In some cases, the lexicon may include, for example, all words in the English language, or a more specialized lexicon for a particular domain. In some cases, the output stream may be a plain text stream or a file of characters. In some cases, the OCR process can preserve the original layout of the image data. In some cases, near-neighbor analysis can utilize co-occurrence frequency to correct errors by noting that certain words frequently appear together. For example, "Washington, DC" is typically much more common in English than "Washington DOC." In some cases, the OCR process can utilize a priori knowledge of the grammar of the language being recognized. For example, grammatical rules may be used to determine whether a word is a verb or a noun. Distance conceptualizations can be employed for recognition and classification. For example, the Levenshtein distance algorithm can be used in post-OCR processing to further optimize results.

[0076] 1, in some embodiments, the device 100 can remove artifacts from an image. As used herein, an "artifact" is a visual imperfection, an element of an image that distracts from an element of interest, an element of an image that obscures an element of interest, or any other undesirable element of an image.

[0077] Still referring to FIG. 1 , device 100 may include an image processing module. As used in this disclosure, an “image processing module” is a component designed to process digital images. In one embodiment, the image processing module may include software algorithms that can analyze, manipulate, or otherwise improve an image, such as, but not limited to, image processing techniques described below. In another embodiment, the image processing module may include hardware components, such as, but not limited to, one or more graphics processing units (GPUs) that can accelerate the processing of large numbers of images. In some cases, the image processing module may be implemented using one or more image processing libraries, such as, but not limited to, OpenCV, PIL / Pillow, ImageMagick, etc.

[0078] 1 , the image processing module may be configured to receive images from the optical sensor 120. One or more images may be transmitted from the optical sensor 120 to the image processing module via any suitable electronic communication protocol, including, but not limited to, packet-based protocols such as Transmission Control Protocol / Internet Protocol (TCP-IP), File Transfer Protocol (FTP), etc. Receiving the images may include retrieving the images from a data store containing the images, as described below. For example, but not by way of limitation, the images may be retrieved using a query that specifies a timestamp that the images are required to match.

[0079] Still referring to FIG. 1 , the image processing module may be configured to process images. In one embodiment, the image processing module may be configured to compress and / or encode the images to reduce file size and storage requirements while maintaining essential visual information necessary for further processing steps as described below. In one embodiment, compressing and / or encoding the images may facilitate high-speed transmission of the images. In some cases, the image processing module may be configured to perform lossless compression on the images, which can maintain the original image quality of the image. In non-limiting examples, the image processing module may use compression techniques such as, but not limited to, Huffman coding, Lempel-Ziv-Welch (LZW), Run-Length Compression (RLC), and the like. One or more lossless compression algorithms, such as Regression Level Encoding (RLE), may be utilized to identify and remove image redundancy without losing information. In such embodiments, compressing and / or encoding each image may include converting the file format of each image to PNG, GIF, lossless JPEG2000, etc. In one embodiment, an image compressed via lossless compression may be fully reconstructed to the image's original form (e.g., the original image's resolution, dimensions, color representation, format, etc.). In other cases, the image processing module may be configured to perform lossy compression on the image, which may sacrifice some image quality to achieve a higher compression ratio. In a non-limiting example, the image processing module may utilize one or more lossy compression algorithms, such as, but not limited to, the Discrete Cosine Transform (DCT) for JPEG or the Wavelet Transform for JPEG2000, to discard less important information in the image, resulting in a smaller file size with only a small loss in image quality. In such embodiments, compressing and / or encoding the images may include converting the file format of each image to JPEG, WebP, lossy JPEG2000, etc.

[0080] Still referring to FIG. 1 , in one embodiment, processing the image can include determining a degree of representation quality of the region of interest in the image. In one embodiment, the image processing module can determine the blurriness of the image. In a non-limiting example, the image processing module may perform blur detection by taking a Fourier transform, or an approximation such as a fast Fourier transform (FFT), of the image and analyzing the distribution of low and high frequencies in the resulting frequency domain representation of the image. For example, but not by way of limitation, the number of high frequency values ​​below a threshold level may indicate blur. In another non-limiting example, blur detection may be performed by convolving the image, a channel of the image, or the like, with a Laplacian kernel, which may generate a numerical score reflecting the number of abrupt changes in intensity present in each image, such that a high score indicates sharpness and a low score indicates blur. In some cases, blur detection can be performed using a gradient-based operator, which measures the operator based on the gradient or first derivative of the image, based on the hypothesis that abrupt changes indicate sharp edges in the image and therefore a low degree of blur. In some cases, blur detection may be performed using a wavelet-based operator that exploits the ability of discrete wavelet transform coefficients to describe the frequency and spatial content of an image. In some cases, blur detection may be performed using a statistics-based operator that utilizes some image statistics as texture descriptors to calculate the focus level. In other cases, blur detection may be performed using discrete cosine transform (DCT) coefficients to calculate the focus level of an image from its frequency content. Additionally or alternatively, the image processing module may be configured to rank images according to a degree of quality of depiction of the region of interest and select the highest-ranked image from the plurality of images.

[0081] Still referring to FIG. 1 , processing the image may include enhancing the image or at least one region of interest through a number of image processing techniques to improve the quality of the image (or a degree of quality of depiction) for better processing and analysis, as described further in this disclosure. In one embodiment, the image processing module may be configured to perform a noise reduction operation on the image, which may remove or minimize noise (caused by various causes, such as sensor limitations, poor lighting conditions, image compression, etc.), resulting in a cleaner, more visually consistent image. In some cases, the noise reduction operation may be performed using one or more image filters; for example, but not limited to, noise reduction operations may include Gaussian filtering, median filtering, bilateral filtering, etc. The noise reduction process may be performed by the image processing module by averaging or filtering neighboring pixel values ​​of each pixel in the image to reduce random fluctuations.

[0082] Still referring to FIG. 1 , in another embodiment, the image processing module may be configured to perform a contrast enhancement operation on the image. In some cases, an image may exhibit low contrast, e.g., features may be difficult to distinguish from the background. A contrast enhancement operation may improve the contrast of the image by expanding the intensity range of the image and / or redistributing intensity values ​​(i.e., the degree of brightness of pixels in the image). In a non-limiting example, the intensity values ​​represent the gray level or color of each pixel and may scale from an intensity range of 0 to 255 for an 8-bit image and from 0 to 16,777,215 for a 24-bit color image. In some cases, the contrast enhancement operation may include, but is not limited to, histogram equalization, adaptive histogram equalization (CLAHE), contrast stretching, etc. The image processing module may be configured to adjust the brightness and darkness levels in the image to make features more distinguishable (i.e., to enhance the degree of rendering quality). Additionally or alternatively, the image processing module may be configured to perform a brightness normalization operation to correct for variations in lighting conditions (i.e., uneven brightness levels). In some cases, the image may include consistent brightness levels across an entire region after the brightness normalization operation performed by the image processing module. In a non-limiting example, the image processing module may perform global normalization or local mean normalization, where a mean intensity value for the entire image or a region of the image is calculated and used to adjust the brightness level.

[0083] Still referring to FIG. 1 , in other embodiments, the image processing module may be configured to perform color space conversion operations to enhance the degree of rendering quality. In a non-limiting example, in the case of a color image (i.e., an RGB image), the image processing module may be configured to convert the RGB image to grayscale or HSV color space. Such conversion may enhance the difference in intensity values ​​between regions or features of interest and the background. The image processing module may further be configured to perform image sharpening operations, such as, but not limited to, unsharp masking, Laplacian sharpening, and high-pass filtering. The image processing module may use image sharpening operations to enhance edges and fine details associated with regions or features of interest in the image by emphasizing high-frequency components in the image.

[0084] Still referring to FIG. 1 , processing the image may include isolating a region or feature of interest from the remainder of the image as a function of multiple image processing techniques. The image may include the highest-ranked image selected by the image processing module as described above. In one embodiment, the multiple image processing techniques may include one or more morphological operations, which are techniques developed based on set theory, lattice theory, topology, and random functions used to manipulate geometric structures using structuring elements. For purposes of this disclosure, a “structuring element” is a small matrix or kernel that defines the shape and size of a morphological operation. In some cases, a structuring element may be placed at the center of each pixel of the image and used to determine the output pixel value at that location. In a non-limiting example, isolating a region or feature of interest from the image may include applying a dilation operation, which is a basic morphological operation configured to expand or grow the boundaries of objects (e.g., cells, dust particles, etc.) in the image. In another non-limiting example, isolating a region or feature of interest from the image may include applying an erosion operation, which is a basic morphological operation configured to shrink or contract the boundaries of objects in the image. In another non-limiting example, isolating a region of interest or feature from an image may include applying an opening operation, which is a basic morphological operation configured to remove small objects or thin structures from an image while preserving larger structures. In a further non-limiting example, isolating a region of interest or feature from an image may include applying a closing operation, which is a basic morphological operation configured to fill small gaps or holes in objects in an image while preserving the overall shape and size of the objects. These morphological operations can be performed by an image processing module to enhance object edges, remove noise, or fill gaps in regions of interest or features before further processing.

[0085] Still referring to FIG. 1 , in one embodiment, isolating a region of interest or feature from an image can include utilizing an edge detection technique that can detect one or more shapes defined by edges. As used in this disclosure, “edge detection technique” includes mathematical methods that identify points in a digital image where the image brightness changes abruptly and / or discontinuities occur. In one embodiment, such points can be organized into straight and / or curved line segments called “edges.” Edge detection techniques can be performed by the image processing module using any suitable edge detection algorithm, including, but not limited to, Canny edge detection, Sobel operator edge detection, Prewitt operator edge detection, Laplacian operator edge detection, and / or differential edge detection. Edge detection techniques may include edge detection based on phase congruency, which finds all locations in an image where all sinusoids in the frequency domain, generated using, for example, Fourier decomposition, have a consistent phase, which can indicate the location of an edge. Edge detection techniques can be used to detect the shape of features of interest, such as cells, which indicate cell membranes or cell walls. In one embodiment, edge detection techniques can be used to find closed shapes formed by edges.

[0086] Still referring to FIG. 1 , in a non-limiting example, separating the feature of interest from the image can include determining the feature of interest using edge detection techniques. The feature of interest can include a specific region within the digital image that contains information relevant to further processing, as described below. In a non-limiting example, image data located outside the feature of interest can contain irrelevant or redundant information. Portions of the image containing irrelevant or redundant information can be ignored by the image processing module, thereby allowing resources to be focused on the feature of interest. In some cases, the feature of interest can vary in size, shape, and / or location within the image. In a non-limiting example, the feature of interest can be depicted as a circle surrounding a cell nucleus. In some cases, the feature of interest can specify one or more coordinates, distances, etc., such as the center and radius of the circle surrounding the cell nucleus in the image. The image processing module can then be configured to separate the feature of interest from the image based on the feature of interest. In a non-limiting example, the image processing module can crop the image according to a bounding box surrounding the feature of interest.

[0087] Still referring to FIG. 1 , the image processing module may be configured to perform connected component analysis (CCA) on the image for feature-of-interest isolation. As used in this disclosure, “connected component analysis (CCA),” also known as connected component labeling, is an image processing technique used to identify and label connected regions in a binary image (i.e., an image in which each pixel has only two possible values: 0 or 1, black or white, or foreground or background). As described herein, a “connected region” is a group of adjacent pixels that share the same value and are connected based on a predefined neighborhood system, such as, but not limited to, a 4-connected neighborhood or an 8-connected neighborhood. In some cases, the image processing module may convert the image to a binary image by thresholding, which may include setting a threshold that separates pixels of the image corresponding to the feature of interest (foreground) from pixels corresponding to the background. Pixels with intensity values ​​above the threshold may be set to 1 (white), and pixels below the threshold may be set to 0 (black). In one embodiment, CCA may be employed to detect and extract features of interest by identifying multiple connected regions that exhibit specific characteristics or features of the feature of interest. The image processing module can then filter the plurality of connected regions by analyzing characteristics of the plurality of connected regions, such as, but not limited to, area, aspect ratio, height, width, perimeter, etc. In a non-limiting example, the image processing module may retain connected components that closely resemble the dimensions and aspect ratio of the feature of interest as the feature of interest, while discarding other components. The image processing module may further be configured to extract the feature of interest from the image for further processing, as described below. Still referring to FIG. 1 , in one embodiment, isolating the feature of interest from the image may include dividing a region depicting the feature of interest into a plurality of sub-regions.

[0088] Dividing a region into subregions can include dividing the region as a function of the feature of interest and / or the CCA via an image segmentation process. As used in this disclosure, "image segmentation" refers to a process that divides a digital image into one or more segments, each representing a different portion of the image. The image segmentation process can change the representation of the image. The image segmentation process can be performed by an image processing module. In a non-limiting example, the image processing module can perform region-based segmentation, which includes growing regions from one or more seed points or pixels on the image based on similarity criteria. Similarity criteria may include, but are not limited to, color, intensity, texture, etc. In a non-limiting example, the region-based segmentation can include region growing, region merging, watershed algorithms, etc.

[0089] 1, in some embodiments, the device 100 can remove artifacts identified by the machine vision system or optical character recognition system described above. Non-limiting examples of artifacts that may be removed include dust particles, air bubbles, cracks in the slide 116, writing on the slide 116, shadows, visual noise such as grainy images, etc. In some embodiments, the artifacts may be partially removed and / or may be less visible.

[0090] Still referring to FIG. 1 , in some embodiments, artifacts can be removed using an artifact-removal machine learning model. In some embodiments, the artifact-removal machine learning model can be trained on a dataset including images associated with artifact-free images. In some embodiments, the artifact-removal machine learning model can accept as input an image including artifacts and output an image free of artifacts. For example, the artifact-removal machine learning model can accept as input an image including air bubbles in a slide and output an image free of the air bubbles. In some embodiments, the artifact-removal machine learning model can include a generative machine learning model, such as a diffusion model. A diffusion model can learn the structure of a dataset by modeling how data points diffuse through a latent space. In some embodiments, artifact removal can be performed locally. For example, the device 100 can include a pre-trained artifact-removal machine learning model and apply the model to the images. In some embodiments, artifact removal can be performed externally. For example, the device 100 can transmit image data to another computing device and receive the artifact-removed image. In some embodiments, artifacts can be removed in real time. In some embodiments, artifacts can be removed based on user identification. For example, a user can use a mouse cursor to drag a box around the artifacts, and the device 100 can remove the artifacts within the box.

[0091] 1 , in some embodiments, device 100 may display an image to a user. In some embodiments, the first image may be displayed to the user in real time. In some embodiments, the image may be displayed to the user using output interface 132. For example, the first image may be displayed on a display such as a screen. In some embodiments, the first image may be displayed to the user in the context of a graphical user interface (GUI). For example, the GUI may include controls for navigating the image, such as zoom controls and controls for changing the displayed location. The GUI may include a touch screen.

[0092] Still referring to FIG. 1 , in some embodiments, the device 100 can receive a parameter set from a user. As used herein, a “parameter set” is a set of values ​​that identifies how to capture an image. The parameter set may be implemented as a data structure, as described below. In some embodiments, the device 100 can receive the parameter set from a user using the input interface 128. The parameter set can include X and Y coordinates that indicate where the user wants to look. The parameter set can include a desired magnification level. As used herein, a “magnification level” is a data item that indicates how much to zoom in or out to capture an image. The magnification level can account for optical zoom and / or digital zoom. As a non-limiting example, the magnification level can be “8x.” The parameter set may also include a desired depth of focus and / or focal length of the sample. In a non-limiting example, the user can manipulate the input interface 128 so that the parameter set includes X and Y coordinates and a magnification level corresponding to a more magnified view of a particular region of the sample. In some embodiments, the parameter set corresponds to a more magnified view of a particular region of the sample included in the first image. This may be done, for example, to obtain a more detailed view of a small object. As used herein, unless otherwise specified, "X coordinate" and "Y coordinate" refer to coordinates along the vertical axis, and the plane defined by these axes is parallel to the plane of the slide 116 surface. In some cases, setting the magnification may involve changing one or more optical elements within the optical system. For example, setting the magnification may involve replacing a first objective lens with a second objective lens having a different magnification. Furthermore, replacing one or more optical components "down-beaming" from the objective lens can change the overall magnification of the optical system and set the magnification. In some cases, setting the magnification may involve changing the digital magnification. Digital magnification may involve outputting an image at a different resolution, i.e., after rescaling the image, using an output interface.In some embodiments, device 100 can capture images with a particular field of view, which can be determined, for example, based on the level of magnification or how wide the camera captures the image.

[0093] Still referring to FIG. 1 , in some embodiments, the apparatus 100 can move one or more of the slide port 140, the slide 116, and the at least one optical system 120 to a second position. In some embodiments, the location of the second position can be based on a parameter set. In some embodiments, the location of the second position can be based on identifying a region of interest, a line, or a point, as described above. The second position can be determined, for example, by changing the position of the optical system 120 relative to the slide 116 based on the parameter set. For example, the parameter set can indicate that the second position is achieved by changing the X coordinate by 5 mm in a particular direction. In this example, the second position can be determined by changing the original position of the optical system by 5 mm in that direction. In some embodiments, such movement can be performed using the actuator mechanism 124. In some embodiments, the actuator mechanism 124 can move the slide port 140 such that the slide 116 is in a position relative to the at least one optical system 120 that allows the optical sensor 120 to capture images as directed by the parameter set. For example, the slide 116 can rest on the slide port 140, and movement of the slide port 140 can move the slide 116. In some embodiments, the actuator mechanism 124 can move the slide 116 so that the slide 116 is in a position relative to the at least one optical system 120 such that the optical sensor 120 can capture images as directed by the set of parameters. For example, the slide 116 can be connected to the actuator mechanism 124 such that the actuator mechanism 124 can move the slide 116 relative to the at least one optical system 120. In some embodiments, the actuator mechanism 124 can move the at least one optical system 120 so that the slide 116 is in a position relative to the slide 116 such that the optical sensor 120 can capture images as directed by the set of parameters.For example, the slide 116 may be stationary, and the actuator mechanism 124 may move the at least one optical system 120 to a position relative to the slide 116. In some embodiments, the actuator mechanism 124 may move more than one of the slide port 140, the slide 116, and the at least one optical system 120 so that they are in the correct relative positions. In some embodiments, the actuator mechanism 124 may move the slide port 140, the slide 116, and / or the at least one optical system 120 in real time. For example, user input of a set of parameters may result in substantially instantaneous movement of an item by the actuator mechanism 124.

[0094] Still referring to FIG. 1 , in some embodiments, the device 100 may capture a second image of the slide 116 at a second position. In some embodiments, the device 100 may capture the second image using at least one optical system 120. In some embodiments, the second image may include an image of an area of ​​the sample. In some embodiments, the second image may include an image of the area captured in the first image. For example, the second image may include a magnified image of the area in the first image, with a higher resolution per unit area. In another example, the second image may be captured using a focal length based on a focus pattern. This allows the second image to be displayed, allowing the user to detect smaller details within the imaged area.

[0095] 1, in some embodiments, the second image includes an X and Y coordinate shift relative to the first image. For example, the second image may partially overlap the first image.

[0096] Still referring to FIG. 1 , in some embodiments, device 100 may capture the second image in real time. For example, a user may manipulate input interface 128 to create a parameter set, actuator mechanism 124 may begin moving slide 116 relative to optical system 120 substantially immediately after input interface 128 is manipulated, and optical system 120 may capture the second image substantially immediately after actuator mechanism 124 completes its movement. In some embodiments, artifacts may also be removed in real time. In some embodiments, images may also be annotated in real time. In some embodiments, a focus pattern may also be determined in real time, and images captured after the first image may be captured using a focal length according to the focus pattern.

[0097] Still referring to FIG. 1 , in some embodiments, the device 100 can display a second image to a user. In some embodiments, the second image can be displayed to the user using the output device 132. The second image can be displayed as described above with respect to the output device and displaying the first image. In some embodiments, displaying the second image to the user can include replacing an area of ​​the first image with the second image to generate a hybrid image and displaying the hybrid image to the user. As used herein, a “hybrid image” is an image constructed by combining the first image and the second image. In some embodiments, generating the hybrid image in this manner can preserve the second image. For example, if the second image has a higher resolution per unit area and covers a smaller area than the first image, the second image can replace a lower resolution per unit area segment of the first area corresponding to the area covered by the second image. In some embodiments, image adjustments can be made at the boundary between the first and second images in the hybrid image to offset visual differences between the first and second images. In a non-limiting example, colors may be adjusted so that the background color of an image is consistent along the image's boundaries. As another non-limiting example, image brightness may be adjusted so that there are no noticeable differences between the brightness of the images. In some embodiments, artifacts may be removed from the second image and / or the hybrid image, as described above. In some embodiments, the second image may be displayed to the user in real time. For example, adjustments (such as annotations and / or artifact removal) may begin substantially immediately after the second image is captured, and an adjusted version of the second image may be displayed to the user substantially immediately after the adjustments are made. In some embodiments, an unadjusted version of the second image may be displayed to the user while the adjustments are made. In some embodiments, if there are multiple images covering a particular area, a lower-resolution image of the area may be displayed when the user zooms out using user interface 136.

[0098] 1 , in some embodiments, apparatus 100 may transmit a data structure including the first image, the second image, the hybrid image, and / or the plurality of images to an external device. Such external devices may include, in non-limiting examples, a phone, a tablet, or a computer. In some embodiments, such transmission may configure the external device to display the images.

[0099] Still referring to FIG. 1 , in some embodiments, device 100 can annotate an image. In some embodiments, device 100 can annotate a first image. In some embodiments, device 100 can annotate a second image. In some embodiments, device 100 can annotate a hybrid image. For example, during creation of a hybrid image, device 100 can recognize cells depicted in the hybrid image as a particular type of cell and annotate the hybrid image indicating the cell type. In a non-limiting example, device 100 can associate text with a particular location within the image, where the text describes a feature present at that location within the image. In some embodiments, device 100 can annotate an image selected from a list consisting of the first image, the second image, and the hybrid image.

[0100] Still referring to FIG. 1 , in some embodiments, annotation may be performed as a function of user input of annotation instructions into input interface 128. As used herein, “annotation instructions” refers to data generated based on user input indicating whether to create an annotation or describing an annotation to be created. For example, a user may select an option that controls whether device 100 annotates an image. In another example, a user may manually annotate an image. In some embodiments, annotation may be performed automatically. In some embodiments, annotation may be performed using an annotation machine learning model. In some embodiments, the annotation machine learning model may include an optical character recognition model, as described above. In some embodiments, the annotation machine learning model may be trained using a dataset including image data associated with text depicted in the image. In some embodiments, the annotation machine learning model may accept input image data and output annotated image data and / or annotations to apply to the image data. In some embodiments, the annotation machine learning model may be used to convert text written on slide 116 into annotations on the image. In some embodiments, the annotation machine learning model may include a machine vision model, as described above. In some embodiments, an annotation machine learning model, including a machine vision model, may be trained on a dataset including image data associated with annotations that indicate features of the image data. In some embodiments, the annotation machine learning model may accept input image data and may output annotated image data and / or annotations to apply to the image data. Non-limiting examples of features that the annotation machine learning model may be trained to recognize include cell types, cell features, and objects within the slide 116, such as air bubbles. In some embodiments, images may be annotated in real time. For example, annotation may begin immediately after the image is captured and / or immediately after a command to annotate the image is received, and the annotated image may be displayed to the user immediately after annotation is completed.

[0101] 1 , in some embodiments, device 100 may determine a visual element data structure. In some embodiments, device 100 may display visual elements to a user as a function of the visual element data structure. As used herein, a "visual element data structure" is a data structure that describes a visual element. As non-limiting examples, the visual element may include a first image, a second image, a hybrid image, and an element of a GUI.

[0102] Still referring to FIG. 1 , in some embodiments, the visual element data structure may include visual elements. As used herein, “visual element” is data that is visually displayed to a user. In some embodiments, the visual element data structure may include rules for displaying the visual elements. In some embodiments, the visual element data structure may be determined as a function of the first image, the second image, and / or the hybrid image. In some embodiments, the visual element data structure may be determined as a function of items from a list consisting of the first image, the second image, the hybrid image, a GUI element, and an annotation. In a non-limiting example, the visual element data structure may be generated such that visual elements describing features of the first image, such as annotations, are displayed to a user.

[0103] 1, in some embodiments, the visual element may include one or more elements such as text, images, shapes, charts, particle effects, interactable features, etc. As a non-limiting example, the visual element may include a touchscreen button that sets a magnification level.

[0104] 1, the visual element data structure may include rules that govern whether or when a visual element is displayed. In a non-limiting example, the visual element data structure may include a rule that causes a visual element including annotations describing the first image, the second image, and / or the hybrid image to be displayed when a user selects a particular region of the first image, the second image, and / or the hybrid image using the GUI.

[0105] 1, the visual element data structure can include multiple visual elements or rules for presenting multiple visual elements at one time. In one embodiment, approximately 1, 2, 3, 4, 5, 10, 20, or 50 visual elements are displayed simultaneously. For example, multiple annotations can be displayed simultaneously.

[0106] Still referring to FIG. 1 , in some embodiments, device 100 can transmit visual elements to a display, such as output interface 132. The display can convey the visual elements to a user. The display can include, for example, a smartphone screen, a computer screen, a tablet screen, etc. The display can be configured to provide a visual interface. The visual interface can include one or more virtual interactive elements, such as, but not limited to, buttons, menus, etc. The display can include one or more physical interactive elements, such as buttons, a computer mouse, or a touch screen, that allow a user to input data into the display. The interactive elements can be configured to enable interaction between a user and a computing device. In some embodiments, the visual element data structure is determined as a function of data input by a user into the display.

[0107] Still referring to FIG. 1 , variables and / or data described herein may be represented as data structures. In some embodiments, a data structure may include one or more functions and / or variables, such as a class in object-oriented programming. In some embodiments, a data structure may include data in the form of a Boolean, integer, floating point, string, date, etc. In a non-limiting example, an annotation data structure may include string values ​​representing the text of the annotation. In some embodiments, the data of the data structure may be organized as a linked list, tree, array, matrix, tensor, etc. In a non-limiting example, an annotation data structure may be organized in an array. In some embodiments, a data structure may include one or more elements of metadata or may be associated with one or more elements of metadata. A data structure may include one or more self-referencing data elements that the processor 104 can use in interpreting the data structure. In a non-limiting example, a data structure may include a "*" tag indicating that the content between the tags is a date. <date> "and"< / date> " tag.

[0108] Still referring to FIG. 1 , the data structure may be stored, for example, in memory 108 or a database. The database may be implemented as, without limitation, a relational database, a key-value database such as a NOSQL database, or any other format or structure for use as a database that one of ordinary skill in the art would recognize as appropriate upon reviewing this disclosure in its entirety. The database may alternatively or additionally be implemented using a distributed data storage protocol and / or data structure, such as a distributed hash table. The database may include multiple data entries and / or records, as described above. Data entries in the database may be flagged or linked to one or more additional information elements, which may be reflected in data entry cells and / or in linked tables, such as tables related by one or more indexes in a relational database. One of ordinary skill in the art, upon reviewing this disclosure in its entirety, will recognize various ways in which data entries in a database may store, search, organize, and / or reflect the data and / or records, and categories and / or populations of data, used herein, consistent with this disclosure.

[0109] 1, in some embodiments, the data structure may be read and / or manipulated by the processor 104. In a non-limiting example, the image data structure may be read and displayed to a user. In another non-limiting example, the image data structure may be modified to remove artifacts, as described above.

[0110] Still referring to FIG. 1 , in some embodiments, the data structure may be calibrated. In some embodiments, the data structure may be trained using a machine learning algorithm. In a non-limiting example, the data structure may include an array of data representing biases in the connections of a neural network. In this example, the neural network may be trained on a set of training data, and a backpropagation algorithm may be used to correct the data in the array. Machine learning models and neural networks are further described herein.

[0111] One or more features of the device 100 may correspond to features disclosed in one or more of the following: (A) U.S. Patent Application No. 18 / 217,378, filed July 25, 2023, entitled "APPARATUS AND A METHOD FOR DETECTING ASSOCIATIONS AMONG DATASETS OF DIFFERENT TYPES," which is incorporated herein by reference in its entirety; (B) U.S. Patent Application No. 18 / 226,017, filed July 25, 2023, entitled "APPARATUS AND A METHOD FOR GENERATING A CONFIDENCE SCORE ASSOCIATED WITH A SCANNED LABEL," which is incorporated herein by reference in its entirety; or (C) U.S. Patent Application No. 18 / 226,058, filed July 25, 2023, entitled "IMAGING DEVICE AND A METHOD FOR IMAGE GENERATION OF A SPECIMEN" (incorporated herein by reference in its entirety). (D) U.S. Patent Application No. 18 / 226,100, filed July 25, 2023, entitled "APPARATUS AND METHODS FOR REAL-TIME IMAGE GENERATION" (incorporated herein by reference in its entirety).

[0112] 2, an exemplary embodiment of a machine learning module 200 that may perform one or more machine learning processes as described in this disclosure is illustrated. The machine learning module may use the machine learning processes to perform the determining, classifying, and / or analyzing steps, methods, processes, etc. as described in this disclosure. As used in this disclosure, a "machine learning process" is a process that automatically uses training data 204 to generate algorithms instantiated in hardware or software logic, data structures, and / or functions executed by a computing device / module to generate output 208 when data is provided as input 212, as opposed to a non-machine learning software program in which the commands to be executed are predetermined by a user and written in a programming language.

[0113] Still referring to FIG. 2 , as used herein, “training data” refers to data containing correlations that a machine learning process can use to model relationships between two or more categories of data elements. For example, without limitation, training data 204 may include multiple data entries, also known as “training examples,” each representing a set of data elements recorded, received, and / or generated together, where the data elements may be correlated by shared presence in a given data entry, proximity in a given data entry, etc. The multiple data entries of training data 204 may exhibit one or more trends in correlations between categories of data elements. For example, without limitation, high values ​​of a first data element belonging to a first category of data elements tend to correlate with high values ​​of a second data element belonging to a second category of data elements, indicating the possibility of a proportional or other mathematical relationship linking values ​​belonging to the two categories. Multiple categories of data elements may be associated in training data 204 according to various correlations. Correlations may indicate causal and / or predictive links between categories of data elements and may be modeled as relationships, such as mathematical relationships, by machine learning processes, as described in further detail below. Training data 204 may be formatted and / or organized by categories of data elements, for example, by associating the data elements with one or more descriptors that correspond to the categories of the data elements. As a non-limiting example, training data 204 may include data entered into standardized forms by a person or process, such that entry of a given data element in a given field of the form maps to one or more descriptors of the category. Elements of training data 204 may be linked to descriptors of the category by tags, tokens, or other data elements.For example, but not by way of limitation, the training data 204 may be provided in a fixed-length format, a format that links the location of the data to a category, such as a comma-separated values ​​(CSV) format, and / or a self-describing format, such as extensible markup language (XML) or JavaScript Object Notation (JSON), that allows a process or device to detect the category of the data.

[0114] Alternatively or additionally, and continuing to refer to FIG. 2 , the training data 204 may include one or more uncategorized elements. That is, the training data 204 may be unformatted or may not include descriptors for some elements of the data. Machine learning algorithms and / or other processes may sort the training data 204 according to one or more categorizations, for example, using natural language processing algorithms, tokenization, detecting correlation values ​​in the raw data, etc. Categories may be generated using correlation and / or other processing algorithms. As a non-limiting example, in a corpus of text, phrases consisting of “n” compounds, such as nouns modified by other nouns, may be identified according to the statistically significant prevalence of n-grams containing such words in a particular order. Such n-grams may be categorized as linguistic elements, such as “words,” that are tracked similarly to single words, generating new categories as a result of the statistical analysis. Similarly, in a data entry containing text data, a person's name may be identified by reference to a list, dictionary, or other glossary, allowing for ad-hoc categorization by a machine learning algorithm and / or automatic association of the data in the data entry with a descriptor or a given format. The ability to automatically categorize data entries allows the same training data 204 to be applied to two or more different machine learning algorithms, as described in further detail below. The training data 204 used by the machine learning module 200 may correlate any input data, as described in this disclosure, with any output data, as described in this disclosure. As a non-limiting example, the input may include an image of a region of interest, and the output may include a determination of whether a sample is present.

[0115] 2 , the training data may be filtered, sorted, and / or selected using one or more supervised and / or unsupervised machine learning processes and / or models, as described in further detail below, including, but not limited to, the training data classifier 216. The training data classifier 216 may include a “classifier,” which as used in this disclosure is a machine learning model, as defined below, and is a data structure that represents and / or uses a mathematical model, neural net, or program generated by a machine learning algorithm known as a “classification algorithm,” as described in further detail below, for example, that sorts inputs into categories or bins of data and outputs the categories or bins of data and / or their associated labels. The classifier may be configured to output at least one datum that labels or otherwise identifies sets of data, such as those clustered and found to be close under a distance metric, as described below. The distance metric may include any norm, such as, but not limited to, the Pythagorean norm. The machine learning module 200 may generate a classifier using a classification algorithm, which is defined as a process in which a computing device and / or any module and / or component operating on the computing device derives a classifier from the training data 204. The classification may be performed using, but is not limited to, a linear classifier such as, but not limited to, a logistic regression and / or a naive Bayes classifier, a nearest neighbor classifier such as a k-nearest neighbor classifier, a support vector machine, a least-squares support vector machine, Fisher's linear discriminant, a quadratic classifier, a decision tree, a boosted tree, a random forest classifier, learning vector quantization, and / or a neural network-based classifier. As a non-limiting example, the training data classifier 216 may classify elements of the training data as either present or absent in the sample.

[0116] With further reference to FIG. 2 , training examples for use as training data may be selected from a population of potential examples according to a cohort relevant to the analytical problem, classification task, etc. to be solved. Alternatively or additionally, training data may be selected to span a set of situations or inputs that the machine learning model and / or process is likely to encounter upon deployment. For example, for each category of input data to a machine learning process or model that may exist across a range of values ​​in a population of phenomena, such as, but not limited to, images, user data, process data, physical data, etc., the computing device, processor, and / or machine learning model may select training examples representing each possible value in such range and / or a representative sample of values ​​in such range. Selecting a representative sample may include selecting training examples in a proportion that matches a statistically determined and / or predicted distribution of such values ​​according to relative frequency, e.g., such that more frequently encountered values ​​in the population of analyzed data are represented by more training examples than less frequently encountered values. Alternatively or additionally, the set of training examples may be compared to a collection of representative values ​​in a database and / or presented to a user, allowing the process to detect, automatically or via user input, one or more values ​​not included in the set of training examples. A computing device, processor, and / or module may automatically generate missing training examples by receiving and / or acquiring missing input and / or output values ​​and associating the missing input and / or output values ​​with corresponding output and / or input values ​​that coexist in a data record with the acquired values, such as provided by a user and / or other device.

[0117] Still referring to FIG. 2 , the computer, processor, and / or module may be configured to sanitize the training data. As used in this disclosure, "sanitizing" training data refers to a process in which training examples that prevent a machine learning model and / or process from converging to a useful result are removed. For example, but not limited to, the training examples may include input and / or output values ​​that are outliers from typically encountered values, so that the machine learning algorithm using the training examples adapts to unlikely quantities as inputs and / or outputs. For example, values ​​that are more than a threshold number of standard deviations away from the average, mean, or expected value may be removed. Alternatively or additionally, one or more training examples may be identified as having low-quality data, where "low-quality" is defined as having a signal-to-noise ratio below a threshold.

[0118] By way of non-limiting example, and with further reference to FIG. 2 , images used to train an image classifier or other machine learning model and / or process that takes images as input or produces images as output may be rejected if their image quality is below a threshold. For example, without limitation, a computing device, processor, and / or module may perform blur detection and eliminate one or more blurs. Blur detection may be performed by, by way of non-limiting example, taking a Fourier transform or an approximation, such as a fast Fourier transform (FFT), of the image and analyzing the distribution of low and high frequencies in the resulting frequency domain representation of the image. The number of high frequency values ​​below a threshold level may indicate blur. In a further non-limiting example, blur detection may be performed by convolving the image, a channel of the image, or the like, with a Laplacian kernel, which may generate a numerical score reflecting the number of abrupt changes in intensity shown in the image, with a high score indicating sharpness and a low score indicating blur. Blur detection can be performed using gradient-based operators, which measure the operator based on the gradient or first derivative of the image, based on the hypothesis that abrupt changes indicate sharp edges in the image and therefore less blur. Blur detection may also be performed using wavelet-based operators, which exploit the ability of discrete wavelet transform coefficients to describe the frequency and spatial content of an image. Blur detection may also be performed using statistics-based operators, which utilize some image statistics as texture descriptors to calculate the focus level. Blur detection may also be performed using discrete cosine transform (DCT) coefficients to calculate the focus level of an image from its frequency content.

[0119] Still referring to FIG. 2 , a computing device, processor, and / or module may be configured to be preconditioned on one or more training examples. For example, but not by way of limitation, if a machine learning model and / or process has one or more inputs and / or outputs that require, transmit, or receive a certain number of bits, samples, or other units of data, elements of one or more training examples used as or compared to the inputs and / or outputs may be modified to have such units of data. For example, a computing device, processor, and / or module may convert a smaller number of units, such as an image with a low number of pixels, into a desired number of units, such as by upsampling or interpolation. As a non-limiting example, an image with a low number of pixels may have 100 pixels, but the desired number of pixels may be 128 pixels. The processor may interpolate the image with a low number of pixels to convert the 100 pixels into 128 pixels. It should also be noted that, after reading this disclosure, one skilled in the art will know various methods for interpolating a smaller number of data units, such as samples, pixels, or bits, into a desired number of such units. In some examples, the set of interpolation rules may be trained with highly detailed inputs and / or outputs, a corresponding set of inputs and / or outputs downsampled to a smaller number of units, and a neural network or other machine learning model trained to predict interpolated pixel values ​​using the training data. As a non-limiting example, sample inputs and / or outputs, such as a sample image having sample-augmented data units (e.g., additional pixels between original pixels), may be input to a neural network or machine learning model, which may output a pseudo-replica sample image in which pixels between the original pixels are assigned dummy values ​​based on the set of interpolation rules.As a non-limiting example, in the context of an image classifier, a machine learning model may have a set of interpolation rules trained with a set of high-definition images and images downsampled to a smaller number of pixels, and a neural network or other machine learning model trained using those examples to predict interpolated pixel values ​​in the context of facial images. As a result, inputs with sample-expanded data units (with dummy values ​​added between the original data units) can be run through the trained neural network and / or model to fill in values ​​to replace the dummy values. Alternatively or additionally, the processor, computing device, and / or module may utilize a sample expander technique, a low-pass filter, or both. A "low-pass filter," as used in this disclosure, is a filter that passes signals with frequencies below a selected cutoff frequency and attenuates signals with frequencies above the cutoff frequency. The exact frequency response of the filter depends on the filter design. The computing device, processor, and / or module may use averaging, such as luma averaging or chroma averaging, in the image to fill in data units between the original data units.

[0120] In some embodiments, with continued reference to FIG. 2 , a computing device, processor, and / or module may downsample elements of a training example to a desired smaller number of data elements. As a non-limiting example, a high-pixel-count image may have 256 pixels, but a desired number of pixels may be 128. The processor may downsample the high-pixel-count image to convert the 256 pixels to 128 pixels. In some embodiments, the processor may be configured to perform downsampling on the data. Downsampling, also known as decimation, may involve removing every Nth entry, all but every Nth entry, etc., in a sequence of samples, a process known as “compression,” which may be performed, for example, by an N-sample compressor implemented using hardware or software. Anti-aliasing and / or anti-imaging filters, and / or low-pass filters may be used to clean up compression artifacts.

[0121] Still referring to FIG. 2 , the machine learning module 200 may be configured to execute a lazy learning process 220 and / or protocol. This may alternatively be referred to as a “lazy loading” or “call-when-needed” process and / or protocol, in which machine learning is performed by combining the input and a training set upon receipt of the input to be converted into an output, and deriving an algorithm used to generate the output on demand. For example, an initial set of simulations may be run to cover initial heuristics and / or “first guesses” on outputs and / or relationships. As a non-limiting example, the initial heuristics may include ranking associations between the input and elements of the training data 204. The heuristics may include selecting several highest-ranked associations and / or elements of the training data 204. Lazy learning may be implemented using any suitable lazy learning algorithm, including, but not limited to, a K-nearest neighbor algorithm, a lazy Naive Bayes algorithm, etc., and upon reviewing this disclosure in its entirety, one skilled in the art will recognize a variety of lazy learning algorithms that may be applied to generate outputs as described in this disclosure, including, but not limited to, lazy learning applications of machine learning algorithms, as described in more detail below.

[0122] Alternatively or additionally, and continuing to refer to FIG. 2 , a machine learning process as described herein can be used to generate the machine learning model 224. As used herein, a “machine learning model” is a data structure that represents and / or instantiates a mathematical and / or algorithmic representation of a relationship between inputs and outputs, generated using any machine learning process, including, but not limited to, any of the processes described above, and stored in memory; once created, the inputs are submitted to the machine learning model 224, which generates an output based on the derived relationship. For example, without limitation, a linear regression model generated using a linear regression algorithm can calculate a linear combination of input data using coefficients derived during the machine learning process to calculate output data. As a further non-limiting example, the machine learning model 224 may be generated by creating an artificial neural network, such as a convolutional neural network, that includes an input layer of nodes, one or more hidden layers, and an output layer of nodes. Connections between nodes may be created through a process of "training" the network, in which elements from a set of training data 204 are applied to the input nodes, and an appropriate training algorithm (such as Levenberg-Marquardt, conjugate gradient, simulated annealing, or other algorithm) is then used to adjust the connections and weights between nodes in adjacent layers of the neural network to produce desired values ​​at the output nodes. This process is sometimes referred to as deep learning.

[0123] Still referring to FIG. 2 , the machine learning algorithm can include at least one supervised machine learning process 228. At least one supervised machine learning process 228, as defined herein, includes an algorithm that receives a training set relating a number of inputs to a number of outputs and attempts to generate one or more data structures that represent and / or instantiate one or more mathematical relationships relating the inputs to the outputs, each of the one or more mathematical relationships being optimal according to some criteria specified to the algorithm using some scoring function. For example, the supervised learning algorithm can include an image of a region of interest, as described above, as input, a determination of whether a sample is present as output, and a scoring function that represents the desired form of the relationship to be found between the input and the output. The scoring function can, for example, attempt to maximize the probability that a given input and / or combination of elements of the inputs is associated with a given output and minimize the probability that a given input is not associated with a given output. The scoring function may be expressed as a risk function that represents the "expected loss" of the algorithm relating inputs to outputs, where the loss is calculated as an error function that represents the degree to which the prediction produced by the relationship is inaccurate when compared to a given input-output pair provided in the training data 204. Those skilled in the art will recognize, upon reviewing this disclosure in its entirety, various possible variations of at least one supervised machine learning process 228 that may be used to determine the relationship between inputs and outputs. The supervised machine learning process may include a classification algorithm as defined above.

[0124] With further reference to FIG. 2 , training a supervised machine learning process can include iteratively updating coefficients, biases, and weights based on, but not limited to, an error function, an expected loss, and / or a risk function. For example, outputs generated by a supervised machine learning model using example inputs from training examples can be compared to example outputs from the training examples, and an error function can be generated based on the comparison, which may include any error function suitable for use with any machine learning algorithm described in this disclosure, including the squared difference between one or more sets of compared values, etc. Such error functions can then be used to update one or more weights, biases, coefficients, or other parameters of the machine learning model through any suitable process, including, but not limited to, a gradient descent process, a least-squares process, and / or other processes described in this disclosure. This may be done iteratively and / or recursively to gradually adjust the weights, biases, coefficients, or other parameters. The updates may be performed using one or more backpropagation algorithms in neural networks. Iterative and / or recursive updates to weights, biases, coefficients, or other parameters as described above may be performed until the currently available training data is exhausted and / or a convergence test is passed, where a "convergence test" is a test on a condition selected as indicating that the model and / or its weights, biases, coefficients, or other parameters have reached a certain degree of accuracy. The convergence test may, for example, compare the difference between two or more consecutive error or error function values, where a difference below a threshold may be considered to indicate convergence. Alternatively or additionally, one or more error and / or error function values ​​evaluated in a training iteration may be compared to a threshold.

[0125] Still referring to FIG. 2 , a computing device, processor, and / or module may be configured to perform the methods, method steps, sequences of method steps, and / or algorithms described with reference to this figure in any order and to any degree of repetition. For example, a computing device, processor, and / or module may be configured to repeatedly perform a single step, sequence, and / or algorithm until a desired or commanded result is achieved. The repetition of a step or sequence of steps may be performed iteratively and / or recursively using the output of a previous iteration as input for a subsequent iteration, aggregating the inputs and / or outputs of an iteration to produce an aggregate result, decreasing or decrementing one or more variables such as global variables, and / or dividing a large processing task into a set of smaller processing tasks that are addressed iteratively. The computing device, processor, and / or module may execute any step, sequence of steps, or algorithm in parallel, such as performing a step simultaneously and / or nearly simultaneously two or more times using two or more parallel threads, processor cores, etc., and the division of tasks among parallel threads and / or processes may be performed according to any protocol suitable for dividing tasks among iterations. Those skilled in the art will recognize, upon reviewing this disclosure in its entirety, various ways in which steps, sequences of steps, processing tasks, and / or data may be subdivided, shared, or otherwise processed using iterative, recursive, and / or parallel processing.

[0126] 2, the machine learning process can include at least one unsupervised machine learning process 232. An unsupervised machine learning process, as used herein, is a process that draws inferences in a dataset without regard to labels, such that the unsupervised machine learning process is free to discover any structure, relationships, and / or correlations provided in the data. The unsupervised machine learning process 232 may not require a response variable, and the unsupervised machine learning process 232 can be used to find interesting patterns and / or inferences between variables, determine the degree of correlation between two or more variables, etc.

[0127] Still referring to FIG. 2 , the machine learning module 200 can be designed and configured to create the machine learning model 224 using techniques for developing linear regression models. The linear regression model may include ordinary least squares regression, which aims to minimize the squared difference between predicted and actual results according to an appropriate norm (e.g., a vector space distance norm) that measures such difference. The coefficients of the resulting linear equation can be modified to improve the minimization. The linear regression model may include a ridge regression method, in which the function being minimized may include, in addition to the least squares function, a term that multiplies the square of each coefficient by a scalar to penalize large coefficients. The linear regression model may also include a least absolute shrinkage selection operator (lasso) model, in which ridge regression is combined with multiplying the least squares term by a factor of 1 divided by twice the number of samples. The linear regression model may also include a multitasking lasso model, in which the norm applied to the least squares term in the lasso model is the Frobenius norm, which corresponds to the square root of the sum of the squares of all terms. The linear regression model may include an elastic net model, a multitask elastic net model, a least-angle regression model, a LARS lasso model, an orthogonal matching pursuit model, a Bayesian regression model, a logistic regression model, a stochastic gradient descent model, a perceptron model, a passive-aggressive algorithm, a robustness regression model, a Huber regression model, or other suitable models that would occur to one of ordinary skill in the art upon reviewing this disclosure in its entirety. In embodiments, the linear regression model may be generalized to a polynomial regression model, thereby finding a polynomial (e.g., quadratic, cubic, or higher order) that provides the best predicted output / actual output fit. As would be apparent to one of ordinary skill in the art upon reviewing this disclosure in its entirety, methods similar to those described above can be applied to minimize the error function.

[0128] Still referring to FIG. 2 , the machine learning algorithm may include, but is not limited to, linear discriminant analysis. The machine learning algorithm may include quadratic discriminant analysis. The machine learning algorithm may include kernel ridge regression. The machine learning algorithm may include support vector machines, including, but not limited to, regression processes based on support vector classification. The machine learning algorithm may include stochastic gradient descent algorithms, including classification and regression algorithms based on stochastic gradient descent. The machine learning algorithm may include nearest neighbor algorithms. The machine learning algorithm may include various forms of latent space regularization, such as variational regularization. The machine learning algorithm may include Gaussian processes, such as Gaussian process regression. The machine learning algorithm may include cross-decomposition algorithms, including partial least squares and / or canonical correlation analysis. The machine learning algorithm may include naive Bayes methods. The machine learning algorithm may include decision tree-based algorithms, such as decision tree classification and regression algorithms. The machine learning algorithm may include ensemble methods, such as bagging meta-estimators, forests of random trees, AdaBoost, gradient tree boosting, and / or voting classifier methods. The machine learning algorithms may include neural net algorithms, including convolutional neural net processing.

[0129] Still referring to FIG. 2 , the machine learning model and / or process may be deployed or instantiated by incorporation into a program, device, system, and / or module. For example, without limitation, the machine learning model, neural network, and / or some or all of its parameters may be stored and / or deployed in any memory or circuit. Parameters such as coefficients, weights, and / or biases may be stored as circuit-based constants, such as arrays of wires and / or binary inputs and / or outputs set to logic “1” and “0” voltage levels in a logic circuit to represent numbers according to any suitable encoding system, including binary complement, or may be stored in any volatile and / or non-volatile memory. Similarly, mathematical operations and data inputs and / or outputs to or from the model, neural network layers, etc. may be instantiated in hardware circuitry and / or in the form of firmware, machine code such as binary opcode instructions, assembly language, or instructions in any higher-level programming language. To instantiate a machine learning process and / or model, any technique for hardware and / or software instantiation of memory, instructions, data structures, and / or algorithms may be used, which may include, but is not limited to, the manufacture and / or configuration of non-reconfigurable hardware elements, circuits, and / or modules, such as, but not limited to, ASICs; the manufacture and / or configuration of reconfigurable hardware elements, circuits, and / or modules, such as, but not limited to, FPGAs; the manufacture and / or configuration of non-reconfigurable memory elements, circuits, and / or modules, such as, but not limited to, non-reconfigurable ROMs; the manufacture and / or configuration of reconfigurable and / or reconfigurable memory elements, circuits, and / or modules, such as, but not limited to, reconfigurable ROMs or other memory technologies described in this disclosure; and / or any combination of the manufacture and / or configuration of any computing device and / or its components, as described in this disclosure.Such deployed and / or instantiated machine learning models and / or algorithms may receive inputs from, and generate outputs for, any other processes, modules, and / or components described in this disclosure.

[0130] 2 , any process of training, retraining, deployment, and / or instantiation of a machine learning model and / or algorithm may be performed and / or repeated after initial deployment and / or instantiation to modify, refine, and / or improve the machine learning model and / or algorithm. Such retraining, deployment, and / or instantiation may be performed as a periodic or periodic process, for example, after a quantity measure such as the number of bytes or other measure of processed data, the number of uses or executions of the processes described in this disclosure, and / or according to a software, firmware, or other update schedule, with retraining, deployment, and / or instantiation at regular elapsed time periods. Alternatively or additionally, retraining, deployment, and / or instantiation may be event-based and may be triggered, without limitation, by user input indicating suboptimal or otherwise problematic performance, and / or by an automated field testing and / or audit process that may compare the output of the machine learning model and / or algorithm, and / or error and / or its error function, to any threshold value, convergence determination, etc., and / or compare the output of the processes described herein to similar threshold values, convergence determinations, etc. Event-based retraining, deployment, and / or instantiation may alternatively or additionally be triggered by the receipt and / or generation of one or more new training examples, where the number of new training examples may be compared to a pre-configured threshold, and exceeding the pre-configured threshold may trigger retraining, deployment, and / or instantiation.

[0131] Still referring to FIG. 2 , retraining and / or additional training may be performed using any of the processes for training described above, using any version of a current or previously deployed machine learning model and / or algorithm as a starting point. Training data for retraining may be collected, preprocessed, sorted, classified, sanitized, or otherwise processed according to any process described in this disclosure. Training data may include, but is not limited to, training examples including inputs and associated outputs used, received, and / or generated from any version of any system, module, machine learning model or algorithm, apparatus, and / or method described in this disclosure; such examples may be modified and / or labeled according to user feedback or other processing to indicate desired results, and / or may have actual or measured results from a process being modeled and / or predicted by a system, module, machine learning model or algorithm, apparatus, and / or method as “desired” results to be compared with outputs for such training processes.

[0132] The redeployment may be performed using any reconfiguration and / or rewriting of reconfigurable and / or rewritable circuitry and / or memory elements; alternatively, the redeployment may be performed by fabrication of new hardware and / or software components, circuits, instructions, etc., which may be added to and / or replace existing hardware and / or software components, circuits, instructions, etc.

[0133] 2 , one or more of the processes or algorithms described above may be performed by at least one dedicated hardware unit 236. For purposes of this figure, a “dedicated hardware unit” is a hardware component, circuitry, etc., other than the main control circuitry and / or processor that performs the method steps described in this disclosure, that is specifically designated or selected to perform one or more particular tasks and / or processes described with reference to this figure, such as, but not limited to, preconditioning and / or sanitizing training data and / or training machine learning algorithms and / or models. The dedicated hardware unit 236 may include hardware units capable of efficiently performing repetitive or intensive calculations, such as, but not limited to, matrix-based calculations that update or adjust parameters, weights, coefficients, and / or biases of machine learning models and / or neural networks, using pipelined processing, parallel processing, etc., and such hardware units may be optimized for such processing by including dedicated circuitry for matrix operations and / or signal processing operations, including, for example, multiple arithmetic and / or logic circuit units, such as multipliers and / or adders, that can operate simultaneously and / or in parallel, etc. Such special-purpose hardware units 236 may include, but are not limited to, graphics processing units (GPUs), special-purpose signal processing modules, FPGAs, or other reconfigurable hardware configured to instantiate parallel processing units for one or more specific tasks. A computing device, processor, apparatus, or module may be configured to direct one or more special-purpose hardware units 236 to perform one or more operations described herein, such as evaluation of model and / or algorithm outputs, one-time or iterative updates of parameters, coefficients, weights, and / or biases, and / or any other operations, such as vector and / or matrix operations, described in this disclosure.

[0134] Referring now to FIG. 3, an exemplary embodiment of a neural network 300 is illustrated. A neural network 300, also known as an artificial neural network, is a network of "nodes" or data structures having one or more inputs, one or more outputs, and a function that determines the output based on the inputs. Such nodes may be organized into networks such as, but not limited to, a convolutional neural network, including an input layer of nodes 304, one or more hidden layers 308, and an output layer of nodes 312. Connections between nodes may be created through a process of "training" the network, in which elements from a set of training data are applied to the input nodes, and an appropriate training algorithm (such as Levenberg-Marquardt, conjugate gradient, simulated annealing, or other algorithms) is then used to adjust the connections and weights between nodes in adjacent layers of the neural network to produce desired values ​​at the output nodes. This process is sometimes referred to as deep learning. Connections may only run from input nodes to output nodes in a "feedforward" network, or the output of one layer may be fed back to the input of the same or a different layer in a "recurrent network." As a further non-limiting example, a neural network can include a convolutional neural network that includes an input layer of nodes, one or more hidden layers, and an output layer of nodes. As used in this disclosure, a "convolutional neural network" is a neural network in which at least one hidden layer is a convolutional layer that convolves the input to that layer with a subset of the input known as a "kernel," along with one or more additional layers such as a pooling layer, a fully connected layer, etc.

[0135] 4, an exemplary embodiment of a neural network node 400 is illustrated. The node has multiple inputs x that can receive values ​​from, but are not limited to, inputs to the neural network that contains the node and / or from other nodes. iA node may implement one or more activation functions to generate its output given one or more inputs. Activation functions include, but are not limited to, binary step functions that compare an input to a threshold and output a logic 1 or logic 0 output or the like, linear activation functions where the output is directly proportional to the input, and / or non-linear activation functions where the output is not proportional to the input. Non-linear activation functions include, but are not limited to, binary step functions that compare an input x to a threshold and output a logic 1 or logic 0 output or the like.

number

number

number

number

number

number

[0136] 4, a "convolutional neural network," as used in this disclosure, is a neural network in which at least one hidden layer is a convolutional layer that convolves the input to that layer with a subset of the input known as a "kernel," along with one or more additional layers, such as a pooling layer, a fully connected layer, etc. CNNs can include, but are not limited to, extensions of deep neural networks (DNNs), which are defined as neural networks with two or more hidden layers.

[0137] Still referring to FIG. 4 , in some embodiments, a convolutional neural network can learn from images. In non-limiting examples, a convolutional neural network may perform tasks such as image classification, detecting objects depicted in images, image segmentation, and / or image processing. In some embodiments, a convolutional neural network may operate such that each node in the input layer is connected only to regions of nodes in the hidden layer. In some embodiments, the regions may collectively create a map of features from the input layer to the hidden layer. In some embodiments, a convolutional neural network may include layers in which all node weights and biases are identical. In some embodiments, this allows the convolutional neural network to detect features, such as edges, at different locations within an image.

[0138] Referring now to FIG. 5, an exemplary embodiment of a slide imaging method 500 is illustrated, which may help reduce errors due to misfocusing. In step 505, an entire slide image may be captured at a low resolution (e.g., 1x). In step 510, image segmentation may be performed to detect regions where material is present. This may delimit all regions on the slide, including dust areas, pen marks, and printed text (e.g., annotations), with bounding boxes. In step 515, a sample presence probability score may be calculated for each bounding box encompassing a region of interest. This may be performed using various algorithms, such as K-means, NLMD (non-local means denoising), features such as hue color space for each row, or segmentation models such as U-Net. This may be used to find the best row 520. The best row may include rows sandwiched between the rows above and below, with the sum of the weighted scores for that row representing the highest score for sample presence. Once the best row is determined, the best (x,y) location may be determined. The best (x,y) location may include the (x,y) point that maximizes the probability of a sample being present at that point based on specimen boundaries, color, etc. The best focus of that bounding box may be determined by collecting a Z-stack at the best (x,y) point 525. As used herein, a "Z-stack" refers to multiple images taken at a particular (x,y) location using varying focal lengths. In some embodiments, a Z-stack may have a height less than the height of the slide. In some embodiments, a Z-stack may include 2, 3, 4, 5, 6, 7, 8, 9, 10, or more images captured at various focal lengths. In some embodiments, images captured as part of a Z-stack may be spaced less than 1 millimeter apart. For example, capturing fewer images may increase efficiency by increasing speed. Capturing images closer together may allow for more accurate determination of the optimal focal length. In some embodiments, these factors may make it desirable to capture a relatively small number of images over a relatively small height.This may increase the importance of selecting a focal distance for capturing a Z-stack so that the object to be focused is located within the Z-stack height. The focal pattern may be used to estimate the focal distance at the (x,y) location where the Z-stack is captured. Using such an estimate can improve confidence that the object to be focused is within the Z-stack height. In some embodiments, the distance of the Z-stack may be identified and / or captured using a plane in the camera field of view and / or at least two points along its row. Such a plane may include a focal pattern as described herein. In steps 530 and 535, the Z level may be scanned through the row and used to identify a focal pattern, such as a plane. A focal pattern, such as a plane, may be identified using a set of points along the row as a function of the best focus of the points in that row. Such a plane may be used to estimate the focal distance at other locations, such as the location of an adjacent row. Such a plane may be recalculated based on additional data as new data is acquired. For example, best focus may be identified at an additional point, and the plane may be updated based on this additional data. Such a plane may also be used to estimate the focal distance in another region of interest 540. This procedure can be repeated for all regions of interest on slide 545 .

[0139] Referring now to FIGS. 6A-6C, a slide progression through various steps described herein is illustrated. Slide 604 may include annotations 608A-608B, dust 612A-612B, and samples 616A-616B. The steps described herein address the challenge of selecting the correct focus for regions where samples are present, despite the presence of dust and annotations. By dividing slide 604 into separate regions of interest 620A-620F, each segment can be scanned at a different focus optimal for that region. This makes focus decisions independent of the spatial distribution of samples, dust, and annotations. In some embodiments, image classification to detect dust and annotations to avoid scanning is best performed downstream rather than during scanning. In some embodiments, this addresses the challenge of running advanced models live on the scanning device during scanning and the risk of false positives when classifying regions as dust or annotations and skipping scanning the associated regions. Additionally, annotations that may be useful for downstream tasks may need to be scanned for use in downstream multimodal learning. In some embodiments, segmentation followed by a best focus determination can be a useful approach to find all regions of interest at the optimal focus obtained with the best row estimate.

[0140] 7, an exemplary embodiment of a slide imaging method 700 is illustrated. One or more steps of method 700 may be implemented as, but not limited to, described herein with reference to other figures. One or more steps of method 700 may be implemented using, but not limited to, at least one processor.

[0141] Still referring to FIG. 7, in some embodiments, the method 700 may include receiving at least one region of interest 705 .

[0142] Still referring to FIG. 7, in some embodiments, method 700 may include capturing a first image of a slide at a first position within at least one region of interest 710 using at least one optical system.

[0143] 7, in some embodiments, method 700 may include identifying a focus pattern as a function of the first image and the first location 715. In some embodiments, identifying the focus pattern includes identifying a row that includes the first location, capturing a plurality of first images at the first location, each of the plurality of first images having a different focal length, determining an optimally focused first image among the plurality of best-focus first images, and identifying the focus pattern using the focal lengths of the plurality of optimally focused images at a set of points along the row. In some embodiments, the row further includes a second location, and the method further includes capturing, using at least one processor and at least one optical system, a plurality of second images at the second location, each of the plurality of second images having a different focal length; determining, using the at least one processor, an optimally focused second image among the plurality of best-focused second images; identifying, using the at least one processor, a focus pattern using the focal lengths of the optimally focused first image and the optimally focused second image; and extrapolating, using the at least one processor, a third focal length for the third location as a function of the focus pattern. In some embodiments, the third location is located outside the row. In some embodiments, the third location is located within a different region of interest than the first location. In some embodiments, identifying the row includes identifying the row based on a first row sample presence score from a first set of row sample presence scores. In some embodiments, identifying a row based on the first row sample presence score includes determining a row from a second set of sample presence scores whose adjacent row has the highest sample presence score, wherein the second set of sample presence scores is determined using machine vision. In some embodiments, identifying a plane includes identifying a plurality of points and a plurality of best focuses at the plurality of points, and generating the plane as a function of a subset of the plurality of points and a corresponding subset of the best focuses at the points.In some embodiments, identifying the focus pattern further includes updating the focus pattern, and updating the focus pattern includes identifying the additional point and a best focus at the additional point, and updating the focus pattern as a function of the additional point and the best focus at the additional point.

[0144] Still referring to FIG. 7, in some embodiments, the method 700 may include extrapolating the focal length of the second position as a function of the focal pattern 720.

[0145] 7, in some embodiments, method 700 may include capturing a second image of the slide using at least one optical system at a second position and focal length 725. In some embodiments, capturing the second image includes capturing multiple images taken at focal lengths based on the focus pattern and constructing the second image from the multiple images.

[0146] Still referring to FIG. 7, in some embodiments, the method 700 may further include moving the movable element to the second position using an actuator mechanism.

[0147] Still referring to FIG. 7, in some embodiments, the method 700 may further include using machine vision to identify the point within the row having the maximum point sample presence score.

[0148] Still referring to FIG. 7, in some embodiments, method 700 may further include capturing a low magnification image of the slide using at least one processor and an optical system, the low magnification image having a magnification lower than that of the first image, identifying at least one region of interest in the low magnification image using at least one processor and machine vision, and determining, using at least one processor, whether the sample is included in either the first image or the second image.

[0149] It should be noted that any one or more of the aspects and embodiments described herein may be suitably implemented using one or more machines (e.g., one or more computing devices utilized as user computing devices for electronic documents, one or more server devices, such as document servers) programmed in accordance with the teachings herein, as would be apparent to those skilled in the computer arts. Appropriate software coding may be readily achieved by skilled programmers based on the teachings of the present disclosure, as would be apparent to those skilled in the software arts. The above-described aspects and implementations employing software and / or software modules may also include appropriate hardware to support the implementation of the machine-executable instructions of the software and / or software modules.

[0150] Such software may be a computer program product employing a machine-readable storage medium. A machine-readable storage medium may be any medium capable of storing and / or encoding a sequence of instructions for execution by a machine (e.g., a computing device) and causing the machine to perform any one of the methodologies and / or embodiments described herein. Examples of machine-readable storage media include, but are not limited to, magnetic disks, optical disks (e.g., CDs, CD-Rs, DVDs, DVD-Rs, etc.), magneto-optical disks, read-only memory "ROM" devices, random-access memory "RAM" devices, magnetic cards, optical cards, solid-state memory devices, EPROMs, EEPROMs, and any combination thereof. As used herein, machine-readable medium is intended to include not only single media but also physically separate collections of media, such as, for example, a collection of compact discs, one or more hard disk drives in combination with computer memory, etc. As used herein, machine-readable storage medium does not include transitory forms of signal transmission.

[0151] Such software may also include information (e.g., data) carried in a data signal on a data carrier such as a carrier wave. For example, the machine-executable information may be included in a data carrier signal embodied in a data carrier, the signal encoding a sequence of instructions, or portions thereof, for execution by a machine (e.g., a computing device), and any associated information (e.g., data structures and data) that cause the machine to perform any one of the methodologies and / or embodiments described herein.

[0152] Examples of computing devices include, but are not limited to, e-book reading devices, computer workstations, terminal computers, server computers, handheld devices (e.g., tablet computers, smartphones, etc.), web appliances, network routers, network switches, network bridges, any machine capable of executing a sequence of instructions that specify actions to be performed by that machine, and any combination thereof. In one example, a computing device may include and / or be included in a kiosk.

[0153] 8 shows a diagrammatic representation of one embodiment of a computing device in the exemplary form of a computer system 800 upon which a set of instructions may be executed that causes a control system to perform any one or more of the aspects and / or methodologies of the present disclosure. It is also contemplated that multiple computing devices may be utilized to execute a set of instructions specifically configured to cause one or more of the devices to perform any one or more of the aspects and / or methodologies of the present disclosure. Computer system 800 includes a processor 804 and memory 808 that communicate with each other and with other components via a bus 812. Bus 812 may include any of several types of bus structures, including, but not limited to, a memory bus, a memory controller, a peripheral bus, a local bus, and any combination thereof, using any of a variety of bus architectures.

[0154] Processor 804 may include any suitable processor, such as, but not limited to, a processor incorporating logic circuitry for performing arithmetic and logical operations, such as an arithmetic logic unit (ALU), which may be controlled by a state machine and directed by operational input from memory and / or sensors. Processor 804 may be configured according to, by way of non-limiting example, the von Neumann architecture and / or the Harvard architecture. Processor 804 may include, incorporate, and / or be incorporated into, without limitation, a microcontroller, a microprocessor, a digital signal processor (DSP), a field programmable gate array (FPGA), a complex programmable logic device (CPLD), a graphical processing unit (GPU), a general-purpose GPU, a tensor processing unit (TPU), an analog or mixed signal processor, a trusted platform module (TPM), a floating-point unit (FPU), and / or a system-on-chip (SoC).

[0155] Memory 808 may include a variety of components (e.g., machine-readable media), including, but not limited to, random-access memory components, read-only components, and any combination thereof. In one example, a basic input / output system 816 (BIOS), containing the basic routines that help to transfer information between elements within computer system 800, such as during start-up, may be stored in memory 808. Memory 808 may also include (e.g., stored on one or more machine-readable media) instructions (e.g., software) 820 that embody any one or more of the aspects and / or methodologies of the present disclosure. In another example, memory 808 may further include any number of program modules, including, but not limited to, an operating system, one or more application programs, other program modules, program data, and any combination thereof.

[0156] Computer system 800 may also include a storage device(s) 824. Examples of storage devices (e.g., storage device 824) include, but are not limited to, hard disk drives, magnetic disk drives, optical disk drives in combination with optical media, solid-state memory devices, and any combination thereof. Storage device 824 may be connected to bus 812 by an appropriate interface (not shown). Examples of interfaces include, but are not limited to, SCSI, Advanced Technology Attachment (ATA), Serial ATA, Universal Serial Bus (USB), IEEE 1394 (FIREWIRE), and any combination thereof. In one example, storage device 824 (or one or more components thereof) may be removably interfaced with computer system 800 (e.g., via an external port connector (not shown)). In particular, storage device 824 and associated machine-readable media 828 may provide nonvolatile and / or volatile storage of machine-readable instructions, data structures, program modules, and / or other data for computer system 800. In one example, the software 820 may reside, completely or partially, within the machine-readable medium 828. In another example, the software 820 may reside, completely or partially, within the processor 804.

[0157] Computer system 800 may also include input devices 832. In one example, a user of computer system 800 may input commands and / or other information into computer system 800 via input devices 832. Examples of input devices 832 include, but are not limited to, alphanumeric input devices (e.g., keyboards), pointing devices, joysticks, gamepads, audio input devices (e.g., microphones, voice response systems, etc.), cursor control devices (e.g., mice), touchpads, optical scanners, video capture devices (e.g., still cameras, video cameras), touch screens, and any combination thereof. Input devices 832 may interface with bus 812 via any of a variety of interfaces (not shown), including, but not limited to, a serial interface, a parallel interface, a game port, a USB interface, a FIREWIRE interface, a direct interface to bus 812, and any combination thereof. Input devices 832 may include a touchscreen interface, which may be part of or separate from display 836, as described below. Input devices 832 may be utilized as a user selection device for selecting one or more graphical representations in a graphical interface, as described above.

[0158] A user may also input commands and / or other information into computer system 800 via storage device 824 (e.g., a removable disk drive, flash drive, etc.) and / or network interface device 840. A network interface device, such as network interface device 840, may be utilized to connect computer system 800 to one or more of various networks, such as network 844, and one or more remote devices 848 connected thereto. Examples of network interface devices include, but are not limited to, a network interface card (e.g., a mobile network interface card, a LAN card), a modem, and any combination thereof. Examples of networks include, but are not limited to, a wide area network (e.g., the Internet, an enterprise network), a local area network (e.g., a network associated with an office, building, campus, or other relatively small geographic space), a telephone network, a data network associated with a telephone / voice provider (e.g., a mobile communications provider's data and / or voice network), a direct connection between two computing devices, and any combination thereof. A network, such as network 844, may employ wired and / or wireless communication modes. In general, any network topology may be used. Information (eg, data, software 820 , etc.) may be communicated to and / or from computer system 800 via network interface device 840 .

[0159] Computer system 800 may further include a video display adapter 852 that communicates displayable images to a display device, such as display device 836. Examples of display devices include, but are not limited to, a liquid crystal display (LCD), a cathode ray tube (CRT), a plasma display, a light emitting diode (LED) display, and any combination thereof. Display adapter 852 and display device 836 can be utilized in combination with processor 804 to provide graphical representations of aspects of the present disclosure. In addition to a display device, computer system 800 may include one or more other peripheral output devices, including, but not limited to, audio speakers, a printer, and any combination thereof. Such peripheral output devices may be connected to bus 812 via peripheral interface 856. Examples of peripheral interfaces include, but are not limited to, a serial port, a USB connection, a FIREWIRE connection, a parallel connection, and any combination thereof.

[0160] The foregoing is a detailed description of exemplary embodiments of the present invention. Various modifications and additions may be made without departing from the spirit and scope of the present invention. Features of each of the various embodiments described above may be combined with features of other described embodiments as appropriate to provide various combinations of features in related new embodiments. Moreover, while a number of separate embodiments have been described above, what has been described herein is merely illustrative of the application of the principles of the present invention. Furthermore, while certain methods herein may be illustrated and / or described as being performed in a particular order, such orders may be varied considerably within the skill of ordinary skill in order to implement methods, systems, and software according to the present disclosure. Accordingly, the present description is intended to be illustrative only, and not to otherwise limit the scope of the present invention.

[0161] Exemplary embodiments are disclosed above and illustrated in the accompanying drawings. Those skilled in the art will appreciate that various modifications, omissions, and additions may be made to what is specifically disclosed herein without departing from the spirit and scope of the invention.

Claims

1. 1. An apparatus for imaging a slide, comprising: at least one optical system including an optical sensor; at least one processor; a memory communicatively coupled to the at least one processor; The memory includes a plurality of memory areas, each of which ... capturing a plurality of first images at a first position of the slide using the at least one optical system, each of the plurality of first images having a different focal length; identifying a focus pattern using the focal lengths of each of the plurality of optimally focused images at the set of points along the row; and Including, extrapolating a focal length at a second position as a function of the focal pattern; capturing a second image of the slide at the second position and at the focal length using the at least one optical system; A device that stores instructions that cause processing to occur.

2. 10. The apparatus of claim 1, further comprising an actuator mechanism mechanically connected to the at least one optical system, the actuator mechanism configured to move the at least one optical system to the second position.

3. 3. The device of claim 2, wherein the actuator mechanism is in electronic communication with an actuator control, the actuator control configured to operate the actuator mechanism based on input received from a user interface including at least an input interface.

4. The apparatus of claim 1 , wherein the at least one processor is configured to determine a region of interest on a slide, and the first location is located within the region of interest.

5. The device of claim 4, wherein determining the region of interest includes using a sample identification machine learning model trained based on a dataset associating exemplary images of slides and segments of slide images with the presence or absence of samples, accepting a second image as input, and outputting a determination result.

6. 5. The apparatus of claim 4, wherein determining the region of interest includes using a machine learning model for identifying regions of interest (ROI identification machine learning model) trained based on a dataset including example images of slides associated with example regions of the images in which features are present, accepting slide images as input and outputting data regarding the location of the region of interest.

7. The apparatus of claim 4 , further comprising an image processing module configured to determine a measure of depiction quality in the region of interest.

8. The apparatus of claim 7 , wherein determining the measure of rendering quality comprises performing a blur detection process.

9. identifying a focus pattern includes identifying rows from the plurality of first images; The apparatus of claim 1 , wherein identifying the row comprises identifying the row based on a first row sample presence score from a first set of row sample presence scores.

10. identifying a focus pattern includes identifying rows from the plurality of first images; The apparatus of claim 1 , wherein identifying the row comprises identifying a point within the row having a maximum point sample presence score.

11. 1. A method of imaging a slide, comprising: capturing, with at least one processor and at least one optical system, a plurality of first images at a first position of the slide, each of the plurality of first images having a different focal length, the optical system including an optical sensor; identifying, by the at least one processor, a focus pattern using focal lengths for each of a plurality of optimally focused images at a set of points along the row; and and extrapolating, by the at least one processor, a focal length at a second position as a function of the focus pattern; capturing a second image of the slide at the second position and the focal length with the at least one processor and the at least one optical system; A method comprising:

12. 12. The method of claim 11 , wherein capturing the second image further comprises using an actuator mechanism mechanically connected to the at least one optical system, the actuator mechanism configured to move the optical system to the second position.

13. The method of claim 12 , wherein the actuator mechanism is in electronic communication with an actuator control, the actuator control operating the actuator mechanism based on input received from a user interface including at least an input interface.

14. The method of claim 11 , further comprising determining, by the at least one processor, a region of interest on a slide, the first location being located within the region of interest.

15. The method of claim 14, wherein determining the region of interest includes using a sample identification machine learning model trained based on a dataset associating exemplary images of slides and segments of slide images with the presence or absence of samples, accepting a second image as input, and outputting a determination result.

16. 15. The method of claim 14, wherein determining the region of interest includes using a machine learning model for identifying regions of interest (ROI identification machine learning model) trained based on a dataset including example images of slides associated with example regions of the images in which features are present, accepting slide images as input, and outputting data regarding the location of the region of interest.

17. The method of claim 14 , further comprising using an image processing module configured to determine a measure of depiction quality in the region of interest.

18. The method of claim 17 , wherein determining the measure of rendering quality comprises performing a blur detection process.

19. identifying a focus pattern includes identifying rows from the plurality of first images; The method of claim 11 , wherein identifying the row comprises identifying the row based on a first row sample presence score from a first set of row sample presence scores.

20. identifying a focus pattern includes identifying rows from the plurality of first images; The apparatus of claim 1 , wherein identifying the row comprises identifying a point within the row having a maximum point sample presence score.