LEARNING AUTOFOCUS
Patent Information
- Application Number
- DE502019013628
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-11-20
- Filing Date
- 2019-11-20
- Publication Date
- 2025-08-07
- Estimated Expiration
- 2039-11-20
AI Technical Summary
Conventional autofocus systems struggle with adapting to different object characteristics, require multiple images for focus determination, and are inefficient in defining optimal 3D scanning areas, leading to potential sample damage and increased costs.
Utilizing trained models, particularly neural networks, to analyze captured images for rapid focus position determination, enabling precise focus prediction with minimal image acquisition and adaptability across various imaging systems.
The method allows for accurate and efficient focus positioning with reduced sample damage and improved adaptability, supporting a wide range of imaging modalities and conditions, including confocal microscopes and digital cameras.
Description
[0001] The invention relates to a method and a device for determining a focus position using trained models.
[0002] Visual perception is an important component of human physiology. Optical stimuli in the human eye allow us to gain information from our surroundings. The human eye can focus on objects in the near and far field. This dynamic adjustment of the eye's refractive power is also called accommodation. Focusing on objects is also important in technical imaging processes for a sharp representation of the object. The technique of automatically focusing on objects is also known as autofocus. However, the current state of the art reveals some disadvantages and problems with existing systems, which are discussed below.
[0003] Conventional autofocus systems expect a clearly defined plane of focus, so that a single plane is in focus and all surrounding parallel planes are out of focus. The excellent plane of focus differs from the surrounding planes in that only it contains a maximum degree of sharpness. This means that all other parallel planes produce a blurred image of an object. This type of focus system has several disadvantages. For example, it cannot be used when several, or all, image planes are in focus. This is the case by definition with confocal microscope systems, for example. With confocal microscopes, every image plane is in focus, so the statement "in focus" is not true. There are no "sharp" and "blurred" planes. Another disadvantage of state-of-the-art autofocus systems is that they cannot be adapted to different object characteristics.In addition, scan areas, or the start and end points of scan areas, cannot be defined for microscope images.
[0004] Another well-known method for automatic focusing is the laser reflection method. Laser reflection methods require a clearly measurable reflection from a reference surface in order to work accurately. This type of focus system has the disadvantage that the degree of optical reflection depends on the optical conditions. The use of immersion fluids can reduce the reflection intensity and push it below a measurable value. An optically precisely adjusted transition between the objective lens, immersion fluid, and specimen slide would reduce the reflection component R to near zero, thus blocking reflection-based autofocus systems. Furthermore, laser reflection methods require that the object under investigation always remains at the same distance from the reflection surface. However, this is not the case with moving objects.In living cells, for example, the cell can "migrate," meaning its distance from the measurable reflection surface is not equidistant. Furthermore, laser reflection techniques depend on the wavelength used. Especially in confocal systems, the wavelength can influence the signal-to-noise ratio. Furthermore, it is impossible to determine whether an object is actually present in a scan area. Reflection techniques always capture a stack of images, even if there is no object in the scanned area that can be scanned. For large numbers of scans, this results in a loss of time and storage space, which can result in high costs.
[0005] Another method used in autofocus systems is the cross-sectional imaging method. Cross-sectional imaging methods have the disadvantage that they cannot be used in microscopy and are limited to standard camera technology.
[0006] A problem with all known automatic focusing methods is that none of them is capable of predicting not only a suitable focal plane but also an optimal 3D scanning area. Furthermore, standard methods are often slow and are only partially suitable for processes involving scanning a larger area. Furthermore, existing methods cannot distinguish between a large number of different objects. Thus, the current state of the art presents the problem of not clearly defining which of a large number of different objects should be focused on.
[0007] Another disadvantage of existing systems is that a large number of images must often be acquired to determine a focal plane. The light exposure of a sample during focusing can be significant. If such methods are applied to living cells or animals, the light exposure can lead to phototoxic effects, causing cell death or a change in cell behavior. The autofocus method can therefore have a significant impact on the object being examined and thus distort measurement values if too many images are acquired to determine a focal plane.
[0008] US 4 965 443 A discloses a focus detection device with neural network means.
[0009] US 9 516 237 B1 discloses focus-based shutter.
[0010] US 2010 / 0177189 A1 discloses a method and apparatus for controlling a microscope.
[0011] GUOJIN CHEN ET AL: "Research on image autofocus system based on neural network", SPIE - INTERNATIONAL SOCIETY FOR OPTICAL ENGINEERING. PROCEEDINGS, US, Vol. 5253, 2 September 2003 DOI:10.1117 / 12.521703, ISSN 0277-786X, page 461, discloses an autofocus method using a neural network.
[0012] It is the object of the present invention to provide an improved method and device for precisely determining a focus position.
[0013] This object is achieved by the subject matter of the independent claims. Embodiments are defined by the dependent claims.
[0014] Other aspects or embodiments fall within the scope of the claims if they have all the features specified in the independent claims.
[0015] Parts of the description and drawings relating to embodiments or aspects not covered by the claims are not presented as embodiments of the inventions, but as examples useful for understanding the invention.
[0016] The object of the present invention is therefore to determine a focus position efficiently and quickly. The present invention solves the problems and the object mentioned by a method and a device for determining a focus position by analyzing captured images using a trained model.
[0017] The method according to the invention for determining a focus position comprises recording at least one first image, wherein image data of the at least one recorded first image depend on at least one first focus position when recording the at least one first image, and determining a second focus position based on an analysis of the at least one recorded first image by means of a trained model.
[0018] The method and device according to the invention have the advantage that trained models, which can be based on neural networks, e.g., in the sense of deep learning, can be applied to acquired data to determine a focus position. Deep learning models can make predictions about the optimal focus that are as precise as or better than an expert in imaging techniques. The model can be applied to a wide range of samples, image acquisition conditions, and acquisition modalities and enables greater accuracy compared to state-of-the-art autofocus systems.
[0019] Furthermore, the method and device according to the invention have the advantage that, when using a single image as input value for the neural network, the focus position can be determined much faster and with less damage to the sample than is possible with previous image-based focusing systems, which require a large number of images to find a focus position. Furthermore, trained models can be used to focus on specific objects in image data that are in a context with a training of the trained model. This enables focusing on an imaged object regardless of the location of the object. A further advantage of the method and device according to the invention is that they can be used in different systems, such as confocal microscopes and digital cameras.
[0020] The method and device according to the invention can each be further improved by specific embodiments. Individual technical features of the embodiments of the invention described below can be combined with one another as desired and / or omitted, provided that the technical effect achieved by the omitted technical feature is not important.
[0021] In one embodiment, the method may comprise the step of capturing at least one second image with the determined second focus position, which is shifted, for example, along an optical axis relative to the first focus position. This may occur automatically, thus enabling rapid adaptation to specific conditions during a measurement and optimization of a focus position while a measurement is still in progress. The at least one first image may be captured at a lower resolution than the at least one second image. This allows the speed at which focused images can be captured to be increased.
[0022] In one embodiment, the at least one first image and the at least one second image can contain information that is in a context with training of the trained model. For example, if at least one object is in a context with training of the trained model, at least one of one or more objects depicted in one or more of the at least one recorded first image can be depicted more sharply in the at least one recorded second image than in the at least one recorded first image. This enables continuous refocusing and a sharp depiction of moving objects across multiple images at a time interval.
[0023] In one embodiment of the method, determining the second focus position may include determining a 3D scan area. This may be based on information with which the model was trained. For example, the model may have been trained on an object or sample of a specific size and may determine a focus position and an extent of the object based on the analysis of the at least one first image. Determining a scan area enables rapid execution of experiments in which an object is to be scanned along an axis.
[0024] In one embodiment, the method can comprise the step of adjusting a device, for example a microscope or a digital camera, to the second focus position. The device can comprise an optical system and a sensor by means of which the adjustment of the device to the second focus position can be carried out. One embodiment of the method according to the invention can be achieved in that adjusting the device to the second focus position comprises displacing the optical system relative to the sensor. Adjusting the device to the second focus position can additionally or alternatively also comprise displacing the optical system and the sensor relative to an object or along an optical axis. The adjustment of the device can be carried out by a focus drive.
[0025] In one or more embodiments of the method according to the invention, the method can comprise the step of acquiring data. This can comprise acquiring a third image. The image data of the acquired third image can depend on a user-defined focus position set by a user. The user-defined focus position can differ from the second focus position, and the acquired data can comprise the acquired third image, a representation of the deviation from the second focus position to the user-defined focus position, and / or a representation of the user-defined focus position. By acquiring a user-defined focus position, the system can conclude that the user did not agree with the focus position suggested by the trained model.This data can help to continuously improve a trained model by being used to fine-tune the trained model or to train new models. To obtain sufficient data for training a trained model, acquiring the data may involve acquiring an image stack. The image data of each image in the image stack may depend on a different focus position that has a known distance from the user-defined focus position. The distance may correspond to an integer multiple of an axial resolution of a lens used to the user-defined focus position along an optical axis. The third acquired image corresponding to the user-defined focus position may be marked as the target state for training that includes the acquired data.
[0026] In one embodiment of the method, the acquired data can include metadata. The metadata can include information on image acquisition modalities, lighting, a sample, an object, system parameters of an image acquisition device, image acquisition parameters, a context, and / or lens parameters. This allows models to be trained on specific conditions and thus deliver very precise predictions compared to models trained on data with high variability.The trained model for determining a focus position can, for example, be selected from a plurality of models, wherein the plurality of trained models can be classified by an area of application, each of the plurality of trained models has been trained in a specific way, the plurality of trained models is hierarchically organized, and / or individual trained models from the plurality of trained models are specialized for individual types of samples, experiments, measurements or device settings.
[0027] In one or more embodiments of the method according to the invention, the method can further comprise the step of adapting the trained model using the acquired data. This can be done by training a portion of the trained model using the acquired data, or by retraining the trained model using aggregated data, wherein the aggregated data originates from one or more sources and / or includes the acquired data. Thus, the prediction accuracy of trained models, such as deep learning models, can be continuously improved, as the method can adapt to new circumstances using feedback and is thus more flexible than prior art methods. Because the system improves by "learning" from user behavior, the accuracy of the autofocus becomes increasingly higher. Furthermore, the continuous improvement of the model opens up new fields of application.Furthermore, the method becomes increasingly robust to new, previously unseen samples. In contrast, previous approaches are either so general that they compromise accuracy or so specific that they cannot be applied to new sample types or imaging conditions. Thus, the inventive method can be applied to a broad and ever-expanding spectrum of samples, imaging conditions, and imaging modalities.
[0028] In one embodiment of the device according to the invention, the device for determining a focus position can comprise one or more processors and one or more computer-readable storage media. Computer-executable instructions are stored on the storage media, which, when executed by the one or more processors, cause the above-described inventive method to be carried out.
[0029] In one embodiment, the device according to the invention can be part of an image recording system for adjusting a focus position. The image recording system can comprise a digital camera (for example, a compact camera or a camera for use in manufacturing, with robots, or autonomously operating machines), a computer with a camera, which can also be portable (for example, a notebook, a cell phone, a smartphone, or a tablet), or a microscope or microscope system. Alternatively, the device can be spatially separated from the image recording system and connected to the image recording system via a network. In one embodiment, the image recording system can comprise at least one sensor, an optical system for imaging one or more objects onto one of the at least one sensor, and at least one actuator.The at least one actuator, which may be, for example, a Z-drive of a tripod, a Z-galvanometer, a piezo focus on a lens of the optical system, a direct drive, a linear motor, a stepper motor, an ultrasonic motor, a ring motor, a micromotor, and a piezo stage, can adjust the image acquisition system to the second focus position. In embodiments, the at least one actuator can be part of a focusable lens, wherein the focusable lens is based on focus technologies without a piezo drive. The image acquisition system can comprise means configured to download trained models from a cloud via a network and / or to upload acquired data to the cloud. This allows models to be continuously updated and trained models to be made available quickly.A further embodiment of the image recording system according to the invention, which can be combined with the previous ones, can comprise a user interface which is configured so that a user-defined focus position can be set and / or input parameters can be recorded by means of the user interface.
[0030] The present invention will be described in more detail below with reference to exemplary drawings. The drawings show examples of advantageous embodiments of the invention.
[0031] They show: Figure 1 a schematic representation of a system according to the invention for determining a focus position according to one embodiment, Figure 2 a schematic representation of a system according to the invention for determining a focus position according to one embodiment, Figure 3a schematic representation of a method according to the invention for determining a focus position according to one embodiment, Figure 4 a schematic representation of a trained model for determining a focus position according to one embodiment, and Figure 5 a schematic flow diagram of a method according to the invention.
[0032] Figure 1shows a system 100 for determining one or more focus positions for imaging methods. The system 100 includes a device 110 for capturing images. The device 110 for capturing images may include a camera, such as a digital camera, a camera built into a portable computer, such as a smartphone camera, or a microscope, such as a wide-field microscope, a confocal microscope, or a light-sheet microscope. The device 110 for capturing images may include one or more sensors 112 and one or more actuators 114. Furthermore, the device 110 may include an optical system consisting of at least one of the following optical components: one or more lenses, one or more mirrors, one or more apertures, and one or more prisms, wherein the lenses may include various lens types. The optical system may include an objective lens.In addition, the device 110 may include one or more light sources for illumination and / or fluorescence excitation.
[0033] The one or more sensors 112 and one or more actuators 114 can receive and / or transmit data. One of the sensors 112 can comprise an image sensor, for example, a silicon sensor such as a CCD ("charge-coupled device") sensor or a CMOS ("complementary metal-oxide-semiconductor") sensor, for recording image data, wherein an object can be imaged on the image sensor using the optical system. The one or more sensors can capture image data, preferably in digital form, and metadata. Image data and / or metadata can be used, for example, to predict a focus position.Metadata may include data related to the acquisition of one or more images, for example, image acquisition modalities such as bright-field illumination, epifluorescence, differential interference contrast, or phase contrast; image acquisition parameters such as the intensity of the light source(s), image sensor gain, sampling rate; and lens parameters such as the axial resolution of the lens. Image data and metadata may be processed in a device 130, such as a workstation, an embedded computer, or a microcomputer, in an image acquisition device 110.
[0034] Even if Figure 1While the schematic representation delineates areas 110, 120, and 130, these areas may be part of a single device. Alternatively, in other embodiments, they may also comprise two or three devices that are spatially separated from one another and connected to one another via a network. In one embodiment, the device 130 may include one or more processors 132, a volatile data memory 134, a neural network or trained model 136, a permanent data memory 138, and a network connection 140.
[0035] The one or more processors 132 may include computing accelerators, such as graphical processing units (GPUs), TensorFlow processing units (TPUs), application-specific integrated circuits (ASICs) or field-programmable gated arrays (FPGAs) specialized for machine learning (ML) and / or deep learning (DL), or at least one central processing unit (CPU). Using the one or more processors 132, trained models can be applied to determine focus. The use of trained or predictive models, which are used in image acquisition devices to analyze the acquired images, helps to make predictions (inference). Inference involves predicting a dependent variable (called y) based on an independent variable (called X) using a previously trained neural network or model.An application machine or device 130 that is not (primarily) used to train neural networks can thus be configured to make predictions based on trained networks. A transfer of a trained neural network to the application machine or device 130 can be carried out in such a way that the application machine or device 130 gains additional "intelligence" through this transfer. This can enable the application machine or device 130 to solve a desired task independently. In particular, inference, which requires orders of magnitude less computing power than training, i.e., the development of a model, also works on conventional CPUs. With the help of the one or more processors 132, models or parts of models can be trained using artificial intelligence (AI) in embodiments. The trained models themselves can be executed by the one or more processors.This results in a cognitively enhanced device 130. Cognitively enhanced means that the device can be enabled to semantically recognize and process image content or other data by using neural networks (or deep learning models) or other machine learning methods.
[0036] Data can be uploaded to or downloaded from a cloud 150 via the network connection 140. The data can include image data, metadata, trained models, their components, or hidden representations of data. Thus, the computer 130 can load new, improved models from the cloud 150 via the network connection 140. To improve models, new data can be loaded into the cloud 150 automatically or semi-automatically. In one embodiment, an overwriting of a focus position predicted by the neural network is detected. By overwriting, new value pairs, such as image data and a focus position corresponding to a desired focus position or target state, can be generated for training or fine-tuning the neural network 136 or a neural network stored in the cloud.The data generated in this way continuously expands the available training dataset for training models in the cloud 150. This creates a feedback loop between user and manufacturer, through which neural networks can be continuously improved for specific applications. The cloud 150 can include an AI component 152, which can be used to train models and / or neural networks.
[0037] Furthermore, the system 100 comprises a user interface 120, which can be part of the device 110, the computer 130, or another device, such as a control computer or a workstation. The user interface 120 comprises one or more switching elements 122, such as a focus control, and a software-implemented user interface 124, via which additional parameters for the neural network 136 can be entered. This can include information about a sample type and experimental parameters, such as staining and culture conditions. In one embodiment, the user can overwrite a result from a neural network and thus contribute to its fine-tuning. For example, the user can enter the correct or desired focus position for an application.
[0038] The neural network 136 can make a prediction based on input values. The input values can include image and / or metadata. For example, the neural network 136 can make a prediction about a correct focus position based on input values. The prediction can be used to control at least one of the one or more actuators 114 accordingly. Depending on the prediction or the analysis of the input values by the neural network, a focus drive can be adjusted. The one or more actuators can include a Z-drive of the tripod, a Z-galvanometer, a piezo focus on the lens ("PiFoc"), a direct drive, a linear motor, a stepper motor, an ultrasonic motor, a ring motor, a micromotor, or a piezo stage.
[0039] The one or more actuators 114 can be set to a predicted plane or focus position. The process for determining a plane or optimal focus position can be repeated any number of times. One of the sensors 112 can capture another image, and the new image can be analyzed using the neural network 136. Based on the result, an actuator can set a new position, or it can be determined that the optimal position for a focus has been reached. The number of repetitions for determining an optimal focus position can be determined by a termination criterion. For example, the reversal of the direction of the difference vector between the current and predicted focus position can be used as the termination criterion.
[0040] Figure 2shows, in one embodiment, communication of an image acquisition system 210, for example a microscope, with AI-enabled devices 220 and 230. A single microscope can itself comprise hardware acceleration and / or a microcomputer, which enable it to be AI-enabled and enable the execution of trained models (e.g., neural networks). Trained models can comprise deep learning result networks. These neural networks can represent results learned through at least one deep learning process and / or at least one deep learning method. These neural networks condense collected knowledge into a specific task ensemble in a suitable manner through automated learning, such that a specific task can henceforth be automated and performed with the highest quality.
[0041] The microscope may comprise one or more components 260 and 270. Various components of the microscope, such as actuators 260 and sensors 270, may themselves be integrated circuits and comprise microcomputers or FPGAs. The microscope 210 is configured to communicate with an embedded system 220 and with a control computer 230. In one example, the microscope 210 communicates simultaneously or in parallel with one or more embedded systems 220 having hardware-accelerated AI (artificial intelligence) and with its control computer 230 via bidirectional communication links 240. Data (such as images, device parameters, experimental parameters, biological data) and models, their components, or hidden representations of data can be exchanged via the bidirectional communication links 240, e.g., a deep learning bus.The models can be modified during the course of an experiment (through training or adapting parts of a model). Furthermore, new models can be loaded onto a microscope or device from a cloud 250 and deployed. This can happen based on the recognition and interpretation of data generated by a model itself or by capturing user input.
[0042] For the continuous improvement of models, one embodiment involves collecting data from as many users as possible. Users can specify their preferences regarding which data may be collected and processed anonymously. Furthermore, there may be an option to evaluate model predictions at the appropriate point in the user interface for the experiment. In one example, the system could determine a focus position. The user has the option to overwrite this value. The focus position overwritten by the user can represent a new data point for fine-tuning a model, which is then available for model training. The user thus has the advantage of being able to continually download and / or run improved models. In turn, this opens up the possibility for the manufacturer to continually improve its offering, as new models can be developed based on data from many users.
[0043] In Figure 2 Furthermore, the communication between image acquisition system 210 and embedded computer 220 or system computer 230 is shown, which can communicate with each other as well as with a cloud 250 or a server. Image acquisition system 210 and its attached components can communicate with each other and with workstations via a deep learning bus system. Specialized hardware and / or a TCP / IP network connection or an equivalent can be used for this purpose. The deep learning bus system can include the following features: Networking of all subsystems—i.e., components of the image acquisition system, sensors, and actuators—with each other and with suitable models. These subsystems can be intelligent, i.e., possess neural networks or machine intelligence themselves, or non-intelligent. Networking all subsystems and modules of image acquisition systems with one or more image acquisition systems results in a hierarchical structure with domains and subdomains. All domains and subdomains, as well as the associated systems and subsystems, can be centrally recorded and searchable so that a model manager can distribute models across them. For communication in time-critical applications, specialized hardware for the bus system can be used. Alternatively, a network connection based on TCP / IP or a suitable web standard can be used.The bus system preferably manages the following data or part of it: ID for each component (actuators, sensors, microscopes, microscope systems, computing resources, workgroups, institutions); rights management with author, institution, read / write permissions of the executing machine, desired payment system; metadata from experiments and models; image data; models and their architecture with learned parameters, activations, and hidden representations; required interfaces; required runtime environment with environment variables, libraries, etc.; all other data as required by the model manager and rights management. The multitude of different AI-enabled components creates a hierarchical structure with domains and subdomains. All components are recorded and findable in a directory.Rights management takes place at every level (i.e., component attached to the microscope, microscope / system, workgroup, computer resources, institution). This creates a logical and hierarchical functional structure that facilitates collaboration among all stakeholders.
[0044] According to one embodiment, the method and device according to the invention for determining a focus position can be used in all microscopic imaging modalities, in particular wide-field microscopy (with and without fluorescence), confocal microscopy, and light-sheet microscopy. All focusing mechanisms used in microscopy and photography can be considered as focus drives, in particular the Z-drive of the stand, Z-galvanometer stages, but also piezo stages or piezo focusing units on the objective ("PiFoc"). Alternatively, the focus drive can comprise an actuator that is part of a focusable objective, wherein the focusable objective is based on focus technologies without piezo drive. Possible illumination modalities include transmitted light illumination, phase contrast, differential interference contrast, and epifluorescence.Different illumination modalities may require training a separate neural network for each. In this case, an automatic pre-selection of the appropriate model would take place. System parameters (so-called image metadata) can be used for this purpose. Alternatively, another neural network is used for image classification. This uses the image stack used for focusing and makes a prediction about the illumination modality, e.g., transmitted light, phase contrast, etc.
[0045] Figure 3 shows a schematic representation of the functionality of determining a focus position according to one embodiment. The method according to the invention can be used to locate a sample plane at the beginning of an experiment.
[0046] In a first step, a first image 310 and a second image 320 can be recorded. The images 310 and 320 can be recorded at a distance equal to the axial resolution of the lens used. The first image 310 and / or the second image 320 are then analyzed using a trained model, and a prediction 340 is made about a position. The prediction 340 (Δz, in the case shown Δz > 0) consists of the difference between the last recorded and / or analyzed image 320 and the focus position 330 predicted by the model. In one embodiment, this position can be moved to using a focus drive. From there, another image 350 can be recorded at a distance equal to the axial resolution of the lens used. The further image 350 can then be analyzed again using the trained model in order to make another prediction 360 for the new position moved to.If the predicted delta (Δz ≤ 0) of the further prediction 360 is negative, as in . Figure 3 can be seen, the correct focus position 330 has already been found. Otherwise, the procedure can be continued until the correct position is found.
[0047] The correct position can be verified in both directions along the axial axis. For example, an image can be captured in one direction and the other along the optical axis around a predicted position, e.g., at a distance equal to the axial resolution of the lens used or at a predefined distance. These images can then be analyzed using the trained model. Once the correct position has been found, a prediction from one image should yield a positive delta, while a prediction from the other image from the opposite direction should yield a negative delta.
[0048] In an embodiment that does not fall within the scope of the claims but is helpful for understanding the invention, the device according to the invention operates in two different modes. In Mode I, a plane in which a sample is visible or in focus is found without user intervention. In Mode II, the user can fine-tune the focus for their specific application to achieve even more precise results or an area with the highest information content in the acquired images.
[0049] In addition to basic training, which can take place in a cloud or during development, the model can learn from experience. The user can determine what they believe to be the optimal focus and communicate this to the system. In this case, the image stack used for focusing, along with metadata (information on image acquisition modalities, lighting, etc.), as well as the focus position determined by the user, are transmitted. Both represent new value pairs, X i and yi, that can be added to the training dataset. In another short training round ("fine-tuning"), the model can improve its prediction accuracy and learn sample types and lighting modalities. If this is done by the manufacturer or in the cloud, the user can download a new model. This can happen automatically. In one embodiment, the user or the device according to the invention can perform the fine-tuning training itself.The manufacturer can use a customer data management system that stores user-specific data for fine-tuning. This also allows the manufacturer to offer customized neural networks. The manufacturer can use email, social media, or push notifications to inform the user about the availability of newly trained neural networks. The manufacturer can also offer these to other customers via an app store.
[0050] In one embodiment, which does not fall within the scope of the claims but is helpful for understanding the invention, a customer's consent can be obtained to store their image data. Instead of the original image data, a transformation of the data can also be stored. For this purpose, the image data is processed by a portion of the neural network and transferred into a "feature space." This typically has a lower dimensionality than the original data and can be viewed as a form of data compression. The original image can be reconstructed by knowing the exact version of the neural network and its parameters.
[0051] In one embodiment, a focus position can be determined based on an analysis of only a single image 320. This image 320 can be captured, and a prediction can be made based on an analysis of the captured image 320 using a trained model. This enables a jump prediction to an optimal focus position 330. The system "jumps" directly to the predicted focus position or a position along the optical axis (Z-position) without focusing in the conventional sense. The output image used for this purpose can be blurred, i.e., out of focus. In the case of any confocal image, which is by definition in "focus," a plane along the optical axis, for example, a z-plane, can be predicted with the desired information.The image used does not have to contain any of the information sought, but the image must fit a context that is part of the learning model used.
[0052] Using a single image also opens up expanded applications. These include refocusing in time stacks or applications in focus maps. A focus map consists of a set of spatial coordinates in the lateral (stage position) and axial (Z-drive) directions. Based on one or more focus positions in a focus map, neighboring focus positions in the focus map can be predicted. This means that a smaller number of support points in a focus map is sufficient to achieve the same accuracy as with previous methods. The focus map can also change dynamically and become increasingly accurate the more support points are available after an acquisition progress. In a next step, this can be used to automatically determine a curved surface in 3D space. For large samples, this can be used to capture the informative spatial elements and thus saves acquisition time and storage space.
[0053] The trained or learned model used can be based on deep learning methods to determine the desired jump prediction of the focus position or a Z-plane in focus (for non-confocal imaging techniques), or a plane containing the desired information (for confocal imaging techniques) based on one or a few images. Conventional image-based methods always require the acquisition and analysis of a complete image stack in the axial direction. This is time-consuming and computationally expensive and can damage samples during certain measurements.
[0054] The prediction of the optimal focus position can be performed using a deep convolutional neural network (CNN), which is a form of deep learning. The model is used in two ways: 1) It "learns" to predict optimal focus positions using a training dataset, and 2) It predicts focus positions using an existing stack of images or a single image. The former is called supervised training, and the latter is called inference. Training can take place in a cloud or on a computer designed for training neural networks. The development or training of the neural networks can be done by a third party or locally.Training can also be performed automatically using conventional autofocus methods with one or more suitable calibration samples, with the calibration samples being adapted to the intended use, i.e., the sample type to be examined later. The inference can be performed by the user. In one embodiment, however, the user can also improve a model or neural network through repeated, less extensive training. This is referred to as "fine-tuning."
[0055] In Figure 4 A schematic representation of a trained model is shown. Data 400 input to the model may include image data and / or metadata. The trained model can make a prediction 430 about the position of a focus or the distance to the correct focus position from the input data.
[0056] The training of a model consists of finding internal model parameters (e.g., weights W and / or thresholds B) for existing input values X, such as image data or metadata, which result in an optimal prediction of the output values ŷ i. The target output values yi are known for training. In this application, X, W, and B are viewed as tensors, while ŷ i can be a scalar. For example, ŷ i can be the focus position of the Z-drive or a value that is related to a focus position and uniquely identifies it (e.g., a relative distance Δz to the current image position). In one embodiment, a prediction, as in Figure 3 shown, based on a Z-stack of acquired images. The predicted focus position is passed to the Z-drive of a microscope, and the microscope is adjusted accordingly.
[0057] The model consists of two functionally distinct parts. The first part 410 of the model computes a pyramidally hierarchically arranged cascade of convolutions, each followed by nonlinearity and dimensionality reduction. The parameters for the convolution are learned during training of the trained model. Thus, the model "learns" a hidden representation of the data during training, which can be understood as data compression as well as the extraction of image features.
[0058] The second part 420 of the model orders and weights the extracted image features. This enables the most accurate prediction possible, such as a focus position. Training the image feature extraction 410 is time- and computationally intensive, in contrast to training the weighting of the image features 420. This last part of the model can be trained locally with little effort (fine-tuning). Fine-tuning can be triggered by user input and / or automatically. For example, a user can overwrite a focus position predicted by a trained model. If the user overwrites a focus position, it can be assumed that the user is not satisfied with the prediction or wants to focus on a different area or object.Another example is when a conventional or a second autofocus system automatically finds a focus position that deviates from the focus position predicted by the system. In this case, the second autofocus system can overwrite the predicted focus position. Fine-tuning in these and similar cases occurs automatically and improves the prediction accuracy in this application. Through fine-tuning, the individual user can quickly and locally improve the accuracy of predictions of a trained model. The user can make their data available. Data such as the overwritten focus position and the associated image data and metadata can be uploaded to a cloud via a network connection. In the cloud, models can be trained and / or improved using the data. The image feature extraction part 410 in particular benefits from this.The trained models in the cloud can be improved by a variety of users. In the cloud, data from different sources or users can be aggregated and used for training.
[0059] Models can be trained in the cloud for specific applications. When image data with high variability is acquired, models must be trained very generally. This can lead to errors or inaccuracies. An example of image data with high variability is the use of microscopes in basic research or clinical applications, where image data and other data types differ greatly. Here, it may be useful to perform an upstream classification step with a "master model" and automatically select the correct model for the corresponding data domain. This principle of using a hierarchical ensemble of models is not limited to one data domain but can encompass different data types and domains.Likewise, several hierarchically organized model ensembles can be cascaded to organize data domains according to multiple dimensions and to perform automatic inference despite different applications and variability in the data.
[0060] Figure 5shows a schematic representation of a flowchart of a method 500 according to the invention for determining a focus position. The method 500 comprises a first step 510, in which at least one first image is recorded. Image data of the at least one recorded first image depend on at least one first focus position when recording the at least one first image. The at least one first image can be recorded using a device that comprises an optical system and a sensor. In exemplary embodiments, the device for recording images comprises a microscope or microscope system, a digital camera, a smartphone, a tablet, or a computer with a camera. In particular, the method according to the invention can be used in the field of wide-field microscopy (with and without fluorescence), confocal microscopy, and light sheet microscopy.
[0061] In a second step 520, a second focus position is determined. This involves analyzing the at least one captured first image using a trained model or neural network. The second focus position can then be determined based on the analysis or a prediction from the analysis.
[0062] In a third step 530, a device can be adjusted to the second focus position. This can be done by means of a focus drive on the device for recording images. Depending on the device, for example, an optical system can be moved relative to a sensor along an optical axis in order to change a focus position and / or to adjust the device to the specific focus position from step 520. Alternatively, or additionally, the optical system and the sensor can also be moved together at a constant distance along an optical axis, so that the focus point shifts relative to an observed object. For example, a microscope can move a sample stage in a direction that corresponds, for example, to the optical axis of the microscope.In this way, a desired plane of a sample can be brought into focus or the focus plane can be moved into a desired plane of a sample or object.
[0063] A further step 540 may include capturing at least one second image with the second focus position. Image data of the at least one second image may depict one or more objects more sharply than image data of the at least one first image. The method 500 may focus on objects that are in context with the training or training data of a neural network or trained model. As a result, the method 500 is not limited to specific image areas for focusing. Instead of focusing systems that focus on close or large objects, the method 500 may identify objects and focus on them. This may occur regardless of the position of an object or the size of an object in the image. In one embodiment, determining the focus position is based on information from the image data that is in context with the training or training data of a neural network or trained model.
[0064] The focus is the point in an optical imaging system where all incident rays parallel to the optical axis intersect. An image captured in the focal plane of an optical imaging system can exhibit the highest-contrast structures. A region of highest information can coincide with the focal plane, but this is not required. The region of highest information is the area desired by the viewer of an image, in which the viewer perceives the (subjective) highest information of "something." "Something" here refers to a specific visual structure; for example, the viewer of an image can look for "small green circles with a yellow dot" and then perceive the area of an image containing this desired structure as the area of highest information.
[0065] Several forms of feedback can improve a method and device for determining a focus position. This can occur during a measurement or during image acquisition. Two forms of feedback can be distinguished. Both forms of feedback involve an adaptation of the trained model. 1) According to the invention, the first form is a feedback into the model: Since the models can be continuously fine-tuned and their prediction accuracy is thereby continuously improved, a feedback is created that can change and improve the model, which continuously evaluates the data, even during the runtime of an experiment or a measurement.
[0066] With feedback, an ensemble of methods or procedures can be designed in such a way that results obtained through deep learning feed back into the microscopy system, microscopy subsystems, or other image acquisition devices, creating a kind of feedback loop. Through the feedback, the system is asymptotically transitioned to an optimal and stable state or adapts appropriately (system settings) to capture specific objects more optimally.
[0067] In an embodiment that does not fall within the scope of the claims but is helpful for understanding the invention, data is collected to optimize an autofocus method, consisting of steps 510 and 520. Collecting this data can trigger the process for optimizing the autofocus method. Collecting the data can include capturing a third image, wherein the third image can include image data that depends on a user-defined focus position. The user-defined focus position can differ from the second focus position, which can indicate that a user is not satisfied with the focus position suggested by the trained model and / or that optimization of the autofocus method is necessary. For example, the user can focus on an area in the image that contains the most or most important information for them.The optimization of the autofocus process can thus be achieved by optimizing the trained model.
[0068] The acquired data may include the acquired third image, a representation of the deviation of the second focus position from the user-defined focus position, and / or a representation of the user-defined focus position. Furthermore, the third image may be marked. Since the user-defined focus position corresponds to a desired focus position, the acquired third image corresponding to the user-defined focus position may be marked as a target state for training based on the acquired data.
[0069] In one embodiment, a user can be actively asked for additional data to optimize a model. The user can then mark this data. Marking means an evaluation by the user such that they inform the system which Z position they consider "in focus" or "correct," or which focus position produces sharp images of desired objects. In this way, the additional data can be generated by operating the image acquisition device and used as training data. Images that the user evaluates as "in focus" are given the label "in focus." A series of images can then be captured around this user-defined focus point. Depending on the acquired data and the user-defined focus position, an image stack can be captured or acquired.The image data of each image in the image stack may depend on a focus position that is a known distance from the user-defined focus position.
[0070] Based on the acquired data, the model used can be fine-tuned. This may involve training a portion of the trained model. For example, only the portion 420 from Figure 4 This can be done locally by reselecting internal model parameters. Thus, by optimizing a trained model, the entire autofocus process can be optimized.
[0071] Additionally, metadata can also be captured to classify the acquired data. This allows defining a domain of application for the trained model. The metadata can include information about image acquisition modalities, such as bright-field illumination, epifluorescence, differential interference contrast, or phase contrast; information about an illumination; information about a sample; information about an object; information about system parameters of an image acquisition device; information about image acquisition parameters, such as the intensity of the light source(s), gain at the photosensor, sampling rate; context information; and information about lens parameters, such as the axial resolution of the lens. Context information can be related to the acquired data.For example, contextual information may include keywords or explanations about an experiment related to the collected data and be used to classify the data.
[0072] 2) The second form of feedback, which does not fall within the scope of the claims but is helpful for understanding the invention, is based on image recognition and / or evaluation of acquired data. Models can be exchanged and reloaded at runtime to support other object types, other stainings, and generally other applications. This can even happen during the runtime of an experiment or measurement, making the microscope highly dynamic and adaptable. This also includes the case where the sample under investigation changes substantially during the experiment, requiring an adjustment of the model or its complete replacement during the runtime of the experiment. For example, a "master" or world model can classify the application area and automatically select suitable models, which are then executed.
[0073] The use and continuous improvement of predictive models used in measurements taken by microscopes or other image acquisition devices, where they make predictions (inference) and can be fine-tuned if necessary, advantageously by training only a few nodes in the neural network, optimizes methods for determining focus positions and expands the range of applications of models in image acquisition devices, such as microscopes. Applications of inference using these models are diverse and include, for example, the automation of microscopes or experimental procedures, in whole or in part, such as locating objects and determining or setting an optimal focus position. Reference symbols:
[0074] 100System 110Image acquisition device 112One or more sensors 114One or more actuators 120User interface 122Controller 124Software user interface 130Computer 132One or more processors 134One or more volatile storage media 136Neural network 138One or more storage media 140Network connection 150Cloud 152AI component 210Microscope 220Embedded computer 230System computer 240Bidirectional communication links 250Cloud 260Actuator 270Sensor 300Image stack 310, 320, 350Acquired images 330Correct focus position 340, 360Prediction 400Model input 410, 420Parts of the neural network 430Model output 500Procedure 510 - 540 process steps
Claims
1. A method (500) for determining a focus position, with the steps: - recording (510) at least one first image (310, 320), wherein image data of the at least one recorded first image (310, 320) depend on at least one first focus position when recording the at least one first image; - determining (520) a second focus position (340) based on an analysis of the at least one recorded first image (310, 320) by means of a trained model (136); and - recording (540) at least one second image (330) with the second focus position (340), wherein the at least one first image (310, 320) and the at least one second image (330) contain information that is in a context with a training of the trained model (136), wherein at least one object of one or more objects, which are depicted in one or more of the at least one recorded first image, is depicted more sharply in the at least one recorded second image than in the at least one recorded first image, wherein the at least one object is in a context with a training of the trained model, characterized in that the trained model is adapted by feedback that changes the model, which continuously evaluates the image data, during the runtime of an experiment or a measurement.
2. The method according to claim 1, characterized in that the at least one first image (310, 320) is recorded with a lower resolution than the at least one second image (330).
3. The method according to any one of claims 1 and 2, characterized in that the at least one recorded first image (310, 320) depicts one or more objects, wherein the at least one recorded second image (330) depicts at least one of the one or more objects more sharply than the at least one recorded first image (310, 320).
4. The method according to any one of claims 1 to 3, characterized in that the second focus position is determined during a measurement, and / or the at least one first focus position is shifted along an optical axis to the second focus position.
5. The method according to any one of claims 1 to 4, characterized in that the at least one first focus position corresponds to at least one first focal plane and the second focus position corresponds to a second focal plane, wherein the second focal plane approximately coincides with an object plane.
6. The method according to any one of claims 1 to 5, characterized in that the method further comprises the following step: - adjusting (530) a device to the second focus position, wherein the device comprises an optical system and a sensor.
7. The method according to claim 6, characterized in that adjusting the device (110; 210) to the second focus position comprises shifting the optical system relative to the sensor, and / or adjusting the device to the second focus position comprises moving the optical system and the sensor (112) relative to an object.
8. The method according to any one of claims 1 to 7, characterized in that the method further comprises the step of collecting data, optionally, wherein an application range for the trained model is defined by metadata for classifying the collected data, wherein the metadata comprise contextual information that comprises keywords or explanations about an experiment related to the collected data and is used to classify the collected data.
9. The method according to any one of claims 1 to 8, characterized in that the trained model (136) is based on one or more neural networks or one or more deep learning result networks and / or was selected from a variety of models, wherein the plurality of trained models is classified by an application range, each of the plurality of trained models has been trained in a specific way, the plurality of trained models is hierarchically organized, and / or individual trained models from the plurality of trained models are specialized for individual types of samples, experiments, measurements, or instrument settings.
10. A device (130; 210, 220, 230) for determining a focus position, comprising: one or more processors (132); one or more computer-readable storage media (138) having stored thereon computer-executable instructions that, when executed by the one or more processors (132), cause that at least one sensor (112) captures at least one first image (310, 320), wherein image data of the at least one captured first image (310, 320) depend on at least one first focus position during the capture of the at least one first image; a second focus position is determined, wherein the second focus position is determined based on an analysis of the at least one captured first image (310, 320) by means of a trained model (136); and at least one second image (330) is recorded with the second focus position (340), wherein the at least one first image (310, 320) and the at least one second image (330) contain information that is in a context with a training of the trained model (136), wherein at least one object of one or more objects, which are depicted in one or more of the at least one recorded first image is depicted more sharply in the at least one recorded second image than in the at least one recorded first image, wherein the at least one object is in a context with a training of the trained model, characterized in that the trained model is adapted by feedback that changes the model, which continuously evaluates the image data, during the runtime of an experiment or a measurement.
11. An image recording system (100) for setting a focus position, comprising: the device (130; 210, 220, 230) according to claim 10; the at least one sensor (112; 270); an optical system configured to image one or more objects onto one of the at least one sensor (112; 270); and at least one actuator (114; 260), wherein the at least one actuator (114; 260) is configured to adjust the image recording system (100) to the second focus position.
12. The image recording system according to claim 11, characterized in that the at least one actuator (114; 260) comprises at least one of the following: a Z-drive of a tripod, a Z-galvanometer, a piezo-focus on an objective of the optical system, a direct drive, a linear motor, a stepper motor, an ultrasonic motor, a ring motor, a micromotor, and a piezo stage, and / or that the at least one actuator (114; 260) is part of a focusable lens, wherein the focusable lens is based on focus technologies without piezo drive.
13. The image recording system according to any one of claims 11 to 12, characterized in that the image recording system comprises a digital camera, a portable computer with a camera, or a microscope (210) or microscope system, and / or the device (130; 210, 220, 230) is part of a microscope (210) or microscope system.