Microscope automatic focusing system based on reinforcement learning
By developing a microscope autofocus system based on deep reinforcement learning, which utilizes image gradient features and a deep Q-network decision module, a fast, robust, and generalizable autofocus system for microscopes is achieved, solving the problems of low efficiency and reliance on labeled data in traditional methods.
Patent Information
- Application Number
- CN202511204061.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2025-11-14
AI Technical Summary
Traditional microscope autofocusing methods are inefficient and have poor robustness to complex samples, while supervised learning methods that rely on large amounts of labeled data have limited generalization ability in new scenarios.
A microscope autofocus system based on deep reinforcement learning is adopted. Through image acquisition, gradient feature extraction and deep Q network decision module, unsupervised autofocus is achieved by using a preset reward function of gradient energy, and real-time adjustment is performed by edge computing unit.
It achieves fast adaptive autofocus, reduces sampling times and time overhead, improves robustness and generalization ability for complex samples, and reduces dependence on labeled data.
Smart Images

Figure CN120949434A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of microscopic imaging and intelligent control technology, specifically to a microscope autofocus system based on reinforcement learning. Background Technology
[0002] Traditional microscope autofocus methods are mainly divided into two categories: active and passive. Active autofocus relies on additional hardware (such as infrared or ultrasonic sensors) to measure the distance to the subject and adjust the lens position; passive autofocus adjusts the focal length by analyzing acquired image data (such as contrast and sharpness). Common passive autofocus evaluation functions include gradient-based, Laplacian operator, variance, or information entropy-based methods. These focusing operators calculate sharpness indices by analyzing image brightness, edge, or texture features. For example, the Tenengrad operator and gradient metrics such as variance are widely used in real-time focusing due to their high computational efficiency. Then, combining these sharpness evaluations, traditional methods typically employ strategies such as scanning or pyramid search to find peak points in the depth direction. In recent years, supervised methods based on deep learning have also emerged, mainly using offline training of neural networks to quickly estimate the focus position.
[0003] However, the above methods have limitations: the focus evaluation function needs to be manually designed based on the characteristics of the samples, and multiple image acquisitions are required to traverse the focal plane, which is inefficient; traditional search strategies are prone to getting trapped in local optima, are sensitive to non-uniform samples or changes in illumination, and have poor adaptability; while supervised deep learning methods require a large number of labeled samples for training, are difficult to adapt online, and have limited generalization ability in new scenarios. Summary of the Invention
[0004] The purpose of this invention is to address the shortcomings of the prior art by providing a microscope autofocus system based on reinforcement learning to solve the problem of microscope autofocus.
[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows: This invention provides a microscope autofocus system based on reinforcement learning, the system comprising: The image acquisition module is used to acquire images of the current field of view through the microscope camera; The feature extraction module is used to calculate gradient features of the current field-of-view image, select the normalized gradient mean as the state vector, and input the state vector into the deep Q network decision module. The control module is used to control the stepper motor or piezoelectric platform to finely adjust the focal length along the Z-axis according to the adjustment command. This indicates the adjustment command to be executed based on the current focus position: ,in It focuses on step size; The image sharpness evaluation module is used to guide the deep Q-network decision module to adjust towards better image sharpness based on the preset reward function of gradient energy; The deep Q-network decision module utilizes a pre-trained deep Q-network decision model deployed to the edge computing unit of the microscope. Based on the input state vector and the reward function from the image sharpness evaluation module, it generates a focus adjustment command and sends it to the control module to achieve automatic focusing of the microscope.
[0006] Optionally, the image size of the current field of view is 672×672 pixels.
[0007] Optionally, the focusing step size is 1 micrometer.
[0008] Optionally, the gradient feature calculation in the feature extraction module specifically includes: using the Sobel operator Calculate the gradient magnitude map of the image: , in, Used to characterize image edge strength, and to measure image sharpness, I x I represents the gradient value of the image along the x-direction. y This represents the gradient value of the image along the y-direction.
[0009] Optionally, a preset reward function is provided. : .
[0010] Optionally, a preset reward function is provided. : ,in, This is the focusing penalty coefficient.
[0011] Optionally, the structure of the deep Q-network decision model includes: an input layer, which receives the environmental state perceived by the current system. This environmental state is constructed using image sharpness-related features and is formally represented as a state vector with a fixed length of 128 dimensions; and a hidden layer, consisting of two fully connected layers, each with 256 units, using the ReLU function as the activation function. The first hidden layer takes the input 128-dimensional state vector as input. The projection is onto a 256-dimensional high-dimensional feature space, and the second hidden layer further extracts abstract features; the output layer: the output is the action value function, the dimension is the number of actions, and the number of actions is 3.
[0012] Optionally, the training process of the deep Q-network decision model includes the following steps: Initialize the parameters of the Q network and the target network; implement - A greedy strategy samples actions and interacts to obtain a state-action-reward-new state quadruple; Construct an experience replay pool and train it using mini-batch samples, which consist of a preset N samples; The Adam optimizer is used to backpropagate and update the Q network parameters, copying the Q network parameters to the target network every K steps. Once the loss function converges and the focusing performance stabilizes, the trained Q-network model is deployed to the edge computing unit of the microscope.
[0013] The beneficial effects of this invention include: The microscope autofocus system based on reinforcement learning provided by this invention includes: an image acquisition module for acquiring the current field-of-view image through a microscope camera; a feature extraction module for calculating gradient features on the current field-of-view image, selecting the normalized gradient mean as a state vector, which is input to a deep Q-network decision module; and a control module for controlling a stepper motor or piezoelectric platform to finely adjust the focal length along the Z-axis according to adjustment commands. This indicates the adjustment command to be executed based on the current focus position: ,in The system comprises a focusing step size; an image sharpness evaluation module, which guides the deep Q-network decision module to adjust towards better image sharpness based on a preset reward function of gradient energy; and a deep Q-network decision module, which utilizes a pre-trained deep Q-network decision model deployed to the edge computing unit of the microscope, generates focus adjustment commands based on the input state vector and the reward function from the image sharpness evaluation module, and sends them to the control module to achieve automatic focusing of the microscope. This invention utilizes deep reinforcement learning to train intelligent strategies, and within an unsupervised learning framework, rapidly and adaptively finds the optimal focus through closed-loop interaction with the microscope system, while reducing sampling times and time overhead, and achieving good generalization to various sample environments. Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 The diagram shows the overall structure of the microscope autofocus system based on reinforcement learning provided in an embodiment of the present invention. Figure 2 The training flowchart of the deep Q-network decision model provided in the embodiment of the present invention is shown. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0017] Traditional focusing methods require image acquisition and evaluation across multiple focal planes, involving numerous focusing steps and significant time consumption, making them unsuitable for high-speed scanning or dynamic imaging. Existing sharpness evaluation functions lack robustness to complex samples (low contrast, sparse texture, or flare, etc.), easily getting trapped in local optima and struggling to handle unknown samples. Traditional algorithms have fixed parameters, requiring manual adjustment for different targets and experimental environments; they also lack online learning and environmental adaptation mechanisms. Supervised learning methods rely on large amounts of labeled focus data for model training, resulting in high data collection costs and difficulty in scaling up in scenarios with scarce annotations. To address these shortcomings, the core technical problems this invention aims to solve include: how to significantly accelerate the focusing process, improve the stability and robustness of the focusing process, and reduce the focusing model's dependence on labeled data.
[0018] Figure 1 The diagram shows the overall structure of the microscope autofocus system based on reinforcement learning provided in an embodiment of the present invention.
[0019] like Figure 1 As shown, the microscope autofocus system based on reinforcement learning provided by the present invention includes: The image acquisition module is used to acquire images of the current field of view through the microscope camera. For example, the current field of view image has an image size of 672×672 pixels.
[0020] The feature extraction module is used to calculate gradient features of the current field-of-view image and select the normalized gradient mean as the state vector. State vector The input is fed into the decision module of a deep Q-network (DQN), specifically, the state vector. It is fed into the Deep Q-Network (DQN) model.
[0021] The control module is used to control the stepper motor or piezoelectric platform to finely adjust the focal length along the Z-axis according to the adjustment command. This indicates the adjustment command to be executed based on the current focus position: ,in It is the focusing step size; optionally, the focusing step size is 1 micrometer.
[0022] The image sharpness evaluation module is used to guide the deep Q-network decision module to adjust towards better image sharpness based on the preset reward function of gradient energy; Define a pre-defined reward function based on gradient energy. Preset reward function : In other words, the increase in image sharpness is used as a reward signal; if the image becomes sharper after adjustment, the reward is positive, otherwise it is negative. This represents the gradient magnitude map of the image at the current moment. This represents the gradient magnitude map of the image at the next time step. To prevent instability caused by excessively frequent focusing, an action penalty term can be introduced, in which case a preset reward function is used. for: ,in, This is the focusing penalty coefficient.
[0023] The deep Q-network decision module utilizes a pre-trained deep Q-network decision model deployed to the edge computing unit of the microscope. Based on the input state vector and the reward function from the image sharpness evaluation module, it generates a focus adjustment command and sends it to the control module to achieve automatic focusing of the microscope.
[0024] The microscope autofocus system based on reinforcement learning provided in this application mainly relies on a deep Q-Network (DQN) to learn the optimal focusing strategy from the current image state, thereby achieving efficient adaptive focusing of complex samples in an unsupervised environment.
[0025] The gradient feature calculation in the feature extraction module specifically includes: using the Sobel operator Calculate the gradient magnitude map of the image: , in, Used to characterize the edge strength of an image and to measure image sharpness, Ix represents the gradient value of the image along the x-direction, and Iy represents the gradient value of the image along the y-direction.
[0026] Optionally, the structure of the deep Q-network decision model includes: The input layer receives the environmental state perceived by the current system. This environmental state is constructed using image sharpness-related features and is formally represented as a state vector. Its length is fixed at 128 dimensions, used to comprehensively characterize the sharpness information of the image at the current focal position. State vector The composition is as follows: (1) Image gradient magnitude statistics: The Sobel operator is used to extract the gradient maps in the horizontal and vertical directions of the image respectively. and The gradient magnitude map is calculated as follows: , Extract global mean, standard deviation, maximum value, entropy and other statistical features from G (a total of 16-dimensional global gradient features); (2) Local gradient features (blocking): Divide the image into 8*8 sub-blocks, calculate the average gradient intensity in each sub-block, and extract a total of 64 sets of local gradient feature information; (3) Histogram and frequency domain enhancement features: Normalize the gradient map and calculate its histogram, introduce the features of Fourier high-frequency energy distribution, and supplement the texture detail features, for a total of 48-dimensional features; (4) Normalization processing: The final input vector is input into the DQN network after normalization processing, which can ensure the stable distribution of input data and is conducive to model training.
[0027] The hidden layer consists of two fully connected layers, each with 256 units, using the ReLU function as the activation function. The first hidden layer takes the input 128-dimensional state vector as input. The model is projected into a 256-dimensional high-dimensional feature space to enhance its nonlinear modeling capabilities and has the advantages of computational efficiency and stable gradient propagation. The second hidden layer further extracts abstract features and enhances the model's ability to express complex state-action value functions, providing stable and discriminative high-order features for the output layer. Output layer: Output is an action value function The dimension is the number of actions, and the number of actions is 3.
[0028] The entire model optimization process is updated approximately using the Bellman equation: , in For learning rate, This is a discount factor for future rewards. For the new state, The maximum Q value for the next state.
[0029] Figure 2 The training flowchart of the deep Q-network decision model provided in the embodiment of the present invention is shown.
[0030] The Deep Q-Network (DQN) model used in this application achieves policy learning through continuous interaction with the microscope image environment, such as... Figure 2 As shown, the training process of the deep Q-network decision model includes the following steps: Initialize the parameters of the Q network and the target network. Specifically, initialize the main network. and target network The parameters are: the main network is used to output the action value estimate under the current policy, and the target network is a delayed update copy, which helps to stabilize the training process.
[0031] implement - A greedy strategy samples actions, and the interaction yields a state-action-reward-new state quadruple. At each step, based on... Greedy strategy selects actions , , Execute action By controlling the focus motor to update the focus position, a new image is acquired. ; Calculate image sharpness evaluation index As a reward, construct a quadruple. , where the state It is obtained by extracting features from the updated image.
[0032] An experience replay pool is constructed and trained using mini-batch samples, which consist of a preset N samples. All interaction samples are stored in the experience replay pool. During each training iteration, a mini-batch of N samples is randomly sampled from the replay pool for gradient descent to break the correlation between samples. Finally, mean squared error (MSE) is used to construct the target.
[0033] The Adam optimizer is used to backpropagate and update the Q network parameters. Every K steps, the Q network parameters are copied to the target network, thereby enhancing learning stability.
[0034] Once the loss function converges and the focusing performance stabilizes, the trained Q-network model is deployed to the edge computing unit of the microscope, thereby enabling the extraction of state features of real-time images and the generation of focus adjustment commands, ultimately achieving fast and high-precision autofocus.
[0035] When performing autofocus operations using the system provided in this application, steps such as image input, state extraction, motion decision-making, motor focusing, and reward feedback may be included.
[0036] In summary, this invention utilizes deep reinforcement learning to train intelligent strategies. Under an unsupervised learning framework, it rapidly and adaptively finds the optimal focus through closed-loop interaction with the microscope system, while reducing the number of samplings and time overhead, and achieving good generalization to various sample environments.
[0037] The above embodiments are only for illustrating the technical concept and features of the present invention, and are intended to enable those skilled in the art to understand the content of the present invention and implement it. They should not be used to limit the scope of protection of the present invention. All equivalent changes or modifications made in accordance with the spirit and essence of the present invention should be covered within the scope of protection of the present invention.
Claims
1. A microscope autofocusing system based on reinforcement learning, characterized in that, The system includes: The image acquisition module is used to acquire images of the current field of view through the microscope camera; The feature extraction module is used to calculate gradient features of the current field-of-view image, select the normalized gradient mean as the state vector, and input the state vector into the deep Q network decision module. The control module is used to control the stepper motor or piezoelectric platform to finely adjust the focal length along the Z-axis according to the adjustment command. This indicates the adjustment command to be executed based on the current focus position: ,in It focuses on step size; The image sharpness evaluation module is used to guide the deep Q-network decision module to adjust towards better image sharpness based on the preset reward function of gradient energy; The deep Q-network decision module utilizes a pre-trained deep Q-network decision model deployed to the edge computing unit of the microscope. Based on the input state vector and the reward function from the image sharpness evaluation module, it generates a focus adjustment command and sends it to the control module to achieve automatic focusing of the microscope.
2. The microscope autofocus system based on reinforcement learning according to claim 1, characterized in that, The current field-of-view image has a size of 672×672 pixels.
3. The microscope autofocus system based on reinforcement learning according to claim 1, characterized in that, The focusing step size is 1 micrometer.
4. The microscope autofocus system based on reinforcement learning according to claim 1, characterized in that, The gradient feature calculation in the feature extraction module specifically includes: using the Sobel operator Calculate the gradient magnitude map of the image: , in, Used to characterize image edge strength, and to measure image sharpness, I x I represents the gradient value of the image along the x-direction. y This represents the gradient value of the image along the y-direction.
5. The microscope autofocus system based on reinforcement learning according to claim 4, characterized in that, Preset reward function : .
6. The microscope autofocus system based on reinforcement learning according to claim 4, characterized in that, Preset reward function : ,in, This is the focusing penalty coefficient.
7. The microscope autofocus system based on reinforcement learning according to claim 1, characterized in that, The structure of the deep Q-network decision model includes: The input layer is used to receive the environmental state perceived by the current system. This environmental state is constructed through image sharpness-related features and is represented in the form of a state vector with a fixed length of 128 dimensions. The hidden layer consists of two fully connected layers, each with 256 units, using the ReLU function as the activation function. The first hidden layer takes the input 128-dimensional state vector as input. The projection is onto a 256-dimensional high-dimensional feature space, and the second hidden layer further extracts abstract features. Output layer: The output is an action value function, with the dimension being the number of actions, and the number of actions is 3.
8. The microscope autofocus system based on reinforcement learning according to claim 7, characterized in that, The training process of the deep Q-network decision model includes the following steps: Initialize the parameters of the Q network and the target network; implement - A greedy strategy samples actions and interacts to obtain a state-action-reward-new state quadruple; Construct an experience replay pool and train it using mini-batch samples, which consist of a preset N samples; The Adam optimizer is used to backpropagate and update the Q network parameters, copying the Q network parameters to the target network every K steps. Once the loss function converges and the focusing performance stabilizes, the trained Q-network model is deployed to the edge computing unit of the microscope.