An automatic focusing method and system for a cell slide scanner
By using multimodal information fusion and deep learning prediction to generate a focusing strategy map, the problem of long autofocus time and poor adaptability of traditional cell slide scanners is solved, and an efficient and accurate autofocus process is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- JIANGSU MICROCONTROL BIOTECHNOLOGY CO LTD
- Filing Date
- 2026-03-31
- Publication Date
- 2026-07-17
Smart Images

Figure CN121933422B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of optical device focusing technology, and in particular to an automatic focusing method and system for a cell slide scanner. Background Technology
[0002] Cell slide scanners are fundamental tools in modern medical diagnosis and life science research. Their core task is to convert physical slides into high-speed, high-resolution digital panoramic images. In this process, autofocus is crucial for ensuring image quality and directly determines the scanner's throughput and imaging reliability. Traditional autofocus methods are mainly divided into two categories:
[0003] One type of focusing method is based on the image sharpness evaluation function. This method involves driving the z-axis to acquire multiple images at different heights, calculating the sharpness evaluation value of each image, and using a search algorithm to find the peak position of the evaluation value as the optimal focal plane. However, this method requires multiple moves, resulting in slow focusing speed. The single evaluation function used is difficult to adapt to various complex samples such as bright field, fluorescence, and low contrast. It is also time-consuming for scanning large-size glass slides. Another type is the focusing method based on auxiliary sensors, which uses a laser displacement sensor to directly measure the relative distance between the slide surface and the scanner objective, thereby driving the objective to near the theoretical focal plane. This method has a systematic deviation in measurement accuracy and final imaging focal plane for samples with coverslips, sealing materials or uneven thickness, and cannot cope with focal plane changes caused by the characteristics of the sample itself. It must be combined with image methods for secondary fine-tuning, making system integration complex. Summary of the Invention
[0004] To overcome the shortcomings of existing technologies, this application provides an autofocus method and system for a cell slide scanner, which solves the problems of long autofocus time and poor adaptability to complex samples with low contrast, high background fluorescence, and scratches and contamination in traditional autofocus processes, and reduces the dependence of focus parameter settings on subjective experience. This method combines multimodal information fusion and deep learning prediction to generate a focus strategy map, and determines the optimal focus point through conditional verification and adjustment, thereby achieving high efficiency and high accuracy of autofocus in the cell slide scanner.
[0005] This application provides an autofocusing method for a cell slide scanner, comprising: Step S10: Control the stage of the cell slide scanner to perform a preview scan of the cell slide with a first resolution and a first step accuracy, and simultaneously acquire multimodal preview images of the cell slide and physical height sensing data of the corresponding two-dimensional coordinate position. Step S20: Input the multimodal preview image and the corresponding physical height sensing data into the pre-trained focus prediction model, and output a focus strategy map aligned with the glass slide coordinate space. This map includes the predicted optimal focus offset, focus function recommendation label, focus verification confidence and search radius for each field of view. Step S30: Based on the focusing strategy map, the final focus is determined through feedforward positioning, constraint verification, and conditional adjustment. The scanner stage is controlled at the final focus position, and the slide is formally scanned with a second resolution and a second step accuracy to acquire a complete scan image of the cell slide.
[0006] Furthermore, the multimodal preview images include transmitted bright-field images and fluorescence distribution images, and the physical height sensing data comes from a laser displacement sensor integrated next to the objective lens or stage. Transmitted bright-field images provide the physical topology of the cell slides, used to identify coverslip boundaries, tissue section outlines, marked areas, scratches, dust, and air bubbles; fluorescence distribution images provide biological distribution and density information of the cell slide samples, with the intensity and spatial distribution of fluorescence signals indicating enriched areas of cell nuclei and specific protein markers, aiding in the prediction of the optimal focal plane at the cellular level; physical height sensing data represents the relative distance between the sensor spot illumination point and the upper surface of the slide relative to the sensor baseline, reflecting the overall warp of the slide, carrier tilt, and height abrupt changes, providing geometric constraints for the focus prediction model.
[0007] Furthermore, the multimodal preview data obtained through preview scanning is packaged into channel tensors and input into a pre-trained focus prediction model. The model adopts an encoder-decoder structure with skip connections. The encoder fuses image features and height features through an adaptive feature fusion module, and the decoder upsamples the fused features to predict the optimal focus offset, focus function index, focus verification confidence, and search radius for each position on the cell slide. It outputs a focus strategy map corresponding to the slide space. A saliency filtering module is set between the encoder and decoder to highlight the tissue areas in the slide and assist the scanner in focusing on key positions.
[0008] Furthermore, step S30 implements fault-tolerant focusing based on the strategy graph, and for each target field of view, the following detailed steps are taken: Step S31: Read the predicted optimal focus offset at the corresponding position in the focusing strategy map, calculate the optimal focus coordinates, drive the Z-axis motor of the cell slide scanner stage to jump to the coordinate position, and complete the feedforward positioning. Step S32: Acquire a preview verification image at the predicted optimal focus coordinate position for limited verification, select a limited center area in the image, calculate the sharpness evaluation index, and compare the sharpness evaluation index with a preset threshold. Step S33: When the sharpness evaluation index is lower than the preset threshold, a local search is initiated within the Z-axis search range based on the recommended focus function identifier and search radius in the focus strategy diagram to determine the final focus; otherwise, the best focus coordinates predicted by the model are used as the final focus position.
[0009] This application also provides an autofocusing system for a cell slide scanner, comprising: Multimodal data acquisition module: used to control the stage of the cell slide scanner, to preview the cell slide with a first resolution and a first step accuracy, and to simultaneously acquire multimodal preview images of the cell slide, as well as physical height sensing data of the corresponding two-dimensional coordinate position; Focus strategy prediction module: It is used to input multimodal preview images and corresponding physical height sensing data into the pre-trained focus prediction model and output a focus strategy map aligned with the glass slide coordinate space. This map includes the predicted best focus offset, focus function recommendation label, focus verification confidence and search radius for each viewpoint. Strategy verification and focus execution module: Based on the focus strategy map, it determines the final focus through feedforward positioning, constraint verification and conditional adjustment, controls the scanner stage to be at the final focus position, performs formal scanning of the slide with second resolution and second step accuracy, and acquires complete scan images of the cell slide.
[0010] This application discloses the following technical effects: This application provides an autofocus method and system for a cell slide scanner. Based on multimodal preview data of cell slides, a deep learning model is applied to directly predict and locate the focus, omitting the time-consuming global or local search process in traditional methods. For most easily predictable fields of view, the autofocus process can be completed in milliseconds, greatly shortening the scanning time of the entire slide. This method integrates multimodal information, comprehensively utilizing the morphology, optical, and physical height information of the slide sample to assist the deep learning model in predicting a more comprehensive and accurate focusing strategy. The focusing strategy dynamically recommends the most suitable focusing function for different regions in the scanner's field of view, solving the problem of poor universality of a single focusing function. At the same time, it utilizes the adaptive capability of deep learning to improve the focusing success rate and robustness for complex samples. After generating the focusing strategy map, the method establishes a fault-tolerant mechanism, ensuring the slide imaging quality when the model prediction fails through verification and conditional adjustment, achieving a balance between speed and accuracy in the autofocus process. Attached Figure Description
[0011] Figure 1 A flowchart of an autofocusing method for a cell slide scanner provided in an embodiment of this application.
[0012] Figure 2This is a structural diagram of an autofocus system for a cell slide scanner provided in an embodiment of this application. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description of this application will be provided in conjunction with the accompanying drawings. The described embodiments should not be considered as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0014] Example 1: This application provides an automatic focusing method for a cell slide scanner, such as... Figure 1 As shown, the method includes: Step S10: Control the stage of the cell slide scanner to perform a preview scan of the cell slide with a first resolution and a first step accuracy, and simultaneously acquire multimodal preview images of the cell slide, as well as physical height sensing data of the corresponding two-dimensional coordinate position.
[0015] In this embodiment, before acquiring multimodal preview data, it is ensured that the cell slide scanner is equipped with a CMOS camera that can be quickly switched; the illumination system supports rapid switching between bright field transmitted light and fluorescence channels (such as DAPI); and a high-response-speed Z-axis displacement sensor (such as a laser triangulation sensor) has been installed and calibrated, and its measurement point and the center of the imaging field of view have been aligned on the two-dimensional plane.
[0016] Set the scanner's first resolution, set the camera to low-resolution readout mode, and switch the objective lens to low magnification. The purpose is not to see cell details, but to quickly obtain macroscopic distribution information. Set the first step accuracy of the stage preview scan. This step length is greater than the field of view overlap of the actual scan, usually 50μm to 200μm. The specific value should be such that it can cover the entire standard slide within 10-30 seconds.
[0017] During the preview scan, control the stage to move rapidly along a preset serpentine path. At each preview stop point, perform the following three data acquisition tasks: Acquiring transmission bright-field images: Turn on the bright-field light source, and the camera acquires a grayscale image with a preset short exposure time to capture the physical topology and contamination information of the slide. Areas with significantly different brightness in the image directly reflect the coverslip boundary, tissue section outline, marked areas, scratches, dust, and bubbles. Acquire fluorescence distribution images: Switch to a specific wavelength of fluorescence excitation light, set the intensity to 10%-30% of the formal scan, and the camera acquires images of the corresponding emission band to capture the biological distribution and density information of the cell slide sample. The intensity and spatial distribution of the fluorescence signal indicate the enrichment areas of the cell nucleus and specific protein markers, and assist the model in predicting the optimal focal plane at the cell level. Read physical height sensor data: At the same moment of camera exposure, read the height from the displacement sensor. This reading represents the relative distance between the sensor spot illumination point, the upper surface of the slide, and the sensor baseline. It reflects the overall warping of the slide, the tilt of the vehicle, and abrupt changes in height, providing geometric constraints for the model.
[0018] Due to the physical offset between the camera's field of view and the sensor's measurement point, a pre-calibrated coordinate transformation matrix is applied to all acquired preview images and height data sequences, using the stage coordinate system as a reference, to unify them into the same two-dimensional coordinate system. For transmitted bright-field images and fluorescence distribution images, flat-field correction is used to eliminate uneven illumination in the images, and contrast stretching is performed to improve image quality. For height data, the original sensor signal is converted into a micrometer height value, and filtering is performed to eliminate noise and glitches.
[0019] Step S20: Input the multimodal preview image and the corresponding physical height sensing data into the pre-trained focus prediction model, and output a focus strategy map aligned with the glass slide coordinate space. This map includes the predicted optimal focus offset, focus function recommendation identifier, focus verification confidence, and search radius for each field of view point.
[0020] In this embodiment, the preprocessed multimodal preview image and the corresponding height data are resampled onto a unified regular two-dimensional grid on sparse preview points through bilinear interpolation. The granularity of the grid matches the field-of-view spacing of the formal scan, ensuring that the corresponding predicted focal position coordinates can be found on the grid for each formal scan field of view. The pixel values of the transmission bright field image patch and fluorescence distribution image patch at each grid point are normalized and scaled to the range of [0,1]. The corresponding height data scalar values are expanded into a constant matrix with the same size as the image patch. The three types of data are used as input features for the focus prediction model.
[0021] The detailed steps for obtaining a pre-trained focus prediction model include: Collect historical focusing process data from the cell slide scanner, record multimodal preview data for each slide, as well as the corresponding final focus coordinates, focusing function type, confidence level, and search radius, to form training data; A focus prediction model based on an encoder-decoder structure is constructed. The encoder fuses image features and height features through an adaptive feature fusion module, and the decoder receives the fused features through a skip connection and saliency filtering module, upsamples the fused features, and finally outputs the predicted focus strategy map through a linear layer mapping. A multi-task loss function is adopted to calculate the focus regression loss for predicting the optimal focus offset, the cross-entropy loss for predicting the focus function recommendation identifier, and the regularization loss for predicting the focus verification confidence and search radius. The model parameters are updated by back gradient propagation based on the weighted values of the three loss functions. The model is trained by setting the training period and learning rate. When the multi-task loss function converges, the optimal model parameters are saved as the pre-trained focus prediction model.
[0022] The focus prediction model receives input features and feeds them into the adaptive feature fusion module in the encoder. One branch of this module stacks the transmission bright field image patch, fluorescence distribution image patch and height data matrix in the channel dimension and obtains the preliminary fused features through dynamic convolution. Another branch performs convolution processing on the transmission bright field image patch, fluorescence distribution image patch, and height data matrix, then adds them element-wise and generates an attention mask using the Sigmoid activation function. The attention mask is then multiplied with the preliminary fusion features, and key features in the preliminary fusion features are given higher weights. These weights are adaptively adjusted during model training to highlight feature information that helps with focus prediction, thus obtaining the final fusion features.
[0023] The fused features are processed by convolutional and pooling layers in the encoder to extract multi-scale deep features step by step, and then input into the saliency screening module. This module first convolves the original deep features to generate basic features, then sums the spatial dimensions of the convolution kernels to obtain neighborhood weights, which are multiplied with the basic features to obtain neighborhood focusing features; then the center element of the convolution kernel is extracted to obtain the center weights, which are multiplied with the basic features to obtain the center focusing features. In parallel, the original deep features are processed through global average pooling and linear layer mapping to generate channel attention weights. The salient channels in the original deep features are adaptively selected, and the center focus features are multiplied with the channel attention elements to enhance the center focus features of the salient channels, highlighting the relevant information of the tissue region in the cell slide and enabling the scanner to focus on the key position of the slide. The basic features, neighborhood focus features and enhanced center focus features are added together as the input features of the decoder. The decoder consists of multiple transposed convolutional layers and upsampling layers, upsampling the features to the same spatial size as the input and fusing them with the corresponding output features of the encoder at the same resolution to restore spatial details. The decoder ends with three parallel convolutional output heads, each corresponding to one of the three channels of the focus policy graph. The regression head predicts the optimal focus offset and outputs a single-channel feature map. Each pixel value represents the optimal focus predicted for that grid point, and the Z-axis offset (in micrometers) relative to the global reference plane. The classification head predicts the recommended identifier for the focus function and outputs a K-channel feature map, where K is the total number of focus function types. In this embodiment, K=4, which represent the gradient function, variance function, frequency domain function, and entropy function, respectively. The softmax function is activated in the K-channel dimension, and the channel index with the highest probability value is taken as the recommended focus function for that position. The hybrid output head predicts the focus verification confidence and search parameters, and outputs a dual-channel feature map. The first channel outputs the focus verification confidence, which comprehensively reflects the model's grasp of the predicted focus shift. The second channel outputs a non-negative value, representing the local search radius, which is negatively correlated with the confidence. The results from the three output heads are concatenated along the channel dimension to form the final focus strategy map. Each spatial location in this map contains the optimal focus offset, focus function recommendation identifier, focus verification confidence, and search radius.
[0024] Step S30: Based on the focusing strategy map, the final focus is determined through feedforward positioning, constraint verification, and conditional adjustment. The scanner stage is controlled at the final focus position, and the slide is formally scanned with a second resolution and a second step accuracy to acquire a complete scan image of the cell slide.
[0025] In this embodiment, this step implements fault-tolerant focusing based on the strategy graph. For each target field of view, the following detailed steps are taken: Step S31: Read the predicted optimal focus offset at the corresponding position in the focusing strategy map, calculate the optimal focus coordinates, and drive the Z-axis motor of the cell slide scanner stage to jump to the coordinate position to complete the feedforward positioning. When the stage moves to the absolute coordinates of the target's actual scanning field of view, the focusing strategy map is queried, and the corresponding quadruple strategy parameters are obtained through spatial mapping. The optimal focus coordinates are calculated based on the preset global reference focal plane coordinates and the optimal focus offset in the strategy parameters. The motion controller sends a command to the Z-axis motor to drive the stage of the cell slide scanner to move to the optimal focus coordinates at the maximum safe speed. Step S32: Acquire a preview verification image at the predicted optimal focus coordinates for limited verification. Select a limited central region in the image, calculate the sharpness evaluation index, and compare the sharpness evaluation index with a preset threshold. To achieve maximum speed, instead of calculating the sharpness of the entire image, a 128×128 pixel square region is cropped from the center of the preview verification image. This region typically has the highest probability of containing valid sample textures. For this central region, a sharpness evaluation index based on the Brenner gradient is calculated, which is the sum of squares of the differences between adjacent pixels within this region. This sharpness evaluation index is then compared with a preset threshold, which is dynamically adjusted based on the confidence level in the strategy graph.
[0026] in, This indicates a preset threshold that is dynamically adjusted. As the set base threshold, This represents the adjustment coefficient, which is set to 0.5 in this embodiment based on experience. Indicates the confidence level at the corresponding position; Step S33: When the sharpness evaluation index is lower than the preset threshold, a local search is initiated within the Z-axis search range based on the recommended focus function identifier and search radius in the focus strategy diagram to determine the final focus; otherwise, the optimal focus coordinates predicted by the model are used as the final focus position. When performing a local search, the search range is determined based on the search radius in the focus strategy map. Within this range, the golden section search method is used to move the Z-axis and acquire images at 3-7 different Z positions. The image sharpness evaluation index is calculated based on the focus function recommended by the strategy map. The position that maximizes the recommended focus function value is found through the search, and the final focus position is obtained.
[0027] Example 2: The autofocus system for a cell slide scanner provided in this embodiment of the invention can execute the autofocus method for a cell slide scanner provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method, such as... Figure 2 As shown, it has the following modules: Multimodal data acquisition module: used to control the stage of the cell slide scanner, to preview the cell slide with a first resolution and a first step accuracy, and to simultaneously acquire multimodal preview images of the cell slide, as well as physical height sensing data of the corresponding two-dimensional coordinate position; Focus strategy prediction module: It is used to input multimodal preview images and corresponding physical height sensing data into the pre-trained focus prediction model and output a focus strategy map aligned with the glass slide coordinate space. This map includes the predicted best focus offset, focus function recommendation label, focus verification confidence and search radius for each viewpoint. Strategy verification and focus execution module: Based on the focus strategy map, it determines the final focus through feedforward positioning, constraint verification and conditional adjustment, controls the scanner stage to be at the final focus position, performs formal scanning of the slide with second resolution and second step accuracy, and acquires complete scan images of the cell slide.
[0028] Although this application makes various references to certain modules in the system according to the embodiments of this application, any number of different modules can be used and run on user terminals and / or servers. The various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy distinction between each other and are not used to limit the scope of protection of this invention.
[0029] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application. In some cases, the actions or steps described in this application can be performed in a different order than that shown in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
Claims
1. An automatic focusing method for a cell slide scanner, characterized in that, The method includes: Step S10: Control the stage of the cell slide scanner to perform a preview scan of the cell slide with a first resolution and a first step accuracy, and simultaneously acquire multimodal preview images of the cell slide and physical height sensing data of the corresponding two-dimensional coordinate position. The multimodal preview image includes a transmission bright-field image and a fluorescence distribution image, and the physical height sensing data comes from a laser displacement sensor integrated next to the stage. Step S20: Input the multimodal preview image and the corresponding physical height sensing data into the pre-trained focus prediction model, and output a focus strategy map aligned with the glass slide coordinate space. This map includes the predicted optimal focus offset, focus function recommendation label, focus verification confidence and search radius for each field of view. Step S30: Based on the focusing strategy map, the final focus is determined through feedforward positioning, constraint verification, and conditional adjustment. The scanner stage is controlled at the final focus position, and the slide is formally scanned with a second resolution and a second step accuracy to acquire a complete scan image of the cell slide. Read the predicted optimal focus offset at the corresponding position in the focusing strategy diagram, calculate the optimal focus coordinates, drive the Z-axis motor of the cell slide scanner stage to jump to the coordinate position, and complete the feedforward positioning; A preview verification image is acquired at the predicted optimal focus coordinates for limited verification. A limited central region in the image is selected, and a sharpness evaluation index is calculated. The sharpness evaluation index is then compared with a preset threshold. When the sharpness evaluation index is lower than the preset threshold, a local search is initiated within the Z-axis search range based on the recommended focus function identifier and search radius in the focus strategy diagram to determine the final focus; otherwise, the best focus coordinates predicted by the model are used as the final focus position.
2. The automatic focusing method for a cell slide scanner as described in claim 1, characterized in that, In step S10, the transmitted bright-field image provides the physical topology of the cell slide, which is used to identify coverslip boundaries, tissue section outlines, marked areas, scratches, dust, and bubbles. The fluorescence distribution image provides biological distribution and density information of cell slide samples. The intensity and spatial distribution of the fluorescence signal indicate the enrichment areas of cell nuclei and specific protein markers, and help predict the optimal focal plane at the cell level. The physical height sensing data represents the relative distance between the sensor spot illumination point, the upper surface of the glass slide, and the sensor baseline. It reflects the overall warping of the glass slide, the tilt of the carrier, and abrupt changes in height, providing geometric constraints for the focus prediction model.
3. The automatic focusing method for a cell slide scanner as described in claim 1, characterized in that, Step S20, the detailed steps for obtaining the pre-trained focus prediction model include: Collect historical focusing process data from the cell slide scanner, record multimodal preview data for each slide, as well as the corresponding final focus coordinates, focusing function type, confidence level, and search radius, to form training data; A focus prediction model based on an encoder-decoder structure is constructed. The encoder fuses image features and height features through an adaptive feature fusion module, and the decoder receives the fused features through a skip connection and saliency filtering module, upsamples the fused features, and finally outputs the predicted focus strategy map through a linear layer mapping. A multi-task loss function is adopted to calculate the focus regression loss for predicting the optimal focus offset, the cross-entropy loss for predicting the focus function recommendation identifier, and the regularization loss for predicting the focus verification confidence and search radius. The model parameters are updated by back gradient propagation based on the weighted values of the three loss functions. The model is trained by setting the training period and learning rate. When the multi-task loss function converges, the optimal model parameters are saved as the pre-trained focus prediction model.
4. The automatic focusing method for a cell slide scanner as described in claim 3, characterized in that, One branch of the adaptive feature fusion module stacks the transmission bright field image, fluorescence distribution image, and physical height sensing data in the channel dimension, and obtains preliminary fusion features through dynamic convolution processing. Another branch processes the transmission bright-field image, fluorescence distribution image and physical height sensing data through a convolutional layer, adds them element by element, and generates an attention mask by passing them through a Sigmoid activation function; The attention mask is multiplied by the initial fused features to increase the weight of the key features in the initial fused features. This weight is then adaptively adjusted during model training to obtain the final fused features.
5. The automatic focusing method for a cell slide scanner as described in claim 3, characterized in that, The saliency screening module first performs convolution on the original input features of the module to generate basic features, then sums the spatial dimensions of the convolution kernel to obtain neighborhood weights, which are multiplied with the basic features to obtain neighborhood focusing features; then extracts the center element of the convolution kernel to obtain center weights, which are multiplied with the basic features to obtain center focusing features; In parallel, the original input features are processed through global average pooling and linear layer mapping to generate channel attention weights. The salient channels in the original input features are adaptively selected, and the center focus features are multiplied with the channel attention elements to enhance the center focus features of the salient channels, so that the scanner can focus on the key positions of the slide. The basic features, neighborhood focus features and enhanced center focus features are added together as the input features of the decoder.
6. The automatic focusing method for a cell slide scanner as described in claim 1, characterized in that, In step S30, during the formal scan, the second resolution is greater than the first resolution in step S10, and the second step accuracy is less than the first step accuracy, thereby obtaining imaging data with higher quality than the preview scan.
7. The automatic focusing method for a cell slide scanner as described in claim 1, characterized in that, In step S30, the sharpness evaluation index is based on the Brenner gradient and calculates the sum of squares of the differences between adjacent pixels within a defined region; The preset threshold is dynamically adjusted based on the confidence level in the strategy graph; The local search determines the search interval based on the search radius in the focus strategy map. Within this interval, the golden section search method is used to move the Z-axis and acquire images at 3-7 different Z positions. The image sharpness evaluation index is calculated based on the focus function recommended by the strategy map. The position that maximizes the recommended focus function value is found through the search, and the final focus position is obtained.
8. An automatic focusing system for a cell slide scanner, characterized in that, The system is used to implement the autofocusing method for a cell slide scanner according to any one of claims 1-7, the system comprising: Multimodal data acquisition module: used to control the stage of the cell slide scanner, to preview the cell slide with a first resolution and a first step accuracy, and to simultaneously acquire multimodal preview images of the cell slide, as well as physical height sensing data of the corresponding two-dimensional coordinate position; Focus strategy prediction module: It is used to input multimodal preview images and corresponding physical height sensing data into the pre-trained focus prediction model and output a focus strategy map aligned with the glass slide coordinate space. This map includes the predicted best focus offset, focus function recommendation label, focus verification confidence and search radius for each viewpoint. Strategy verification and focus execution module: Based on the focus strategy map, it determines the final focus through feedforward positioning, constraint verification and conditional adjustment, controls the scanner stage to be at the final focus position, performs formal scanning of the slide with second resolution and second step accuracy, and acquires complete scan images of the cell slide.