A method for estimating the depth information of motile cells and a robot micromanipulation device

By converting the depth estimation problem into a multi-category classification problem, using the fine-grained attention fusion module and feature enhancement method, the problem of low accuracy of estimation of depth information in motion cells is solved, and high-precision depth information acquisition is achieved, providing reliable visual feedback for robot micro-operation.

CN119784723BActive Publication Date: 2025-07-08SUZHOU BOUNDLESS MEDICAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411923420.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-07-08
Estimated Expiration
2044-12-25

AI Technical Summary

Technical Problem

In the prior art, the Z-axis depth information estimation method of moving cells under a monocular microscope has the problem of low ability to capture subtle differences in projection characteristics of adjacent focal planes and many-to-one relationships between morphology and depth values, resulting in low accuracy of depth information estimation.

Method used

A moving cell depth information estimation method is adopted to create a data set by discrete the continuous depth value range into multiple depth tags, and a multi-scale feature is extracted using the fine-grained attention fusion module (FGAF module). Combining channel cross-normalization and adaptive normalization methods, the fine-grained and robustness of the feature map is enhanced, and the weighted loss function is introduced to improve the accuracy of the depth estimation network.

Benefits of technology

It improves the accuracy and robustness of the depth information estimation of moving cells, enhances the generalization ability of the depth estimation network, and provides accurate and reliable visual feedback from robot micro-operation equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119784723B_ABST
    Figure CN119784723B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical fields of robot technology and image processing technology, and specifically refers to a method for estimating the depth information of moving cells and a robot micro-operation device, including: collecting multiple cell images at the depth values of each depth label, constructing a data set and dividing it into a training set and a test set; sequentially inputting each training image in the training set into the FGAF module of the depth estimation network, using the FG module to extract the weighted feature map of each training image, and then using the AF module to obtain the final feature map of each training image; then, according to the channel-based feature enhancement method, obtaining the enhanced feature map of each training image, and finally obtaining the depth value of the moving cells in each training image; training the depth estimation network by minimizing the weighted loss function to obtain a trained depth estimation network. The present invention reduces the problem complexity, improves the estimation accuracy of the depth value of moving cells, and enhances the robustness and generalization ability of the depth estimation network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of robotics and image processing technologies, and particularly to a method for estimating the depth information of moving cells and a robotic micro-operation device. Background Art

[0002] In modern biomedical research and clinical treatment, the development of microscopic cell manipulation technology is crucial. This technology is widely applied in multiple key fields, including in vitro fertilization, cell therapy, tissue engineering, and various types of biological experiments. Using robotic system-assisted cell micro-manipulation technology to obtain the three-dimensional position information of moving cells (such as sperm, bacteria, etc.) under a microscope is a key and challenging task, mainly facing the challenge of obtaining the Z-axis depth information of moving cells under a monocular microscope.

[0003] In the prior art, the methods for estimating the Z-axis depth information of moving cells under a monocular microscope include image-based depth estimation algorithms and classification estimation methods.

[0004] The image-based depth estimation algorithms include the DFF (Depth from Focus) algorithm and the DFD (Depth from Defocus) algorithm; among them, the essence of the DFF algorithm is autofocusing. It finds the clearest image by moving the focal plane up and down and using a focus metric algorithm to determine the three-dimensional position of the target object. However, this process based on repeated experiments limits the speed of the algorithm and makes it inapplicable to quickly locate moving cells; the DFD algorithm constructs a defocus model of the imaged object and establishes a mapping relationship between the blurred features in the image and the corresponding depth, thus avoiding the time-consuming process of finding the plane. However, this algorithm highly depends on the accuracy of the defocus model, and the rapidly changing projection shape during the cell movement process increases the difficulty of constructing a reliable defocus model. Therefore, the existing image-based depth estimation algorithms are not applicable to obtaining the depth information of moving cells.

[0005] The classification estimation methods include research on achieving autofocus by classifying the depth of a silica filament and estimating the offset of a micropipette and an adherent cell relative to the focused position through a network. However, these methods have only been tested on stationary target objects with fixed morphologies, and to avoid the problem of multiple morphologies corresponding to a single depth value caused by the rapid morphological changes during the cell movement process, the modeling process is simplified and not applicable to practical applications, resulting in low accuracy in estimating the depth information of moving cells.

[0006] In addition, although the prior art shows that deep networks can establish a mapping relationship between blurred image features and corresponding depth values at different depths, it does not consider improving the ability of deep networks to capture subtle differences between adjacent focal planes. As a result, it is unable to solve the problem that due to the small size of moving cells, lack of texture and structure information, the projection feature differences of moving cells between different focal planes are very small, leading to low accuracy in estimating the depth information of moving cells. Summary of the Invention

[0007] To this end, the technical problem to be solved by the present invention is to overcome the low ability to capture subtle differences in the projection features of moving cells between adjacent focal planes in the prior art, and the one-to-many relationship between the morphology of moving cells and depth values, resulting in low accuracy in estimating the depth information of moving cells.

[0008] To solve the above technical problems, the present invention provides a method for estimating the depth information of moving cells, including:

[0009] Discretize the continuous depth value range into multiple depth labels and their depth values according to the depth interval threshold; at the depth value of each depth label, collect multiple cell images; based on each cell image and the depth value of its corresponding depth label, construct a data set, and divide the data set into a training set and a test set;

[0010] Input each training image in the training set into the fine-grained attention fusion module of the depth estimation network in turn, perform feature extraction on each training image, and obtain the final feature map of each training image, including:

[0011] Input the current training image into the fine-grained feature extraction module of the fine-grained attention fusion module, and use the grouped convolutional layer of the fine-grained feature extraction module to extract each scale feature map of the current training image; input each scale feature map of the current training image into the global pooling layer of the fine-grained feature extraction module, and output the feature weight vector corresponding to each scale feature map of the current training image; input each scale feature map of the current training image and its corresponding feature weight vector into the activation layer of the fine-grained feature extraction module, and output the weighted feature map of the current training image;

[0012] Input the weighted feature map of the current training image into the attention fusion module of the fine-grained attention fusion module. After performing average pooling and max pooling operations on each channel of the weighted feature map of the current training image using the fusion pooling layer of the attention fusion module, concatenate the average pooling matrix and the max pooling matrix of the obtained weighted feature map, and output the fusion pooling matrix of the weighted feature map of the current training image; Input the fusion pooling matrix of the weighted feature map of the current training image into the fully connected layer of the attention fusion module, and output the channel importance weight matrix of the weighted feature map of the current training image; Input the weighted feature map of the current training image and its channel importance weight matrix into the weighted layer of the attention fusion module, and output the corrected feature map of the current training image; Input the corrected feature map of the current training image into the pooling convolutional layer of the attention fusion module, and output the final feature map of the current training image, thereby obtaining the final feature map of each training image;

[0013] Input the final feature map of each training image into the processing layer of the depth estimation network. Using the channel cross normalization method, perform an affine transformation on each channel of the final feature map of each training image, and then randomly exchange the features of any two channels to obtain the enhanced feature map of each training image;

[0014] Input the enhanced feature map of each training image into the output layer of the depth estimation network, and output the depth prediction value corresponding to the depth label of each training image as the depth value of the moving cells in each training image;

[0015] Use the training set to train the depth estimation network by minimizing the weighted loss function to obtain a trained depth estimation network for estimating the depth value of the moving cells.

[0016] Preferably, using the fine-grained feature extraction module of the fine-grained attention fusion module, according to the input current training image, the output of the weighted feature map of the current training image includes:

[0017] Input the current training image into the fine-grained feature extraction module of the fine-grained attention fusion module, and use the grouped convolutional layer of the fine-grained feature extraction module to extract each scale feature map of the current training image;

[0018] Input each scale feature map of the current training image into the global pooling layer of the fine-grained feature extraction module, and output the feature weight vector corresponding to each scale feature map of the current training image, and its expression is:

[0019] ;

[0020] where, represents the feature weight vector corresponding to the th scale feature map of the current training image; Represents the th scale feature map of the current training image; Represents the GeLU activation function; Represents the Sigmoid activation function; Represents the global pooling operation; Represents the global pooling weight; Represents the GeLU activation function weight;

[0021] Input each scale feature map of the current training image and its corresponding feature weight vector into the activation layer of the fine-grained feature extraction module, and output the weighted feature map of the current training image. Its expression is:

[0022] ;

[0023] Where, Represents the weighted feature map of the current training image; Represents the softmax function; Represents element-wise multiplication.

[0024] Preferably, using the attention fusion module of the fine-grained attention fusion module, according to the input weighted feature map of the current training image, the output final feature map of the current training image includes:

[0025] Input the weighted feature map of the current training image into the attention fusion module of the fine-grained attention fusion module. Use the fusion pooling layer of the attention fusion module to perform average pooling and max pooling operations on each channel of the weighted feature map of the current training image respectively, to obtain the average pooling value and max pooling value of each channel. Their expressions are respectively:

[0026] ;

[0027] ;

[0028] Where, Represents the average pooling value of the th channel in the weighted feature map of the current training image; Represents the max pooling value of the th channel in the weighted feature map of the current training image; Represents the weighted feature map of the current training image; Represents that the average pooling operation and the max pooling operation are performed on two dimensions with height and width ;

[0029] Concatenate the average pooling matrix and max pooling matrix of the weighted feature map obtained, and output the fusion pooling matrix of the weighted feature map of the current training image. Its expression is:

[0030] ;

[0031] Among them, represents the fusion pooling matrix of the weighted feature map of the current training image; represents the fusion function; represents the average pooling matrix of the weighted feature map of the current training image; represents the fusion pooling matrix of the weighted feature map of the current training image;

[0032] Input the fusion pooling matrix of the weighted feature map of the current training image into the fully connected layer of the attention fusion module, and output the channel importance weight matrix of the weighted feature map of the current training image. Its expression is:

[0033] ;

[0034] Among them, represents the channel importance weight matrix of the weighted feature map of the current training image; represents the fully connected operation;

[0035] Input the weighted feature map of the current training image and its channel importance weight matrix into the weighted layer of the attention fusion module, and output the corrected feature map of the current training image. Its expression is:

[0036] ;

[0037] Among them, represents the corrected feature map of the current training image; represents element-wise multiplication;

[0038] Input the corrected feature map of the current training image into the pooling convolutional layer of the attention fusion module, and output the final feature map of the current training image. Its expression is:

[0039] ;

[0040] Among them, represents the pooling operation; represents the convolutional operation; represents the final feature map of the current training image.

[0041] Preferably, input the final feature map of the current training image into the processing layer of the depth estimation network. Using the channel cross normalization method, after performing an affine transformation on each channel in the final feature map of the current training image, randomly exchange the features of any two channels to obtain the enhanced feature map of the current training image, including:

[0042] Input the final feature map of the current training image The processing layer of the input depth estimation network performs an affine transformation on each channel of the final feature map of the current training image to obtain new features for each channel in the final feature map of the current training image, and its expression is:

[0043] ;

[0044] where represents the new feature of the th channel in the final feature map of the current training image; represents the feature of the th channel in the final feature map of the current training image; represents the mean value of the th channel in the final feature map of the current training image; represents the variance of the th channel in the final feature map of the current training image; represents the first transformation parameter; represents the second transformation parameter;

[0045] Randomly select the th channel and the th channel in the final feature map of the current training image, and exchange the feature statistical attributes of the th channel and the th

[0046] ;

[0047] ;

[0048] where represents the target feature of the th channel in the final feature map of the current training image; represents the target feature of the th channel in the final feature map of the current training image; represents the new feature of the th channel in the final feature map of the current training image; represents the new feature of the th New features of each channel; Represents the final feature map of the current training image in the mean value of the Represents the final feature map of the current training image in the variance of the Represents the final feature map of the current training image in the mean value of the Represents the final feature map of the current training image in the variance of the

[0049] Based on the target features of all channels in the final feature map of the current training image, an enhanced feature map of the current training image is obtained.

[0050] Preferably, the expression of the weighted loss function is:

[0051] ;

[0052] Among them, represents the weighted loss function; represents the total number of all training images; represents the weight of the ; represents the true depth value of the depth label corresponding to the represents the predicted depth value of the depth label corresponding to the represents the predicted probability distribution corresponding to the

[0053] Preferably, after obtaining the trained depth estimation network, it further includes:

[0054] Input each test image in the test set into the fine-grained attention fusion module of the trained depth estimation network in turn, extract features from each test image, and obtain the final feature map of each test image;

[0055] Input the final feature map of each test image into the processing layer of the trained depth estimation network, and use the adaptive normalization method to perform affine transformation on each channel in the final feature map of each test image to obtain the enhanced feature map of each test image;

[0056] Input the enhanced feature map of each test image into the output layer of the trained depth estimation network, and output the depth prediction value corresponding to the depth label of each test image as the depth value of the moving cells in each test image.

[0057] Preferably, input the final feature map of the current test image into the processing layer of the trained depth estimation network, and use the adaptive normalization method to perform an affine transformation on each channel in the final feature map of the current test image to obtain the enhanced feature map of the current test image, including:

[0058] Input the final feature map of the current test image into the processing layer of the trained depth estimation network, and use the adaptive normalization method to perform an affine transformation on each channel in the final feature map of the current test image to obtain the new feature of each channel in the final feature map of the current test image, and its expression is:

[0059] ;

[0060] Wherein, represents the new feature of the th channel in the final feature map of the current test image; represents the feature of the th channel in the final feature map of the current test image; represents the first attention parameter; The second attention parameter; represents the mean value of the th channel in the final feature map of the current test image; represents the variance of the th channel in the final feature map of the current test image;

[0061] Based on the new features of all channels in the final feature map of the current test image, obtain the enhanced feature map of the current test image.

[0062] Preferably, the depth interval threshold is 0.5 μm; the continuous depth value range is -2 μm to 2 μm; the depth label is the first to the ninth depth labels; the depth values of the first to the ninth depth labels are -2 μm, -1.5 μm, -1 μm, -0.5 μm, 0 μm, 0.5 μm, 1 μm, 1.5 μm, 2 μm in sequence.

[0063] Preferably, before constructing the data set, it is also necessary to preprocess all the collected cell images to obtain the preprocessed cell images; wherein, the preprocessing includes denoising and contrast enhancement operations.

[0064] The present invention also provides a robot micro-operation device, including:

[0065] An inverted microscope, used for magnifying and viewing moving cells at different focal planes on the stage;

[0066] A CCD camera, placed at the imaging port of the inverted microscope and connected to a computer via a cable; used for collecting cell images at different focal planes under the inverted microscope and transmitting the cell images to the computer;

[0067] An XY electric platform, placed under the stage of the inverted microscope and connected to a computer via a cable; used for receiving control signals sent by the computer to control the movement of the microinjection needle in the X-axis and Y-axis directions;

[0068] A Z-axis electric micromanipulator, placed on the XY electric platform and connected to a computer via a cable; used for moving the microinjection needle in the Z-axis direction according to the depth value of the moving cell estimated by the computer;

[0069] A degree-of-freedom micro-operation arm, fixed to the Z-axis electric micromanipulator and adjacent to the inverted microscope; used for fixing the microinjection needle;

[0070] A microinjection needle, placed on the degree-of-freedom micro-operation arm; used for performing micro-operations on moving cells;

[0071] A vacuum suction pump, adjacent to the stage and connected to a computer via a cable; used for sucking moving cells;

[0072] A computer, connected to the inverted microscope via a cable, including:

[0073] A microscope focusing module, used for controlling the focal length of the inverted microscope according to the built-in focusing mechanism program;

[0074] An image processing module, used for performing preprocessing operations of denoising and contrast enhancement on each cell image collected by the CCD camera to obtain each preprocessed cell image; according to each preprocessed cell image, executing the program of a method for estimating the depth information of moving cells as described above built-in to estimate the depth value of the moving cells in each cell image;

[0075] A control module, used for receiving the estimated depth value of the moving cells in each cell image and controlling the XY electric platform and the Z-axis electric micromanipulator according to the depth value of the moving cells in each cell image to move the microinjection needle to the depth value of the moving cells, so as to achieve micro-operations and sucking on the moving cells;

[0076] An operation interface display module, used for displaying the depth value of the moving cells in each cell image and micro-operation guidance information.

[0077] The above technical solutions of the present invention have the following beneficial effects compared with the prior art:

[0078] A method for estimating the depth information of moving cells according to the present invention discretizes the continuous depth value range into a series of depth tags according to the depth interval threshold, and then converts the depth estimation problem into a multi-class depth classification problem, reducing the complexity of the problem, enhancing the robustness and generalization ability of the depth estimation network, facilitating the annotation of data and the understanding of the results; introducing a fine-grained attention fusion module (FGAF module), through the fine-grained feature extraction module (FG module) in the FGAF module, according to the grouped convolution of different scales, extracting multi-scale fine-grained features, through the attention fusion module (AF module) in the FGAF module, integrating the spatial attention and channel attention mechanisms, enhancing the features of the significant region, effectively overcoming the difficulty of depth estimation caused by the subtle difference in cell projection features between adjacent focal planes and the rapid change in cell morphology caused by cell movement, and then improving the accuracy of estimating the depth value of moving cells; in the training stage, using the channel cross-normalization method to enhance the features of the final feature map of the training image, by exchanging the statistical attributes between different channels of the high-dimensional feature map of the training image, expanding the data feature distribution, so that the model can learn the fine-grained feature differences in the cell image; in the test stage, using the adaptive normalization method to enhance the features of the final feature map of the test image, by learning the fully connected network and two attention mechanisms to adjust the affine transformation parameters, making the feature distribution of the test image closer to the feature distribution of the training image, improving the generalization ability and processing speed of the model in practical applications; introducing a weighted loss function, incorporating the penalty coefficient of the wrong prediction into the loss function, giving higher weights to the samples with larger depth category differences, making the depth estimation network focus on the fine-grained features that distinguish adjacent depth categories, improving the ability of the depth estimation network to accurately classify sperm images at different depth gradients, and further improving the accuracy of estimating the depth information of moving cells, thus providing accurate and reliable visual feedback of moving cells for the robot micro-operation device. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] In order to make the content of the present invention easier to be clearly understood, the following further details the present invention according to the specific embodiments of the present invention in conjunction with the drawings, wherein:

[0080] Figure 1 is a flowchart of a method for estimating the depth information of moving cells provided by the present invention;

[0081] Figure 2 is a schematic diagram of the fine-grained attention fusion module;

[0082] Figure 3 is a schematic diagram of the feature enhancement method; wherein, Figure 3 in (a) represents a schematic diagram of feature enhancement using the channel cross-normalization method; Figure 3 in (b) represents a schematic diagram of feature enhancement using the adaptive normalization method;

[0083] Figure 4 is a schematic diagram of a robot micro - operation device provided by the present invention;

[0084] Figure 5 is a schematic flow chart for realizing the depth estimation of moving cells based on the robot micro - operation device. Detailed implementation manners

[0085] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it. However, the specific embodiments cited do not limit the present invention.

[0086] Referring to Figure 1 as shown, Figure 1 is a flow chart of a method for estimating the depth information of moving cells provided by the present invention; aiming to solve the challenges faced in the Z - axis depth estimation of moving cells (such as sperm, bacteria, etc.) under an inverted microscope during the robot - assisted cell micro - operation process, and providing accurate and reliable visual feedback for the robot micro - operation device, specifically including:

[0087] S1: Discretize the continuous depth value range into multiple depth tags and their depth values according to the depth interval threshold; at the depth values of each depth tag, collect multiple cell images; based on each cell image and the depth value of its corresponding depth tag, construct a data set, and divide the data set into a training set and a test set; wherein, the depth interval threshold is 0.5 μm; the continuous depth value range is from - 2 μm to 2 μm; the depth tags are the first to the ninth depth tags; the depth values of the first to the ninth depth tags are - 2 μm, - 1.5 μm, - 1 μm, - 0.5 μm, 0 μm, 0.5 μm, 1 μm, 1.5 μm, 2 μm respectively; before constructing the data set, it is also necessary to pre - process all the collected cell images to obtain pre - processed cell images; the pre - processing includes denoising and contrast enhancement operations;

[0088] Among them, before determining to discretize the continuous depth value range into multiple depth tags and their depth values, it also includes:

[0089] Utilize the depth - of - field characteristic of the microscope to discretize the continuous depth values into a finite series of depth tags, and each depth tag corresponds to a depth focal plane. Thus, the original depth estimation problem can be reshaped into a multi - class depth classification problem;

[0090] The depth - of - field formula of the microscope is:

[0091] ;

[0092] Wherein, represents the light wave wavelength of the microscope; represents the refractive index of the medium between the microscope objective lens and the object; represents the numerical aperture of the microscope objective lens; represents the magnification of the objective lens; represents the aberration;

[0093] In a specific embodiment of the present invention, the numerical aperture is 0.65 and the magnification of the objective lens is 40. At this time, the depth of field of the microscope is about 1.0 μm. This 1.0-μm depth of field determines the resolution of the microscope in the Z-axis direction, that is, within the range of 1.0 μm, the object on the stage can be clearly imaged;

[0094] However, during the sperm movement process, the sperm head moves rapidly up and down along the Z-axis, making it impossible to be completely in the same focal plane, resulting in a phenomenon where the final image is partially clear and partially blurred;

[0095] In order to capture the subtle depth changes of sperm along the Z-axis and further improve the depth information resolution, the present invention discretizes the continuous depth values into a series of depth tags at intervals of 0.5 μm, and with the help of these depth tags, naturally converts the depth estimation problem under the microscope into a multi-class depth classification problem;

[0096] Therefore, the entire image acquisition process is carried out in the range of -2 to +2 μm, discretized into 9 different depth categories with a resolution of 0.5 μm (depth interval threshold), and a total of 2088 images are acquired. The resolution of each image is set to 96×96 pixels; among them, the automated acquisition method ensures the diversity and accuracy of the data, providing a stable and reliable basis for the multi-classification research of fine-grained feature differences;

[0097] In summary, by generating depth tags, the depth estimation problem is converted into a multi-class depth classification problem, and classification is carried out by learning the mapping relationship between image features and discrete depth tags; the multi-class depth classification problem is carried out in a supervised learning manner, and a labeled training set is used to train the depth estimation network to complete the classification task;

[0098] S2: Input each training image in the training set into the fine-grained attention fusion module of the depth estimation network in turn, extract features from each training image, and obtain the final feature map of each training image, including:

[0099] S21: Input the current training image into the fine-grained feature extraction module (FG module) of the fine-grained attention fusion module. Use the grouped convolutional layer of the fine-grained feature extraction module to extract the feature maps of each scale of the current training image. Among them, through grouped convolution, feature maps of different scales are obtained, which is convenient for independently processing feature maps of different scales in subsequent steps and avoids interference between different scales. At the same time, considering that the fine-grained feature extraction module is a plug-and-play module and does not require excessive addition of depth estimation network parameters, global pooling is used after grouped convolution to compress the spatial information of each scale feature map to reduce the number of parameters.

[0100] Input the feature maps of each scale of the current training image into the global pooling layer of the fine-grained feature extraction module, and output the feature weight vectors corresponding to the feature maps of each scale of the current training image. Its expression is:

[0101] ;

[0102] Among them, represents the feature weight vector corresponding to the th scale feature map of the current training image; represents the th scale feature map of the current training image; represents the GeLU activation function; represents the Sigmoid activation function; represents the global pooling operation; represents the global pooling weight; represents the GeLU activation function weight; the feature weight vector contains global semantic information.

[0103] Input the feature maps of each scale of the current training image and their corresponding feature weight vectors into the activation layer of the fine-grained feature extraction module, and output the weighted feature map of the current training image. Its expression is:

[0104] ;

[0105] Among them, represents the weighted feature map of the current training image; represents the softmax function; represents element-wise multiplication; among them, the softmax function is used to balance the importance of features of different scales, recalibrate the feature weight vector, ensure that the sum of feature weight vectors of different scales is 1, and facilitate the integration of multi-scale feature information.

[0106] S22: Input the weighted feature map of the current training image into the attention fusion module (AF module) of the fine-grained attention fusion module. Use the fusion pooling layer of the attention fusion module to perform average pooling and max pooling operations on each channel of the weighted feature map of the current training image respectively, obtaining the average pooling value and max pooling value of each channel. Their expressions are respectively:

[0107] ;

[0108] ;

[0109] Among them, represents the average pooling value of the th channel in the weighted feature map of the current training image; represents the max pooling value of the th channel in the weighted feature map of the current training image; represents the weighted feature map of the current training image; represents that the average pooling operation and max pooling operation are performed on two dimensions with height and width ; among them, typical features within each channel are captured through the average pooling operation, and extreme features within each channel are captured through the max pooling operation;

[0110] Concatenate the average pooling matrix and max pooling matrix of the weighted feature map obtained, output the fusion pooling matrix of the weighted feature map of the current training image, and its expression is:

[0111] ;

[0112] Among them, represents the fusion pooling matrix of the weighted feature map of the current training image; represents the fusion function; represents the average pooling matrix of the weighted feature map of the current training image; represents the fusion pooling matrix of the weighted feature map of the current training image;

[0113] Input the fusion pooling matrix of the weighted feature map of the current training image into the fully connected layer of the attention fusion module, and output the channel importance weight matrix of the weighted feature map of the current training image. Its expression is:

[0114] ;

[0115] Among them, represents the channel importance weight matrix of the weighted feature map of the current training image; Represents a fully connected operation; among them, the fused pooling matrix obtained after splicing is sent to the fully connected layer to analyze and learn the dependence relationship of features between different channels;

[0116] The weighted feature map of the current training image and its channel importance weight matrix are input into the weighted layer of the attention fusion module, and the corrected feature map of the current training image is output. Its expression is:

[0117] ;

[0118] Among them, Represents the corrected feature map of the current training image;

[0119] The corrected feature map of the current training image is input into the pooling convolutional layer of the attention fusion module, and the final feature map of the current training image is output, and then the final feature map of each training image is obtained. Its expression is:

[0120] ;

[0121] Among them, Represents a pooling operation; Represents a convolutional operation; Represents the final feature map of the current training image; among them, in order to build a bridge between the local and global feature spaces in the cell image, the corrected feature map is spatially re-adjusted; during the re-adjustment process, a pooling operation is involved, aiming to reduce the spatial dimension of the corrected feature map, which helps to more significantly emphasize the global features; subsequently, through a convolutional operation, complex connections are established between the local fine-grained features and the broader global features of the cell image, so that the final feature map output encapsulates the fine-grained local features and enhanced global features of the cell image;

[0122] Based on S22, it can be seen that the attention fusion module focuses on extracting key features within a single channel and adjusts the spatial relationship to build a bridge between the local and global feature spaces in the cell image; the AF module, for the FG module, fuses multi-scale weighted feature maps, performs average pooling and max pooling on each channel, captures typical features and extreme features, and splices the pooled statistical information and inputs it into the fully connected layer to learn the dependence relationship of features between different channels, forms a channel importance weight matrix, multiplies it element-wise with the weighted feature map, re-calibrates the channel weights, and obtains the corrected feature map; then, the corrected feature map is spatially adjusted, and the pooling operation is used to compress the spatial dimension to highlight the global features, and then through the convolutional operation, complex relationships are established between the local fine-grained features and the global features, and the final feature map with both fine-grained local features and enhanced global features is output;

[0123] In summary, the fine-grained attention fusion module (FGAF module) consists of a fine-grained feature extraction module and an attention fusion module, which is used to identify the subtle differences between different focal planes. Its structure is as shown in Figure 2 ; The fine-grained feature extraction module extracts feature information such as texture, edge, and morphology in the images of moving cells at different focal planes by analyzing the images of moving cells at different focal planes, that is, uses grouped convolutions of different scales to extract multi-scale fine-grained features; The attention fusion module is used to weight the contributions of features at different focal planes to highlight the features that are most critical for depth determination, that is, integrates spatial attention and channel attention mechanisms to enhance the features of significant regions;

[0124] S3: Input the final feature map of each training image into the processing layer of the depth estimation network. After performing an affine transformation on each channel in the final feature map of each training image using the channel cross normalization method, randomly exchange the features of any two channels to obtain the enhanced feature map of each training image, including:

[0125] Input the final feature map of the current training image into the processing layer of the depth estimation network, perform an affine transformation on each channel in the final feature map of the current training image to obtain the new features of each channel in the final feature map of the current training image. Its expression is:

[0126] ;

[0127] Among them, represents the new feature of the th channel in the final feature map of the current training image ; represents the feature of the th channel in the final feature map of the current training image ; represents the mean value of the th channel in the final feature map of the current training image ; represents the variance of the th channel in the final feature map of the current training image ; represents the first transformation parameter; represents the second transformation parameter;

[0128] Randomly select the th channel and the th channel in the final feature map of the current training image, exchange the feature statistical attributes of the th channel and the th channel, and obtain the th channel, The target features of the th channel and the

[0129] th channel are expressed as follows:

[0130] ;

[0131] where represents the target feature of the th channel in the final feature map of the current training image ; represents the target feature of the th channel in the final feature map of the current training image ; represents the new feature of the th channel in the final feature map of the current training image ; represents the new feature of the th channel in the final feature map of the current training image ; represents the mean value of the th channel in the final feature map of the current training image ; represents the variance of the th channel in the final feature map of the current training image ; represents the mean value of the th channel in the final feature map of the current training image ; represents the variance of the th channel in the final feature map of the current training image ;

[0132] Based on the target features of all channels in the final feature map of the current training image, an enhanced feature map of the current training image is obtained;

[0133] In summary, in the training stage, the channel cross-normalization method (CrossNorm) is used to enhance the fine-grained features in the cell image; by exchanging the feature statistical attributes between different channels of the high-dimensional feature map of the training image, the data feature distribution is extended, enabling the depth estimation network to learn the fine-grained differences in the cell image; the process of the channel cross-normalization method is as shown in Figure 3 (a) in

[0134] S4: Input the enhanced feature map of each training image into the output layer of the depth estimation network, and output the depth prediction value corresponding to the depth label of each training image as the depth value of the moving cells in each training image;

[0135] S5: Using the training set, train the depth estimation network by minimizing the weighted loss function to obtain a trained depth estimation network for estimating the depth value of moving cells;

[0136] Among them, in the fine-grained classification task of cell images with different depth labels, the traditional cross-entropy loss function is not effective in capturing the subtle differences between adjacent depth labels; since the fine-grained feature differences between cell images corresponding to adjacent depth labels are very small, if the cross-entropy loss function is used to train the depth estimation network, misclassification problems are likely to occur in the initial training stage; to solve this limitation and improve the classification accuracy of the depth estimation network, the present invention introduces a weighted loss function, incorporates a penalty coefficient for incorrect predictions into the loss function, and encourages the model to pay attention to the fine-grained feature differences between adjacent depth labels. Then, the expression of the weighted loss function is:

[0137] ;

[0138] Among them, represents the weighted loss function; represents the total number of all training images; represents the weight of the th training image, represents the true depth value of the depth label corresponding to the th training image; represents the predicted depth value of the depth label corresponding to the th training image; represents the predicted probability distribution corresponding to the

[0139] th training image;

[0140] Based on the completion of steps S1-S5, after obtaining a trained depth estimation network, it further includes:

[0141] Input each test image in the test set into the fine-grained attention fusion module of the trained depth estimation network in sequence, extract features from each test image to obtain the final feature map of each test image;

[0142] Input the final feature map of each test image into the processing layer of the trained depth estimation network, and use the adaptive normalization method to perform an affine transformation on each channel in the final feature map of each test image to obtain the enhanced feature map of each test image, including:

[0143] Input the final feature map of the current test image into the processing layer of the trained depth estimation network. Using the adaptive normalization method, perform an affine transformation on each channel in the final feature map of the current test image to obtain new features for each channel in the final feature map of the current test image. The expression is as follows:

[0144] ;

[0145] Among them, represents the new feature of the th channel in the final feature map of the current test image; represents the feature of the th channel in the final feature map of the current test image; represents the first attention parameter; The second attention parameter; represents the mean value of the th channel in the final feature map of the current test image; represents the variance of the th channel in the final feature map of the current test image;

[0146] Based on the new features of all channels in the final feature map of the current test image, obtain the enhanced feature map of the current test image;

[0147] To sum up, in the test stage, use the adaptive normalization method (SelfNorm) to enhance the fine-grained features in the cell image; adjust the affine transformation parameters by learning the fully connected network and two attention mechanisms, so that the data distribution in the test stage is closer to the training data, improve the robustness of the depth estimation network and its generalization ability in practical applications; at the same time, enhance the stability and generalization ability of the depth estimation network, so that the depth estimation network can also extract key fine-grained features when facing unseen cell images; the process of the adaptive normalization method is as shown in Figure 3 in (b); the adaptive normalization method can be applied not only in the test stage, but also in the training stage;

[0148] Combining the training stage and the test stage, it can be seen that the channel-based feature enhancement method includes the channel cross-normalization method and the adaptive normalization method. By normalizing and gain-adjusting the features of each channel, it adapts to different imaging conditions and environmental changes, thereby enhancing the fine-grained features and improving the generalization ability of the model;

[0149] Input the enhanced feature map of each test image into the output layer of the trained depth estimation network, and output the depth prediction value corresponding to the depth label of each test image as the depth value of the moving cells in each test image.

[0150] The present invention also provides a robot micro - operation device, as Figure 4 shown, comprising:

[0151] An inverted microscope for magnifying and viewing moving cells at different focal planes on the stage;

[0152] A CCD camera placed at the imaging port of the inverted microscope and connected to a computer via a cable; for collecting cell images at different focal planes under the inverted microscope and transmitting the cell images to the computer; wherein, the CCD camera can capture the dynamic activities of moving cells and the micro - operation process;

[0153] An XY electric platform placed under the stage of the inverted microscope and connected to the computer via a cable; for receiving control signals sent by the computer and controlling the movement of the micro - injection needle in the X - axis and Y - axis directions;

[0154] A Z - axis electric micro - manipulator placed on the XY electric platform and connected to the computer via a cable; for moving the micro - injection needle in the Z - axis direction according to the depth value of the moving cell estimated by the computer;

[0155] A multi - degree - of - freedom micro - operation arm fixed on the Z - axis electric micro - manipulator and adjacent to the inverted microscope, i.e., close to the inverted microscope; for fixing the micro - injection needle;

[0156] A micro - injection needle placed on the multi - degree - of - freedom micro - operation arm; for performing micro - operations on moving cells;

[0157] A vacuum suction pump adjacent to the stage and connected to the computer via a cable; for sucking moving cells;

[0158] A computer connected to the inverted microscope via a cable, comprising:

[0159] A microscope focusing module for controlling the focal length of the inverted microscope according to a built - in focusing mechanism program to obtain clear cell images at different focal planes;

[0160] An image processing module for performing pre - processing operations such as denoising, contrast enhancement, and other image optimizations on each cell image obtained by the CCD camera to obtain each pre - processed cell image; according to each pre - processed cell image, executing a program of the above - mentioned method for estimating the depth information of moving cells to estimate the depth value of the moving cells in each cell image; wherein, the process of estimating the depth information of moving cells based on the robot micro - operation device is as Figure 5 shown;

[0161] A control module, configured to receive the estimated depth values of moving cells in each cell image, and control the XY electric platform and the Z-axis electric micromanipulator according to the depth values of the moving cells in each cell image, so as to move the microinjection needle to the depth value of the moving cells, and achieve micro-operation and aspiration of the moving cells;

[0162] An operation interface display module, configured to display the depth values of moving cells in each cell image and micro-operation guidance information, so as to assist the operator in precise control.

[0163] In the robot micro-operation system, by combining the depth estimation algorithm with the XY plane positioning algorithm, the three-dimensional spatial information of sperm is obtained, and the automatic tracking and aspiration of sperm are realized.

[0164] In a specific embodiment of the present invention, taking the sperm aspiration task as an example, by moving the electric XY platform, a constant distance of 100 pixels is maintained between the target sperm and the tip of the microinjection needle. The XY plane position information of the target sperm head is provided to the microinjection needle through a pre-trained UNeXt semantic segmentation model. At the same time, the depth estimation network predicts the Z-axis depth information of the sperm. After obtaining the depth information, the tip of the microinjection needle moves to the corresponding depth. During the entire aspiration process, the injection pump maintains a constant flow rate (1.2 nL / s).

[0165] In summary, the method for estimating the depth information of moving cells and the robot micro-operation device provided by the present invention convert the depth estimation problem into a multi-class depth classification problem by adopting a multi-class classification strategy. By introducing a fine-grained attention fusion module (FGAF module), which combines the fine-grained feature extraction module and the attention fusion module, it effectively identifies the subtle differences between focal planes, and solves the problem that it is difficult to accurately obtain depth information due to the limited depth of field of the microscope and the large dynamic changes of cells themselves. The present invention applies a channel-based feature enhancement method to improve the generalization ability and processing speed of the depth estimation network. In addition, the present invention also demonstrates the application potential and practicality in the robot micro-operation system.

[0166] Obviously, the above embodiments are only examples for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.

Claims

1. A method for estimating the depth information of moving cells, characterized in that, Including: Discretize the continuous depth value range into multiple depth tags and their depth values according to the depth interval threshold; Collect multiple cell images at the depth value of each depth tag; Based on each cell image and the depth value of its corresponding depth tag, construct a data set and divide the data set into a training set and a test set; Input each training image in the training set into the fine-grained attention fusion module of the depth estimation network in turn, extract features from each training image, and obtain the final feature map of each training image, including: Input the current training image into the fine-grained feature extraction module of the fine-grained attention fusion module, use the grouped convolutional layer of the fine-grained feature extraction module to extract each scale feature map of the current training image; input each scale feature map of the current training image into the global pooling layer of the fine-grained feature extraction module, and output the feature weight vector corresponding to each scale feature map of the current training image; input each scale feature map of the current training image and its corresponding feature weight vector into the activation layer of the fine-grained feature extraction module, and output the weighted feature map of the current training image; Input the weighted feature map of the current training image into the attention fusion module of the fine-grained attention fusion module. After performing average pooling and max pooling operations on each channel of the weighted feature map of the current training image respectively using the fusion pooling layer of the attention fusion module, splice the average pooling matrix and the max pooling matrix of the obtained weighted feature map, and output the fusion pooling matrix of the weighted feature map of the current training image; input the fusion pooling matrix of the weighted feature map of the current training image into the fully connected layer of the attention fusion module, and output the channel importance weight matrix of the weighted feature map of the current training image; input the weighted feature map of the current training image and its channel importance weight matrix into the weighted layer of the attention fusion module, and output the corrected feature map of the current training image; input the corrected feature map of the current training image into the pooling convolutional layer of the attention fusion module, and output the final feature map of the current training image, thereby obtaining the final feature map of each training image; Input the final feature map of each training image into the processing layer of the depth estimation network, and use the channel cross normalization method to perform an affine transformation on each channel in the final feature map of each training image, and then randomly exchange features between any two channels to obtain the enhanced feature map of each training image; Input the enhanced feature map of each training image into the output layer of the depth estimation network, and output the depth prediction value of the depth tag corresponding to each training image as the depth value of the moving cells in each training image; Use the training set to train the depth estimation network by minimizing the weighted loss function to obtain a trained depth estimation network for estimating the depth value of moving cells.

2. The method for estimating the depth information of moving cells according to claim 1, wherein, Using the fine-grained feature extraction module of the fine-grained attention fusion module, according to the input current training image, output the weighted feature map of the current training image, including: Input the current training image into the fine-grained feature extraction module of the fine-grained attention fusion module, and use the grouped convolutional layer of the fine-grained feature extraction module to extract each scale feature map of the current training image; Input each scale feature map of the current training image into the global pooling layer of the fine-grained feature extraction module, and output the feature weight vector corresponding to each scale feature map of the current training image. Its expression is: ; Among them, represents the feature weight vector corresponding to the -th scale feature map of the current training image; represents the -th scale feature map of the current training image; represents the GeLU activation function; represents the Sigmoid activation function; represents the global pooling operation; represents the global pooling weight; represents the GeLU activation function weight; Input each scale feature map of the current training image and its corresponding feature weight vector into the activation layer of the fine-grained feature extraction module, and output the weighted feature map of the current training image. Its expression is: ; Among them, represents the weighted feature map of the current training image; represents the softmax function; represents element-wise multiplication.

3. The method for estimating the depth information of motile cells according to claim 1, wherein Using the attention fusion module of the fine-grained attention fusion module, according to the input weighted feature map of the current training image, the output final feature map of the current training image includes: Input the weighted feature map of the current training image into the attention fusion module of the fine-grained attention fusion module. Use the fusion pooling layer of the attention fusion module to perform average pooling and max pooling operations on each channel of the weighted feature map of the current training image respectively, to obtain the average pooling value and max pooling value of each channel. Their expressions are respectively: ; ; Among them, represents the average pooling value of the th channel in the weighted feature map of the current training image; represents the maximum pooling value of the th channel in the weighted feature map of the current training image; represents the weighted feature map of the current training image; represents that the average pooling operation and the maximum pooling operation are performed on two dimensions with a height of and a width of ; Concatenate the average pooling matrix and max pooling matrix of the obtained weighted feature map, and output the fusion pooling matrix of the weighted feature map of the current training image. Its expression is: ; Among them, represents the fusion pooling matrix of the weighted feature map of the current training image; represents the fusion function; represents the average pooling matrix of the weighted feature map of the current training image; represents the fusion pooling matrix of the weighted feature map of the current training image; Input the fusion pooling matrix of the weighted feature map of the current training image into the fully connected layer of the attention fusion module, and output the channel importance weight matrix of the weighted feature map of the current training image. Its expression is: ; Among them, represents the channel importance weight matrix of the weighted feature map of the current training image; represents the fully connected operation; Input the weighted feature map of the current training image and its channel importance weight matrix into the weighted layer of the attention fusion module, and output the corrected feature map of the current training image. Its expression is: ; Among them, represents the corrected feature map of the current training image; represents element-wise multiplication; Input the corrected feature map of the current training image into the pooling convolutional layer of the attention fusion module, and output the final feature map of the current training image. Its expression is: ; Among them, represents a pooling operation; represents a convolution operation; represents the final feature map of the current training image.

4. The method for estimating the depth information of moving cells according to claim 1, wherein Input the final feature map of the current training image into the processing layer of the depth estimation network. Using the channel cross normalization method, after performing an affine transformation on each channel of the final feature map of the current training image, randomly exchange the features of any two channels to obtain the enhanced feature map of the current training image, including: The final feature map of the current training image is input into the processing layer of the depth estimation network to perform an affine transformation on each channel in the final feature map of the current training image to obtain new features for each channel in the final feature map of the current training image. The expression is as follows: ; Among them, represents the new feature of the th channel in the final feature map of the current training image; represents the feature of the th channel in the final feature map of the current training image; represents the mean value of the th channel in the final feature map of the current training image; represents the variance of the th channel in the final feature map of the current training image; represents the first transformation parameter; represents the second transformation parameter; Randomly select the final feature map of the current training image in the th channel and the th channel, swap the feature statistical attributes of the th channel and the th channel to obtain the target features of the th channel and the th channel, and their expressions are respectively: ; ; Among them, represents the target feature of the th channel in the final feature map of the current training image; represents the target feature of the th channel in the final feature map of the current training image; represents the new feature of the th channel in the final feature map of the current training image; represents the new feature of the th channel in the final feature map of the current training image; represents the mean value of the th channel in the final feature map of the current training image; represents the variance of the th channel in the final feature map of the current training image; represents the mean value of the th channel in the final feature map of the current training image; represents the variance of the th channel in the final feature map of the current training image; Based on the target features of all channels in the final feature map of the current training image, obtain the enhanced feature map of the current training image.

5. A method for estimating the depth information of moving cells according to claim 1, characterized in that, The expression of the weighted loss function is: ; Among them, represents the weighted loss function; represents the total number of all training images; represents the weight of the -th training image; represents the ground truth depth value of the depth label corresponding to the -th training image; represents the depth prediction value of the depth label corresponding to the -th training image; represents the predicted probability distribution corresponding to the -th training image.

6. The method for estimating the depth information of moving cells according to claim 1, characterized in that, After obtaining the trained depth estimation network, it further includes: Input each test image in the test set into the fine-grained attention fusion module of the trained depth estimation network in sequence, perform feature extraction on each test image, and obtain the final feature map of each test image; Input the final feature map of each test image into the processing layer of the trained depth estimation network, and use the adaptive normalization method to perform an affine transformation on each channel of the final feature map of each test image to obtain the enhanced feature map of each test image; Input the enhanced feature map of each test image into the output layer of the trained depth estimation network, and output the depth prediction value corresponding to the depth label of each test image as the depth value of the moving cells in each test image.

7. A method for estimating the depth information of moving cells according to claim 6, characterized in that, Input the final feature map of the current test image into the processing layer of the trained depth estimation network, and use the adaptive normalization method to perform an affine transformation on each channel of the final feature map of the current test image to obtain the enhanced feature map of the current test image, including: Input the final feature map of the current test image into the processing layer of the trained depth estimation network. Using the adaptive normalization method, perform an affine transformation on each channel in the final feature map of the current test image to obtain new features for each channel in the final feature map of the current test image. The expression is as follows: ; Among them, represents the new feature of the th channel in the final feature map of the current test image; represents the feature of the th channel in the final feature map of the current test image; represents the first attention parameter; The second attention parameter; represents the mean value of the th channel in the final feature map of the current test image; represents the variance of the th channel in the final feature map of the current test image; Based on the new features of all channels in the final feature map of the current test image, obtain the enhanced feature map of the current test image.

8. The method for estimating the depth information of motile cells according to claim 1, wherein The depth interval threshold is 0.5 μm; the continuous depth value range is from -2 μm to 2 μm; the depth labels are the first to the ninth depth labels; the depth values of the first to the ninth depth labels are -2 μm, -1.5 μm, -1 μm, -0.5 μm, 0 μm, 0.5 μm, 1 μm, 1.5 μm, 2 μm in sequence.

9. The method for estimating the depth information of moving cells according to claim 1, wherein Before constructing the dataset, it is also necessary to preprocess all the collected cell images to obtain the preprocessed cell images; among them, the preprocessing includes denoising and contrast enhancement operations.

10. A robot micro-operation device, characterized in that, Including: An inverted microscope for magnifying and viewing moving cells at different focal planes on the stage. A CCD camera placed at the imaging port of the inverted microscope and connected to the computer via a cable. Used to collect cell images at different focal planes under the inverted microscope and transmit the cell images to the computer. An XY electric platform placed under the stage of the inverted microscope and connected to the computer via a cable; used to receive the control signal sent by the computer and control the movement of the microinjection needle in the X-axis and Y-axis directions. A Z-axis electric micromanipulator placed on the XY electric platform and connected to the computer via a cable; used to move the microinjection needle in the Z-axis direction according to the depth value of the moving cell estimated by the computer. A degree-of-freedom micro-operation arm fixed on the Z-axis electric micromanipulator and adjacent to the inverted microscope. Used to fix the microinjection needle. The microinjection needle is placed on the degree-of-freedom micro-operation arm. Used to perform micro-operations on moving cells. A vacuum suction pump adjacent to the stage and connected to the computer via a cable. Used to suck moving cells. A computer connected to the inverted microscope via a cable, including: A microscope focusing module used to control the focal length of the inverted microscope according to the built-in focusing mechanism program. An image processing module used to perform preprocessing operations of denoising and contrast enhancement on each cell image collected by the CCD camera to obtain each preprocessed cell image; according to each preprocessed cell image, execute the program of a method for estimating the depth information of moving cells as described in any one of claims 1 to 9 built-in to estimate the depth value of the moving cell in each cell image. A control module used to receive the estimated depth value of the moving cell in each cell image and control the XY electric platform and the Z-axis electric micromanipulator according to the depth value of the moving cell in each cell image to move the microinjection needle to the depth value of the moving cell to achieve micro-operations and suction on the moving cell. An operation interface display module used to display the depth value of the moving cell in each cell image and the micro-operation guidance information.

Citation Information

Patent Citations

  • Monocular depth estimation method and system based on infrared image

    CN116168070A

  • Unsupervised monocular depth estimation method based on fine-grained prediction and edge extraction

    CN118674762A