An Automatic Focusing Method for a Single-Cell Mass Spectrometry System Based on Deep Learning
By adopting deep learning-based image processing technology in the single-cell mass spectrometry system, a model including regional filtering mechanism, feature extraction network, classifier and distributed tag encoding is constructed, and the problem of poor accuracy and low efficiency of the autofocus method of the single-cell mass spectrometry system is solved in complex environments, achieving high-precision and high-efficiency autofocus effect.
Patent Information
- Application Number
- CN202210271869.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-18
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-03-18
AI Technical Summary
The existing single-cell mass spectrometry system automatic focusing method has poor accuracy and low efficiency in complex environments, and traditional image processing technology is susceptible to environmental factors such as light, resulting in large errors and difficult to meet real-time response requirements.
Using deep learning-based image processing technology, the model is constructed through the Pytorch deep learning framework, including region filtering mechanism, feature extraction network, classifier and distributed label coding, single images are used for block processing and channel attention mechanism screening, high-dimensional semantic features are extracted, and defocusing distance is predicted to achieve automatic focus.
It significantly improves the automatic focus accuracy and efficiency of single-cell mass spectrometry system in complex environments, overcomes the accuracy problems caused by environmental factors in traditional methods, reduces system costs, and meets the real-time response requirements.
Smart Images

Figure CN114764787B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of instrument science, and particularly relates to an automatic focusing method for a single-cell mass spectrometry system based on deep learning. Background Art
[0002] Cells are the most basic structural and functional units of organisms. The heterogeneity among individual cells is crucial for exploring the laws of microscopic life activities. In order to explore single cells more deeply and obtain more accurate cell biological information, it is necessary to conduct measurement research at the single-cell level. At present, single-cell analysis technology has made great breakthroughs in the past decade. As a non-signaturing detection technology with high sensitivity and high specificity, mass spectrometry analysis technology has been widely used in cell metabolomics analysis. Nano-ESI technology is a "soft" ionization technology with high ionization efficiency and high sensitivity. Compared with traditional electrospray ion sources, it consumes less sample, has higher ionization efficiency and more sufficient ionization time. The pulsed direct current nanoelectrospray ionization source technology (Pico-ESI) proposed based on Nano-ESI technology can significantly extend the duration of the mass spectrometry signal. And people often adopt a mass spectrometry analysis method combining droplet microextraction sampling method with Pico-ESI, which realizes high-throughput single-cell metabolomics analysis.
[0003] The key to the above method lies in whether the cell extract can be extracted. Generally, as long as the tip of the pipette reaches the optimal extraction position, the cell extract can be extracted. In order to fully extract the cell extract, a visual servo control system is introduced into the single-cell mass spectrometry system, and it is determined whether the tip of the pipette accurately reaches the optimal extraction position by judging the clarity of the pipette image collected by the imaging module. The better the image focusing effect, the clearer the image, and the closer the pipette is to the optimal extraction position. Thus, this becomes the process of automatic focusing of the imaging system. The focusing effect of the image changes with the movement of the pipette. However, the movement of the pipette is very time-consuming, seriously affecting the operation efficiency of the system. In the early stage, manual operation was adopted, and later traditional digital image processing technology was used to calculate the quality factor of the region of interest (ROI) in the image collected by the imaging module in the servo system to determine the optimal extraction position of the pipette. However, traditional image processing technology is extremely vulnerable to environmental factors such as external light, so there will be large errors and low accuracy. At the same time, the determination of the ROI region of the image and the calculation of the quality factor require a long calculation time, and it is difficult for traditional image processing technology to meet the near-real-time response requirements of the single-cell system. To address the above problem of poor accuracy, a common method is to add additional hardware devices, but this undoubtedly increases the cost of the system. Summary of the Invention
[0004] The first object of the present invention is to propose an automatic focusing method for a single-cell mass spectrometry system based on deep learning in view of the deficiencies of the prior art. This method is based on the image processing technology of deep learning to improve the accuracy and efficiency of the automatic focusing method for the single-cell mass spectrometry system in a complex environment.
[0005] An automatic focusing method for a single-cell mass spectrometry system based on deep learning largely solves the problems of poor automatic focusing accuracy and low efficiency of the system in a complex and variable environment. This method uses the Pytorch deep learning framework to build a model, which mainly consists of three parts: a region filtering mechanism, a feature extraction network, and a distributed label encoding. A single image is used as the input of the model, and the image is divided into blocks according to actual needs. On this basis, a channel attention mechanism is used to screen the region of interest, filter out the background information, and then use the feature extraction network to obtain high-dimensional semantic features. These features are classified and encoded by the distributed label encoding to obtain the finally predicted defocus distance, that is, the distance between the current position and the best imaging position. To solve the problems of data difference and uneven data distribution, the CrossNorm mechanism is adopted in the feature extraction network part, which makes the model effect more stable.
[0006] The method specifically includes the following steps:
[0007] Step 1: Obtain the real-time motion image of the pipette tip in the single-cell mass spectrometry system, and manually label it to construct a data set; part of the data set is used as the training set, and the other part is used as the test set;
[0008] The label is obtained by the following method:
[0009] (1) During the entire focusing process, the corresponding pipette images are obtained at a fixed time interval, and a total of N real-time motion images of the pipette tip at consecutive moments are obtained. In the present invention, N = 40 can be taken.
[0010] (2) Manually select the image with the best focused pipette tip from the above sequence of images, denoted as image x0, and use the position of the pipette tip in this image x0 as the best extraction imaging position.
[0011] Preferably, when the focusing effect of the pipette tip is the best, the position of the pipette tip is located in the middle of the image.
[0012] (3) Take the above image x0 as the reference image, and calculate the defocus distance of each image according to the following formula;
[0013]
[0014] where d n represents the defocus distance of image x n , The moment representing image x0 The moment representing image x n where n ∈ {1, 2, 3, ..., N}, and T represents a fixed interval duration.
[0015] Step 2: Build a Pytorch deep learning model, train it using the training set, and then test it using the test set;
[0016] The Pytorch deep learning model includes a region filtering mechanism module, a feature extraction network, a classifier, and a distributed label encoding module;
[0017] The region filtering mechanism module uses channel attention mechanism to screen the region of interest and filter out background information; specifically:
[0018] (1) Crop the real-time movement image of the pipette tip in the single-cell mass spectrometry system into an integer multiple of the specified image block size according to the center cropping method, and perform grayscale processing on the cropped image to obtain a grayscale image.
[0019] (2) Divide the above grayscale image into blocks, and each image block has the same size; then stack the image blocks along the channel direction to obtain a multi-channel sample data.
[0020] (3) Use the channel attention mechanism to process the obtained multi-channel sample data to obtain a channel mask.
[0021] (4) Use the above channel mask to perform dot multiplication with the multi-channel sample data along the channel direction to obtain the data after background filtering.
[0022] The feature extraction network is used to extract high-dimensional semantic features based on the data after background filtering output by the region filtering mechanism module; specifically, it is based on the MobileNetV2 architecture, and the CrossNorm mechanism is added after each Block in the basic architecture.
[0023] The classifier is used to classify according to the semantic features output by the feature extraction network;
[0024] The distributed label encoding module is used to encode the classification information output by the classifier to obtain the defocus distance of the input image;
[0025] Step 3: Use the trained and tested Pytorch deep learning model to predict the defocus distance of the real-time movement image of the pipette tip in the single-cell mass spectrometry system, and then realize automatic focusing of the single-cell mass spectrometry system.
[0026] Preferably, the feature extraction network is based on MobileNetV2, and the CrossNorm mechanism is added after each bottleneck block in the basic architecture.
[0027] Preferably, the MobileNetV2 is stacked by convolutional blocks, bottleneck blocks, and pooling layers.
[0028] The convolutional blocks in the MobileNetV2 are composed of convolutional operations, BatchNorm operations, and Relu operations; the bottleneck blocks use an inverted residual structure: 1x1 convolution is used for dimensionality increase, 3x3 depthwise separable convolution, and 1x1 convolution for dimensionality reduction.
[0029] Preferably, the defocus distance is expressed as the dot product of the distributed label prediction value and the encoded scale, that is:
[0030]
[0031] In the formula, represents the defocus distance prediction value of the sample data x n , P n represents the K-dimensional encoded scale corresponding to the sample data x n , P n = {p j}, j ∈ {1, 2, 3,..., K}, represents the distributed label prediction value, N represents the number of samples.
[0032] Preferably, the loss function L total used for backpropagation to optimize the model parameters during model training is calculated as follows:
[0033] L total = L kl + L reg
[0034] where the similarity metric loss function L between the distributed label prediction value n and the actual value Y kl is as follows:
[0035]
[0036] In the formula, f represents the KL divergence metric function;
[0037] The difference metric loss function L between the defocus distance prediction value n and the actual value d reg is as follows:
[0038]
[0039] In the formula, g represents the SmoothL1 function.
[0040] The second object of the present invention is an automatic focusing system for a single-cell mass spectrometry system based on deep learning, including a trained and tested Pytorch deep learning model.
[0041] The third object of the present invention is an electronic device, including a processor and a memory, where the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the above method.
[0042] The fourth object of the present invention is a machine-readable storage medium storing machine-executable instructions, which, when called and executed by a processor, cause the processor to implement the above method.
[0043] The present invention has the following beneficial effects:
[0044] The method of the present invention can greatly improve the real-time performance of the system under high-resolution image input. At the same time, compared with traditional methods, the deep learning method can largely overcome the factors that have a greater impact on accuracy in traditional methods, such as complex and variable environmental factors like image noise and illumination. Therefore, the problem of poor accuracy is also largely solved. Description of the Drawings
[0045] Figure 1 It is the overall structure diagram of the model;
[0046] Figure 2 It is the structure diagram of the region filtering mechanism;
[0047] Figure 3 It is the structure diagram of the feature extraction network, where a is the bottleneck block and b is the feature extraction network;
[0048] Figure 4 It is the schematic diagram of distributed label encoding and loss function;
[0049] Figure 5 It is the schematic diagram of distributed labels; where a is the global schematic diagram of the encoding scale, b is the local schematic diagram of the defocus distance -12 encoding, and c is the local schematic diagram of the defocus distance +6 encoding. Detailed Embodiments
[0050] The following further explains the present invention in conjunction with the drawings; however, the protection scope of the present invention is not limited thereto. It should be noted that for processes or symbols not specifically described in detail, those skilled in the art can refer to the prior art for understanding and implementation.
[0051] Based on the Pytorch deep learning framework, according to the actual movement of the pipette in the single-cell mass spectrometry system, a regression model is constructed to predict the defocus distance. The defocus distance is the distance between the current imaging position of the pipette and the optimal imaging position.
[0052] First, artificial labels are assigned to the obtained pipette image data. The process of artificial labeling is as follows: First, in a set of data, the image position corresponding to the optimal extraction position of the pipette is known. This image is used as the reference image x0 for subsequent labeling work, and this reference image is generally located in the middle of a set of data. Then, according to the relative position relationship between other images and the reference image x0 in the same set of data, the defocus distance label of this image is generated, that is, the distance between this image and the reference image. Considering the problem of the pipette movement direction, the positive and negative of the defocus distance are specified. The calculation formula of the defocus distance is as follows:
[0053]
[0054] In the formula, d n represents the defocus distance of the image x n , represents the moment of the image x0, represents the moment of the image x n , n ∈ {1, 2, 3,..., N}, and N represents the number of sampled images in a set of data. According to the formula, the defocus distance of the image on the right side of the reference image is negative, and the defocus distance of the image on the left side of the reference image is positive. The positive and negative of the defocus distance represent different movement directions of the pipette. Thus, the labels of images at different positions can be obtained. On this basis, the data is further divided. 80% of the data is randomly selected from the entire dataset as the training set, and 20% of the data is used as the test set.
[0055] The overall structure of the present invention is as Figure 1 shown, mainly divided into three parts: a region filtering mechanism, a feature extraction network, a classifier, and a distributed label encoding. The region filtering mechanism is as Figure 2 shown. First, the images collected by the imaging module in the system are subjected to basic image transformations, namely central cropping and grayscale conversion. The height and width of the RGB image are centrally cropped from H and W to H' and W', where H' and W' should satisfy:
[0056]
[0057] In the formula, B H , B W respectively represent the height and width of the image block, satisfying B H = BW , N H 、N W respectively represent the number of image patches in the height and width directions of the original image.
[0058] According to the experimental results, the color information in the RGB image has no direct relationship with the quality factor of the image. Therefore, the cropped image is subjected to grayscale transformation to obtain a single-channel grayscale image. The operations of central cropping and grayscale transformation not only reduce the input data volume and the complexity of the model, but also provide a basis for further image processing. Then, the obtained grayscale image is divided into blocks, and the size of each image block is B H ×B W . Stack the above image blocks along the channel direction to obtain multi-channel data, that is, convert the original RGB image data of 1×3×H×W into data. In order to enable the channel attention mechanism to focus on more effective feature information, group convolution is used to separately process the data of each channel. After group convolution, more representative features are obtained. These features are connected to the channel attention mechanism, and thus an output channel mask is obtained. Then, perform a dot product operation on the channel mask and the converted multi-channel data to obtain the data after region filtering. At this time, the data leaves more useful information compared with the converted multi-channel data and filters out most of the useless information, such as background information. The region filtering mechanism enables the subsequent feature extraction network to pay more attention to the information useful for the model, and accelerates the calculation of the model and improves the accuracy of the model to a certain extent.
[0059] The feature extraction network is as shown in Figure 3 (b). This feature extraction network is based on the MobileNetV2 architecture. This network architecture is mainly composed of convolutional blocks, bottleneck blocks, and pooling layers stacked together. Among them, the convolutional block consists of a convolution operation, a BatchNorm operation, and a Relu operation; the bottleneck block uses an inverted residual structure: uses a 1x1 convolution to increase the dimension, a 3x3 depthwise separable convolution, and a 1x1 convolution to reduce the dimension. Because there are large differences between the data, in order to reduce the adverse effects brought by these differences, a CrossNorm mechanism is added after each basic bottleneck, as shown in Figure 3As shown by the dashed box in (a), the addition of CrossNorm is equivalent to increasing the diversity of data during the training phase, thus expanding the data distribution of the training set. Therefore, the robustness of the model is enhanced. The input of the feature extraction network is the output of the region filtering mechanism, and the output of the feature extraction network is a 1×1280×1×1 feature vector (FeatureVector) after adaptive pooling sampling. This vector represents the high-dimensional semantic features of the current input image, which makes the subsequent distributed coding more effective.
[0060] The feature vector output by the feature extraction network is flattened and then input into a classifier with parameter w cls where K is the dimension of the feature vector output by the classifier and also represents the number of distributed coding scales.
[0061] The distributed label coding part is as shown in Figure 4 Adopting distributed coding is mainly to ensure the continuity between data and make the labels used for supervised learning more in line with the actual situation.
[0062] If the true defocus distance corresponding to each sample data x n is d n , n ∈ {1, 2, 3,..., N}, where N represents the number of samples, and d n ∈ [a, b], where a and b are the minimum and maximum values of the actual defocus distance respectively. Divide [a, b] into K - 1 small intervals, and the length of each small interval is S. Therefore, a K-dimensional coding scale P n = {p j}, j ∈ {1, 2, 3,..., K} will be obtained in the interval [a, b], as shown in the last fourth column data in Figure 4 where K = 11. The dimension of this coding scale and its coding interval need to be determined according to the actual data and system parameters. Then d n can be expressed as a linear combination of two scales p i and p j (i ≠ j), that is:
[0063] d n = λ1p i + λ2p j (3)
[0064] In the formula, λ1 and λ2 represent the weight coefficients corresponding to different scales and satisfy the following requirements:
[0065] λ1 + λ2 = 1 (4)
[0066] Similarly, p i and p jIt is also possible to use d n which means:
[0067]
[0068] wherein and respectively represent the floor function and the ceiling function. Therefore, the weights λ1 and λ2 can be calculated by the following formula:
[0069]
[0070] For example Figure 5 as shown, d n = -12 and d n = +6 can be represented using distributed coding as [0,0,0.4,0.6,0,0,0,0,0,0,0] and [0,0,0,0,0,0,0.8,0.2,0,0,0].
[0071] Figure 4 The feature vector output by the classifier in after passing through the Softmax function obtains a probability distribution wherein
[0072]
[0073] The probability distribution represents the predicted value of the distributed label. Thus, the defocus distance can be expressed as the dot product of the predicted value of the distributed label and the coding scale, that is:
[0074]
[0075] In the formula, represents the predicted value of the defocus distance of the nth sample data.
[0076] During the model training process, the loss function used for backpropagation to optimize the model parameters mainly includes the similarity measurement loss function between the predicted value of the distributed label and the actual value and the difference measurement loss function between the predicted value of the defocus distance and the actual value. Among them, the similarity of the distributed label is measured using the KL divergence function, and its corresponding loss value is denoted as L kl , and the smaller the divergence value, the greater the similarity between the two:
[0077]
[0078] where f represents the KL divergence measurement function. The difference of the defocus distance is measured using the SmoothL1 function, denoted as L reg , and the smaller its value, the smaller the difference between the two:
[0079]
[0080] where g represents the SmoothL1 function. Therefore, the final loss function L for model optimization total is expressed as:
[0081] L total = L kl + L reg (11)
[0082] It should be noted that the above embodiments are only used to illustrate the present invention and do not limit the technical solutions described in the present invention. At the same time, although this specification has described the present invention in detail with reference to the above embodiments, those of ordinary skill in the art should understand that the present invention can still be modified or equivalently replaced. Therefore, all technical solutions and their improvements that do not depart from the spirit and scope of the present invention shall be covered by the protection scope of the appended claims of the present invention.
Claims
1. An automatic focusing method for a single-cell mass spectrometry system based on deep learning, characterized in that It includes the following steps: Step 1: Obtain the real-time motion images of the pipette tip in the single-cell mass spectrometry system, manually label them, and construct a dataset; use a part of the dataset as the training set and the other part as the test set; The labels are obtained by the following method: (1) Obtain N consecutive real-time motion images of the pipette tip at fixed time intervals during the entire focusing process; (2) Manually select the image with the best focusing effect of the pipette tip from the images, denoted as image x0, and use the position of the pipette tip in this image x0 as the best extraction imaging position; the best focusing effect of the pipette tip means that the position of the pipette tip is located in the middle of the image; (3) Use the above image x0 as the reference image and calculate the defocus distance of each image according to the following formula; where d n represents the defocus distance of image x n , represents the time of image x0 represents the time of image x n , n ∈ {1, 2, 3,..., N}, and T represents a fixed interval duration; Step 2: Build a Pytorch deep learning model, train it using the training set, and then test it using the test set; The Pytorch deep learning model includes a region filtering mechanism module, a feature extraction network, a classifier, and a distributed label encoding module; The region filtering mechanism module uses a channel attention mechanism to screen the region of interest and filter out background information; Specifically: (1) Crop the real-time motion images of the pipette tip in the single-cell mass spectrometry system to an integer multiple of the specified image block size in the way of central cropping, and perform grayscale processing on the cropped images to obtain grayscale images; (2) Perform block processing on the above grayscale images, and each image block has the same size; then stack the image blocks along the channel direction to obtain a multi-channel sample data; (3) Use the channel attention mechanism to process the obtained multi-channel sample data to obtain a channel mask; (4) Use the above channel mask to perform dot multiplication with the multi-channel sample data along the channel direction to obtain the data after background filtering; The feature extraction network is used to extract high-dimensional semantic features according to the data after background filtering output by the region filtering mechanism module; The classifier is used to classify according to the semantic features output by the feature extraction network; The distributed label encoding module is used to encode the classification information output by the classifier to obtain the defocus distance prediction value of the input image Step 3: Use the trained and tested Pytorch deep learning model to realize the prediction of the defocus distance of the real-time motion images of the pipette tip in the single-cell mass spectrometry system, and then realize the automatic focusing of the single-cell mass spectrometry system.
2. The method according to claim 1, wherein The feature extraction network is based on MobileNetV2, and the CrossNorm mechanism is added after each bottleneck block in the basic architecture.
3. The method according to claim 2, wherein The MobileNetV2 is stacked by convolutional blocks, bottleneck blocks, and pooling layers.
4. The method according to claim 3, wherein The convolutional blocks in the MobileNetV2 are composed of convolutional operations, BatchNorm operations, and Relu operations; the bottleneck blocks use an inverted residual structure: use a 1x1 convolution to increase the dimension, a 3x3 depthwise separable convolution, and a 1x1 convolution to reduce the dimension.
5. The method according to claim 1, wherein Defocus distance Expressed as the dot product of the distributed label prediction value and the encoding scale, i.e.: In formula (8), represents the defocus distance prediction value of the sample data x n , and P n represents the K-dimensional coding scale corresponding to the sample data x n , and P n = {p j}, j ∈ {1, 2, 3,..., K}, represents the distributed label prediction value, and N represents the number of samples.
6. The method according to claim 1, wherein The loss function L used to backpropagate and optimize model parameters during model training total is calculated as follows: L total = L kl + L reg Among them, the distributed label prediction value and the actual value Y n The similarity metric loss function L kl is as follows: In formula (9), f represents the KL divergence metric function; Defocus distance predicted value The difference metric loss function L n from the actual value d reg is as follows: In formula (10), g represents the SmoothL1 function.
7. A system using the automatic focusing method of the single-cell mass spectrometry system based on deep learning according to any one of claims 1-6, characterized in that It includes the trained and tested Pytorch deep learning model.
8. An electronic device, characterized in that, It includes a processor and a memory. The memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the method according to any one of claims 1-6.
9. A machine-readable storage medium, characterized in that, The machine-readable storage medium stores machine-executable instructions. When the machine-executable instructions are called and executed by a processor, the machine-executable instructions cause the processor to implement the method according to any one of claims 1-6.
Citation Information
Patent Citations
Attention mechanism relationship comparison network model method based on small sample learning
CN110020682A
Blood cell microscopic image classification method based on regional confusion mechanism neural network
CN111860406A