Scalable imaging flow cytometer based on self-supervised deep learning

By employing self-supervised deep learning, and utilizing the convolution operation between the mask and the cell motion direction, along with backpropagation of the difference values, real-time high-precision two-dimensional image reconstruction under unlabeled data conditions was achieved. This solves the problems of high data acquisition cost and insufficient scalability of existing imaging flow cytometry cell sorters, and improves the system's adaptability and sorting efficiency.

CN122492881APending Publication Date: 2026-07-31JIANGSU SAIDELI PHARMA MACHINE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGSU SAIDELI PHARMA MACHINE
Filing Date
2026-05-19
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing imaging flow cytometers struggle to achieve real-time, high-precision two-dimensional image reconstruction without labeled training data. Furthermore, the system lacks scalability for different cell types and flow velocities, making it difficult to meet the demands for high-throughput, high-resolution cell sorting.

Method used

A self-supervised deep learning approach is adopted. The one-dimensional temporal light intensity signal sequence of moving cells in the microfluidic channel after being transformed from the spatial domain to the temporal domain by a mask is input into a deep convolutional reconstruction network. The network parameters are updated by using the convolution operation of the mask and the direction of cell movement and the backpropagation of the difference value. This achieves training without manually labeled data. Combined with image feature extraction and classifier output of cell category labels, the sorting execution mechanism is triggered.

Benefits of technology

It significantly reduces data acquisition costs, improves reconstruction accuracy and system scalability, and can adapt to different cell types and mask patterns to meet the needs of real-time high-precision cell sorting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122492881A_ABST
    Figure CN122492881A_ABST
Patent Text Reader

Abstract

This invention discloses a scalable imaging flow cytometry cell sorting instrument based on self-supervised deep learning, belonging to the field of biomedical detection and imaging technology. It includes a self-supervised learning module for performing the following: inputting a one-dimensional temporal light intensity signal sequence obtained by detecting moving cells in a microfluidic channel after a mask spatial-temporal transformation into the input domain into an initialized deep convolutional reconstruction network, outputting a reconstructed two-dimensional cell image; performing a convolution operation between the reconstructed two-dimensional cell image and the mask along the cell motion direction to generate a predicted one-dimensional temporal light intensity signal sequence; calculating the difference between this predicted and actual signals and backpropagating to update the network parameters until convergence to obtain a fully trained network; and inputting the actual one-dimensional temporal light intensity signal sequence of the cells to be sorted into the fully trained network to generate a two-dimensional reconstructed image of the cells to be sorted. This invention can train the reconstruction network without labeled data, is scalable, and has a fast imaging speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biomedical detection and imaging technology, specifically to a scalable imaging flow cytometer for cell sorting based on self-supervised deep learning. Background Technology

[0002] Imaging flow cytometry cell sorting systems can simultaneously acquire flow cytometry cell morphology images and multi-parameter fluorescence information, making them important for rare cell detection and tumor circulating cell analysis. Existing imaging flow cytometry cell sorting systems typically employ line scanning or area array cameras, which are limited by camera frame rate and data transmission bandwidth, making it difficult to simultaneously achieve high throughput and high resolution imaging. To improve imaging speed, a compressed sensing scheme using mask spatial-temporal domain conversion has been proposed: by placing a mask with a random array of light apertures on the microscope image plane, the spatial distribution of the two-dimensional cell image is converted into a one-dimensional temporal light intensity signal, and then a reconstruction algorithm is used to recover the two-dimensional image from the one-dimensional signal. However, traditional reconstruction methods (such as compressed sensing iterative algorithms) are computationally intensive and slow, making it difficult to meet real-time sorting requirements. Deep learning-based reconstruction methods can improve computational efficiency, but existing supervised learning methods rely on a large number of pairs of real two-dimensional images and corresponding one-dimensional signals as training data, making the acquisition of real two-dimensional images difficult and costly. Furthermore, errors exist in the design of the mask pattern and the physical model of the convolution operation, affecting reconstruction accuracy. This invention focuses on solving the problem of real-time high-precision image reconstruction under unlabeled training data conditions, and improving the scalability of the system for different cell types and flow velocities. Summary of the Invention

[0003] The purpose of this invention is to provide a scalable imaging flow cytometry cell sorter based on self-supervised deep learning, in order to solve the technical problem that existing imaging flow cytometry cell sorters are unable to achieve real-time, high-precision two-dimensional image reconstruction under the condition of unlabeled training data, while improving the scalability of the system for different cell types and flow velocities, thereby meeting the requirements of high-throughput, high-resolution cell sorting.

[0004] The objective of this invention can be achieved through the following technical solutions: This invention provides a scalable imaging flow cytometer for cell sorting based on self-supervised deep learning. The sorter includes a self-supervised learning module, which performs the following: A one-dimensional temporal light intensity signal sequence, obtained by detecting moving cells within a microfluidic channel after spatial-temporal transformation via a mask, is input into an initialized deep convolutional reconstruction network. The deep convolutional reconstruction network outputs a reconstructed two-dimensional cell image. The reconstructed two-dimensional cell image is then convolved with the mask along the cell movement direction to generate a predicted one-dimensional temporal light intensity signal sequence. The difference between the predicted and predicted one-dimensional temporal light intensity signal sequences is calculated, and the network parameters of the deep convolutional reconstruction network are updated using backpropagation based on this difference until the difference converges to the predicted value. The update process stops when a threshold is set, resulting in a fully trained deep convolutional reconstruction network. The actual one-dimensional temporal intensity signal sequence of the cells to be sorted is input into the trained deep convolutional reconstruction network to generate corresponding two-dimensional reconstructed images of the cells. Image feature extraction is performed on the two-dimensional reconstructed images of the cells to be sorted to obtain cell morphology feature vectors and cell fluorescence intensity distribution feature vectors. These feature vectors are then input into a classifier, which outputs cell category labels and triggers the sorting mechanism downstream of the microfluidic channel based on these labels. This self-supervised learning approach eliminates the need for manually labeled large amounts of real cell images as training data. The deep convolutional reconstruction network can be automatically trained using only the physical convolutional relationship between the one-dimensional temporal intensity signal and the mask, significantly reducing data acquisition costs. Furthermore, it can adapt to different cell types or mask patterns by adjusting network parameters without changing the hardware structure, demonstrating good scalability.

[0005] As a technical solution of the present invention, before inputting the one-dimensional temporal intensity signal sequence obtained by detecting the moving cells in the microfluidic channel after spatial-temporal transformation via a mask into the initialized deep convolutional reconstruction network, the following operations are performed: a continuous laser beam is reflected by a first bipolar mirror and irradiated onto the cell flow area in the microfluidic channel, exciting the cells to generate fluorescence signals; the fluorescence signals are collected by an objective lens and imaged onto a mask placed on the objective mirror surface, the mask having a randomly distributed array of light-transmitting apertures, the pattern of which is convolved with the two-dimensional image of the cells in the direction of cell movement, converting the spatial distribution of the two-dimensional image of the cells into a one-dimensional temporal intensity distribution; the one-dimensional temporal intensity distribution is then split by wavelength by a second bipolar mirror to multiple single-photon detector arrays, which simultaneously acquire the one-dimensional temporal intensity signal sequence. This solution utilizes the spatial coding characteristics of the mask to achieve parallel compressed sampling of the fluorescence signal, and the high sensitivity of the single-photon detector arrays ensures reliable detection of weak fluorescence signals, providing high-quality one-dimensional temporal input for subsequent deep learning reconstruction.

[0006] Preferably, the array of light-transmitting holes on the mask is distributed in a modified uniform random array, and the projection distances of any two light-transmitting holes in the cell motion direction are different. This design ensures that there is a unique correlation between the columns of the mask convolution kernel, which reduces the ill-conditioning when reconstructing a two-dimensional image from a one-dimensional temporal signal and improves the convergence speed and reconstruction accuracy of the deep convolutional reconstruction network.

[0007] Furthermore, before the continuous laser beam is reflected by the first bipolar mirror and irradiated onto the cell flow area within the microfluidic channel, the following operation is performed: a third laser source is added to the transmission path of the first bipolar mirror. The third laser beam emitted by the third laser source is transmitted through the first bipolar mirror and then combined with the continuous laser beam to jointly irradiate the cell flow area, achieving multi-wavelength excitation. Multi-wavelength excitation can simultaneously acquire cell information with different fluorescent labels, broadening the multi-parameter analysis capability of cell sorting. Moreover, the beam-combining optical path avoids the introduction of additional optical components, resulting in a compact structure.

[0008] In a specific embodiment of the present invention, before performing a convolution operation between the reconstructed two-dimensional cell image and the mask in the cell motion direction, the following operations are performed: A binary matrix representation of the aperture pattern of the mask is obtained, where the number of rows in the binary matrix corresponds to the number of pixels along the length of the mask in the cell motion direction, and the number of columns in the binary matrix corresponds to the number of pixels perpendicular to the width of the mask in the cell motion direction. Elements at the aperture positions in the binary matrix have a value of one, and elements at the non-aperture positions have a value of zero. The reconstructed two-dimensional cell image is represented as a two-dimensional grayscale matrix, where the number of rows in the two-dimensional grayscale matrix corresponds to the number of pixels along the height of the image in the cell motion direction, and the number of columns in the two-dimensional grayscale matrix corresponds to the number of pixels perpendicular to the width of the image in the cell motion direction. The binary matrix is ​​used as the convolution kernel, and a discrete convolution operation with a stride of one is performed along the row direction of the two-dimensional grayscale matrix to obtain the predicted one-dimensional temporal intensity signal sequence. This convolution operation accurately reproduces the linear projection relationship between the physical mask and the cell image, ensuring that the forward process in self-supervised training is completely consistent with physical reality and guaranteeing the accuracy of gradient backpropagation. Specifically, the binary matrix is ​​used as the convolution kernel, and a discrete convolution operation with a stride of one is performed along the row direction of the two-dimensional grayscale matrix to obtain the predicted one-dimensional temporal intensity signal sequence, including the following calculations: Let the two-dimensional grayscale matrix of the reconstructed two-dimensional cell image be... ,in The height of the image along the direction of cell movement is represented by the number of pixels. The width in pixels is perpendicular to the direction of cell movement; the binary matrix of the light-transmitting aperture pattern of the mask is... ,in The length of the mask in pixels along the direction of cell movement, and ; The predicted one-dimensional time-series light intensity signal sequence The Each element is calculated using the following formula: in: Represents the first element in the two-dimensional grayscale matrix. line, number The element values ​​of the column, Represents the first in the binary matrix line, number The column elements are either 0 or 1, the convolution stride is one, and zero-padding is not performed. This calculation method has a clear physical meaning and facilitates efficient batch convolution operations in deep learning frameworks.

[0009] Preferably, the step of calculating the difference between the one-dimensional temporal light intensity signal sequence and the predicted one-dimensional temporal light intensity signal sequence, and using the difference to backpropagate and update the network parameters of the deep convolutional reconstruction network, includes: subtracting the light intensity values ​​at the same time index position in the one-dimensional temporal light intensity signal sequence and the predicted one-dimensional temporal light intensity signal sequence point by point, squaring the results, and then summing them to obtain the difference value; calculating the gradient of the difference value with respect to the feature map of the output layer of the deep convolutional reconstruction network to obtain the output layer gradient tensor; backpropagating the output layer gradient tensor layer by layer to the input layer of the deep convolutional reconstruction network, calculating the gradient of the parameters of that layer during each layer's propagation, and updating the parameters of that layer according to the gradient descent direction. Through the regularization of the point-by-point squared error sum, the network not only fits the signal waveform during training but also suppresses overfitting. Specifically, the difference value is obtained by subtracting the light intensity values ​​at the same time index position in the one-dimensional time-series light intensity signal sequence from the predicted one-dimensional time-series light intensity signal sequence point by point, squaring the results, and then summing them. The difference value is calculated using the following loss function: in: The total length of the one-dimensional temporal light intensity signal sequence; For discrete-time indexing; The one-dimensional temporal intensity signal sequence predicted by the deep convolutional reconstruction network is in the... The light intensity value at each time index; The one-dimensional time-series light intensity signal sequence actually detected in the first... The light intensity value at each time index; is the regularization coefficient, and its value range is . to ; The total number of all weight parameters in the deep convolutional reconstruction network; For the deep convolutional reconstruction network, the first... The values ​​of each weight parameter; The calculated difference value is the loss function value. The L1 regularization term promotes the sparsity of network weights, which helps improve the quality of the reconstructed image and accelerates convergence.

[0010] In a preferred embodiment of the present invention, image feature extraction is performed on the two-dimensional reconstructed image of the cells to be sorted to obtain a cell morphology feature vector and a cell fluorescence intensity distribution feature vector. This includes: performing threshold segmentation on the two-dimensional reconstructed image of the cells to be sorted; extracting the foreground pixel set of the cell region; calculating the equivalent diameter, eccentricity, and ratio of the convex hull area to the foreground area of ​​the foreground pixel set; and concatenating the equivalent diameter, eccentricity, and ratio in sequence to form the cell morphology feature vector. The grayscale value of each pixel in the two-dimensional reconstructed image of the cells to be sorted is normalized to obtain a normalized grayscale value matrix; and the normalized grayscale value matrix is ​​expanded row-wise into a one-dimensional vector, which serves as the cell fluorescence intensity distribution feature vector. This feature extraction method takes into account both global morphological parameters and local grayscale distribution information, and can effectively distinguish cell subpopulations with similar morphology but different fluorescent labels.

[0011] Furthermore, after performing threshold segmentation on the two-dimensional reconstructed image of the cells to be sorted and extracting the foreground pixel set of the cell region, the following operations are performed: eight-neighbor tracking is performed on the boundary of the foreground pixel set to obtain the pixel coordinate sequence of the cell contour; the Fourier descriptor of the pixel coordinate sequence of the cell contour is calculated, and the first few low-frequency Fourier coefficients are truncated to form a Fourier feature vector; the Fourier feature vector is appended to the end of the cell morphology feature vector to form an expanded cell morphology feature vector. The Fourier descriptor can capture the fine shape details of the contour and is invariant to rotation, scaling, and translation, significantly improving the robustness of the classifier in recognizing aberrant cells. Specifically, calculating the Fourier descriptor of the pixel coordinate sequence of the cell contour and truncating the first few low-frequency Fourier coefficients to form a Fourier feature vector specifically includes: representing the pixel coordinate sequence of the cell contour in complex coordinate form. ,in , This represents the total number of outline pixels. and The first The x and y coordinates of each contour pixel The imaginary unit; for Perform Discrete Fourier Transform, Fourier Descriptor The calculation formula is: in: For frequency domain indexing, ; For the first A Fourier descriptor, whose modulus is Indicates the contour at frequency The shape and energy below; before truncation The low-frequency Fourier coefficients constitute the Fourier feature vector. ,in The value is , This indicates a floor operation, discarding the DC component. Low-frequency components correspond to the overall shape of the contour, while high-frequency noise components are discarded, thereby reducing feature dimensionality while maintaining discriminative power.

[0012] Preferably, the cell morphology feature vector and the cell fluorescence intensity distribution feature vector are input into a classifier, and the classifier outputs a cell category label. This includes: concatenating the cell morphology feature vector and the cell fluorescence intensity distribution feature vector along the feature dimension to obtain a joint feature vector; inputting the joint feature vector into a support vector machine (SVM) classifier, which is trained using the joint feature vector from multiple pre-collected labeled cell categories; the SVM classifier outputs the classification confidence score of the joint feature vector relative to each preset cell category; and selecting the preset cell category with the highest classification confidence score as the cell category label. Support vector machines have good generalization ability in high-dimensional feature spaces and require a low number of training samples, allowing for reliable classification results with only a small number of labeled cell samples.

[0013] As a specific embodiment of the present invention, triggering a sorting execution mechanism downstream of the microfluidic channel based on the cell category tag includes: acquiring the movement velocity of the cells to be sorted within the microfluidic channel in real time, and calculating a delay time based on the movement velocity and the distance the cells to be sorted travel from the imaging area to the location of the sorting execution mechanism; starting from the moment the cell category tag is output, timing is initiated, and when the timing reaches the delay time, a trigger pulse signal is sent to the sorting execution mechanism, wherein the pulse width of the trigger pulse signal is greater than the action response time of the sorting execution mechanism. This delayed triggering mechanism ensures that the cells are captured by the sorting execution mechanism precisely when they reach the sorting position, avoiding missorting caused by timing misalignment and improving sorting purity and recovery rate.

[0014] Furthermore, after inputting the actual one-dimensional temporal intensity signal sequence of the cells to be sorted into the fully trained deep convolutional reconstruction network to generate the corresponding two-dimensional reconstructed image of the cells to be sorted, the following operations are performed: inputting the two-dimensional reconstructed image of the cells to be sorted into an auxiliary autoencoder, which has the same encoder structure as the deep convolutional reconstruction network but different decoder parameters, and outputting an auxiliary reconstructed image; calculating the structural similarity index between the two-dimensional reconstructed image of the cells to be sorted and the auxiliary reconstructed image; when the structural similarity index is lower than a preset similarity threshold, marking the two-dimensional reconstructed image of the cells to be sorted as a low-confidence image and triggering the re-acquisition of the actual one-dimensional temporal intensity signal sequence. By introducing an auxiliary autoencoder with the same structure as the original network encoder but an independent decoder, the reconstructed image is validated a second time. The structural similarity index reflects the confidence level of the image quality. When the reconstruction quality is unreliable, the signal is re-acquired, effectively avoiding erroneous sorting decisions caused by transient optical noise or abnormal cell movement, and improving the overall reliability of the system.

[0015] The beneficial effects of this invention are: By inputting the one-dimensional temporal intensity signal sequence obtained from the detection of moving cells within a microfluidic channel after a mask spatial-temporal transformation into the initial deep convolutional reconstruction network, a reconstructed two-dimensional cell image is output. This reconstructed two-dimensional cell image is then convolved with the mask along the cell motion direction to generate a predicted one-dimensional temporal intensity signal sequence. The difference between the actual and predicted one-dimensional signals is calculated, and this difference is used to backpropagate and update the network parameters of the deep convolutional reconstruction network until convergence. This achieves self-supervised training without any real two-dimensional image annotation data. This self-supervised training mechanism directly utilizes the physical imaging process (mask convolution) to generate the supervision signal, avoiding the cumbersome process of manually annotating a large number of two-dimensional cell images or acquiring paired data through other high-resolution imaging devices, as required by traditional supervised learning methods, significantly reducing data acquisition costs. Furthermore, since the training signal originates from the system's own physical model, the network can naturally adapt to systematic errors such as mask pattern deviations and uneven illumination, resulting in higher reconstruction accuracy. In addition, this self-supervised framework does not rely on prior knowledge of specific cell types; after changing the microfluidic channel or mask, only a small amount of one-dimensional signal needs to be collected to retrain the network, demonstrating good scalability. The specific method for generating a predicted one-dimensional temporal intensity signal sequence by convolving the reconstructed two-dimensional cell image with a mask is as follows: A binary matrix representation of the mask's aperture pattern is obtained. The number of rows in this binary matrix corresponds to the length (in pixels) of the mask along the cell's movement direction, and the number of columns corresponds to the width (in pixels) perpendicular to the cell's movement direction. Elements at aperture positions have a value of one, while elements at non-aperture positions have a value of zero. The reconstructed two-dimensional cell image is represented as a two-dimensional grayscale matrix. The number of rows in this grayscale matrix corresponds to the height (in pixels) of the image along the cell's movement direction, and the number of columns corresponds to the width (in pixels) perpendicular to the cell's movement direction. The binary matrix is ​​used as the convolution kernel, and a discrete convolution operation with a stride of one is performed along the row directions of the grayscale matrix to obtain the predicted one-dimensional temporal intensity signal sequence. This convolution operation rigorously simulates the modulation process of the moving cell by the physical mask, ensuring a high degree of consistency between the predicted one-dimensional signal and the actual one-dimensional signal in the temporal domain, thus providing precise physical constraints for self-supervised learning. Compared to directly using neural networks to fit the mapping from one-dimensional signals to two-dimensional images (i.e., black-box learning), the reconstruction network constrained by this physical model can learn a more accurate inverse transform, and the generated two-dimensional cell images maintain higher fidelity in details such as cell morphology and fluorescence intensity distribution. Furthermore, since the convolution kernel (mask matrix) is known and fixed, this scheme does not require the introduction of additional complex physical simulation modules, resulting in high computational efficiency and easy deployment in real-time sorting systems. Attached Figure Description

[0016] The invention will now be further described with reference to the accompanying drawings.

[0017] Figure 1This is a diagram showing the working state of the scalable imaging flow cytometer based on self-supervised deep learning described in this invention. Figure 2 This is a flowchart of the multi-wavelength fluorescence signal acquisition and one-dimensional time-series conversion process; Figure 3 This is a flowchart of the joint feature vector extraction and support vector machine classification output process. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] See Figure 1 This invention provides a scalable imaging flow cytometer for cell sorting based on self-supervised deep learning. The sorter includes a self-supervised learning module, which performs the following: A one-dimensional temporal light intensity signal sequence, obtained by detecting moving cells within a microfluidic channel after spatial-temporal transformation via a mask, is input into an initialized deep convolutional reconstruction network. The deep convolutional reconstruction network outputs a reconstructed two-dimensional cell image. The reconstructed two-dimensional cell image is convolved with the mask along the cell movement direction to generate a predicted one-dimensional temporal light intensity signal sequence. The difference between the predicted one-dimensional temporal light intensity signal sequence and the predicted one-dimensional temporal light intensity signal sequence is calculated, and the network parameters of the deep convolutional reconstruction network are updated using the difference value through backpropagation until the difference value converges to a preset value. The update stops when a threshold is reached, resulting in a fully trained deep convolutional reconstruction network. The actual one-dimensional temporal light intensity signal sequence of the cells to be sorted is input into the fully trained deep convolutional reconstruction network to generate a corresponding two-dimensional reconstruction image of the cells to be sorted. Image feature extraction is performed on the two-dimensional reconstruction image of the cells to be sorted to obtain cell morphology feature vectors and cell fluorescence intensity distribution feature vectors. The cell morphology feature vectors and cell fluorescence intensity distribution feature vectors are input into a classifier, which outputs cell category labels and triggers the sorting execution mechanism downstream of the microfluidic channel based on the cell category labels.

[0020] Example 1: In specific implementation, refer to Figure 2A continuous laser beam, reflected by a first bipolar mirror, illuminates the cell-flowing region within the microfluidic channel, exciting the cells to produce fluorescence signals. The continuous laser beam is emitted by a continuous laser generator, and the wavelength of the laser produced is selected based on the fluorescent dye used to label the cells. The first bipolar mirror has high reflectivity for the wavelength of the continuous laser beam and high transmittance for the wavelength of the fluorescence signal. Cells within the microfluidic channel sequentially pass through the cell-flowing region under the focusing effect of the fluid; this region is located near the focal plane of the objective lens.

[0021] Fluorescence signals are collected by an objective lens and imaged onto a mask placed on the objective lens. The mask has a randomly distributed array of apertures. The pattern of the aperture array is convolved with the two-dimensional image of the cell along the cell movement direction, converting the spatial distribution of the two-dimensional image into a one-dimensional temporal intensity distribution. The objective lens is a high numerical aperture objective lens with a numerical aperture greater than 0.8. The mask is a planar element made of opaque material, and the aperture array is fabricated on the mask using photolithography. The apertures are circular through-holes with a diameter less than 5 micrometers. The distribution of the aperture array satisfies the condition that the projected distances of any two apertures along the cell movement direction are different. This distribution is generated using a modified uniform random array algorithm. This algorithm generates initially uniformly randomly distributed aperture coordinates within the mask plane. The projected distances of any two apertures along the cell movement direction are detected. If equal projection distances exist, the position of one aperture is adjusted to make the projection distances unequal. This process is repeated until all projection distances are different. The area covered by the light-perforated array corresponds to the image field size of the area through which the cells flow on the objective plane.

[0022] A one-dimensional temporal intensity distribution is split wavelength-wise by a second bichromatic mirror and distributed to multiple single-photon detector arrays. These arrays then synchronously acquire the one-dimensional temporal intensity signal sequence. The second bichromatic mirror, located behind a mask, splits the one-dimensional temporal intensity distribution into different optical paths according to different wavelength ranges. Each single-photon detector array contains at least two single-photon detectors, each corresponding to a fluorescence wavelength channel. The single-photon detectors are either single-photon avalanche diodes or photomultiplier tubes, possessing picosecond-level time resolution. Multiple single-photon detector arrays acquire the signal synchronously, with the synchronization signal provided by an external clock source and a sampling rate of at least 1 MHz. The one-dimensional temporal intensity signal sequence is a digital signal sequence; an analog-to-digital converter converts the pulse counts output by the single-photon detectors into digital intensity values.

[0023] In specific implementation, a third laser source is added to the transmission path of the first bipolar mirror. The third laser beam emitted from the third laser source is transmitted through the first bipolar mirror and then combined with the continuous laser beam to jointly irradiate the cell flow area, achieving multi-wavelength excitation. The third laser source is a semiconductor laser or a solid-state laser, and its wavelength is different from that of the continuous laser beam. After being transmitted through the first bipolar mirror, the third laser beam emitted from the third laser source spatially overlaps with the reflected continuous laser beam, forming a collinear combined beam. The combined laser beam is focused by a focusing lens and then irradiates the cell flow area to simultaneously excite different fluorescent dyes on the cells, generating multi-wavelength fluorescence signals. After being collected by the objective lens, the multi-wavelength fluorescence signals are modulated by a mask and then split by the second bipolar mirror according to wavelength to the corresponding single-photon detector array.

[0024] Example 2: In a specific implementation, the binary matrix representation of the mask's aperture pattern is obtained. The number of rows in the binary matrix corresponds to the length (in pixels) of the mask along the cell movement direction, and the number of columns corresponds to the width (in pixels) of the mask perpendicular to the cell movement direction. Elements at the aperture positions in the binary matrix are set to one, while elements at non-aperture positions are set to zero. The binary matrix of the aperture pattern is obtained as follows: Based on the coordinates of the apertures in the mask design file, a zero-matrix is ​​generated, and the element value at each aperture's corresponding pixel position is set to one. The length (in pixels) of the mask along the cell movement direction... and width in pixels The size of the mask is determined by its actual physical dimensions and the pixel resolution of the imaging system. For example, if the physical dimensions of the mask are 100 micrometers × 100 micrometers and the pixel resolution is 1 micrometer / pixel, then... , A binary matrix is ​​denoted as... .

[0025] The reconstructed two-dimensional cell image is represented as a two-dimensional grayscale matrix. The number of rows in the two-dimensional grayscale matrix corresponds to the number of pixels at height along the cell movement direction, and the number of columns corresponds to the number of pixels at width perpendicular to the cell movement direction. The reconstructed two-dimensional cell image is output by a deep convolutional reconstruction network, and its size is fixed during training or inference. The deep convolutional reconstruction network is a convolutional neural network, consisting of an input layer, an encoding module, and a decoding module. The input layer receives a one-dimensional temporal light intensity signal sequence, the length of which is... The encoding module contains three one-dimensional convolutional layers, each followed by a batch normalization layer and a ReLU activation function, compressing the one-dimensional signal into a latent feature vector. The decoding module contains three two-dimensional transposed convolutional layers, each followed by a batch normalization layer and a ReLU activation function, upsampling the latent feature vector into a two-dimensional grayscale matrix. ,in The height of the image along the direction of cell movement is represented by the number of pixels. This represents the width in pixels perpendicular to the direction of cell movement. The element values ​​of the two-dimensional grayscale matrix represent the grayscale value at the corresponding pixel location, normalized to between 0 and 1.

[0026] Using a binary matrix as the convolution kernel, a discrete convolution operation with a stride of one is performed along the row directions of the two-dimensional grayscale matrix to obtain the predicted one-dimensional temporal intensity signal sequence. The specific calculation process is as follows: Let the two-dimensional grayscale matrix of the reconstructed two-dimensional cell image be... The binary matrix of the aperture pattern of the mask is And satisfy Predicting a one-dimensional temporal light intensity signal sequence The Each element is calculated using the following formula: in Represents a two-dimensional grayscale matrix The Middle line, number The element values ​​of the column. Representing a binary matrix The Middle line, number The column's element values ​​are either 0 or 1. The convolution stride is one, and no zero-padding is performed during the convolution process. Therefore, the length of the output sequence is... .

[0027] The difference between a one-dimensional temporal light intensity signal sequence and a predicted one-dimensional temporal light intensity signal sequence is calculated, and the network parameters of the deep convolutional reconstruction network are updated using backpropagation based on this difference. The specific steps include: subtracting the light intensity values ​​at the same time index position in the one-dimensional and predicted one-dimensional temporal light intensity signal sequences point by point, squaring the results, and then summing them to obtain the difference; calculating the gradient of the difference with respect to the feature map of the output layer of the deep convolutional reconstruction network to obtain the output layer gradient tensor; backpropagating the output layer gradient tensor layer by layer to the input layer of the deep convolutional reconstruction network, calculating the gradient of the parameters of that layer during each propagation, and updating the parameters of that layer according to the gradient descent direction. The difference is calculated using the following loss function: in: The total length of the one-dimensional temporal light intensity signal sequence; This is a discrete-time index, with values ​​ranging from 1 to... ; The one-dimensional temporal intensity signal sequence predicted by the deep convolutional reconstruction network is in the... The light intensity value at each time index is the predicted sequence obtained by the above convolution operation; The actual detected one-dimensional time-series light intensity signal sequence is in the first... The light intensity value at each time index; is the regularization coefficient, and its value range is . to The specific values ​​are set before training, for example, set to... ; The total number of all weight parameters in the deep convolutional reconstructed network, including the weights of all convolutional and fully connected layers; For the reconstruction of the deep convolutional network, the first The values ​​of each weight parameter; The calculated difference value is the loss function value. The first term in the loss function is the mean squared error term, and the second term is the L1 regularization term, used to constrain the sparsity of the network weights. In each iteration, according to... Calculate the gradient and update all trainable parameters in the deep convolutional reconstruction network using an optimizer (such as the Adam optimizer with a learning rate of 0.001). Repeat the iteration until the loss function value converges to a preset threshold (e.g., ...). When the update stops, a fully trained deep convolutional reconstruction network is obtained.

[0028] Example 3: In a specific implementation, a threshold segmentation operation is performed on the two-dimensional reconstructed image of the cells to be sorted to extract the foreground pixel set of the cell regions. The threshold segmentation operation uses a global thresholding method, classifying pixels in the two-dimensional reconstructed image of the cells to be sorted with gray values ​​greater than a preset threshold into the foreground pixel set. The preset threshold is calculated using the Otsu algorithm. The Otsu algorithm calculates the gray value that maximizes the inter-class variance as the preset threshold. The foreground pixel set contains the pixel coordinates and gray values ​​of all pixels identified as cell regions.

[0029] Calculate the equivalent diameter of the foreground pixel set. The equivalent diameter is defined as the diameter of a circle with the same area as the foreground pixel set. The formula is: ,in This is the area of ​​the foreground pixel set, which is the number of foreground pixels multiplied by the physical area represented by a single pixel. The physical area represented by a single pixel is determined by the pixel resolution of the imaging system. For example, if the pixel resolution is 0.5 micrometers / pixel, then the area of ​​a single pixel is 0.25 square micrometers.

[0030] The eccentricity of the foreground pixel set is calculated using the covariance matrix of the foreground pixel set. (Covariance matrix...) The elements are: in: This represents the total number of foreground pixels. , The first The x and y coordinates of each foreground pixel. , These are the mean values ​​of the x and y coordinates of all foreground pixels, respectively. The eigenvalues ​​of the covariance matrix are... , The eccentricity is obtained by solving the characteristic equation. Defined as ,in .

[0031] Calculate the ratio of the convex hull area of ​​the foreground pixel set to the foreground area. The convex hull area is obtained by calculating the area of ​​the convex polygon of the coordinate set of all foreground pixels. The Graham scan algorithm is used to obtain the vertex sequence of the convex hull, and then the polygon area formula (Green's function method) is used to calculate the convex hull area. The foreground area is then... The ratio of the convex hull area to the foreground area is denoted as... .

[0032] The equivalent diameter, eccentricity, and the ratio of convex hull area to foreground area are concatenated in order to form a cell morphology feature vector. .

[0033] The grayscale value of each pixel in the 2D reconstructed image of the cells to be sorted is normalized to obtain a normalized grayscale value matrix. The normalization process uses the maximum-minimum normalization method, and the formula is as follows: in: The first cell in the two-dimensional reconstructed image of the cells to be sorted line, number The original grayscale value of the column pixel. and These are the minimum and maximum grayscale values ​​in the entire image, respectively. The normalized gray values ​​range from 0 to 1. The normalized gray value matrix is ​​expanded row-wise into a one-dimensional vector, which serves as the feature vector for the cell fluorescence intensity distribution. The vector length is .

[0034] In a specific implementation, eight-neighbor tracing is performed on the boundaries of the foreground pixel set to obtain the pixel coordinate sequence of the cell contour. The eight-neighbor tracing algorithm starts from the edge pixels of the foreground pixel set and searches for the next contour pixel sequentially along the eight neighbor directions (top, bottom, left, right, and four diagonal directions) until it returns to the starting pixel. The coordinates of all contour pixels are recorded. ,in , This represents the total number of outline pixels.

[0035] Calculate the Fourier descriptor of the pixel coordinate sequence of the cell outline. Represent the pixel coordinate sequence of the cell outline in complex coordinate form. ,in The imaginary unit, .right Perform Discrete Fourier Transform, Fourier Descriptor The calculation formula is: in: For frequency domain indexing, . For the first A Fourier descriptor, whose modulus is Indicates the contour at frequency Shape energy under the current. DC component. The DC component is discarded based on the centroid position of the corresponding contour. .

[0036] Before truncation The low-frequency Fourier coefficients constitute the Fourier feature vector. ,in The value is , This indicates a floor operation. When the total number of outline pixels... When the value is large, take the first 20 low-frequency coefficients; when... When it is less than 40, take There are several coefficients. Low-frequency coefficients correspond to the overall shape characteristics of the contour, while high-frequency coefficients correspond to detail noise.

[0037] The Fourier feature vector is appended to the end of the cell morphology feature vector to form the expanded cell morphology feature vector. The vector length is .

[0038] Example 4: In specific implementation, refer to Figure 3 The cell morphology feature vector and the cell fluorescence intensity distribution feature vector are concatenated along their respective feature dimensions to obtain a joint feature vector. The cell morphology feature vector is denoted as... The characteristic vector of cell fluorescence intensity distribution is denoted as Both are column vectors. The concatenation operation joins the two vectors end-to-end to form a joint feature vector. superscript This indicates transpose. The dimension of the joint feature vector is the sum of the dimensions of the cell morphology feature vector and the dimension of the cell fluorescence intensity distribution feature vector.

[0039] The joint feature vector is input into a support vector machine (SVM) classifier. The SVM classifier is trained using the joint feature vectors from multiple pre-collected, labeled cell categories. The training data is acquired as follows: cell samples of multiple known cell categories are pre-collected. Each cell sample is processed using an imaging flow cytometry cell sorting instrument to acquire a one-dimensional temporal light intensity signal sequence. This sequence is then processed by a well-trained deep convolutional reconstruction network to generate a two-dimensional reconstructed image of the cells to be sorted. Cell morphology feature vectors and cell fluorescence intensity distribution feature vectors are extracted according to the image feature extraction operation described above, and these two vectors are concatenated into a joint feature vector. Simultaneously, the true category label of the cell sample is recorded. The joint feature vectors of all cell samples and their corresponding category labels constitute the training dataset.

[0040] Support Vector Machine (SVM) classifiers employ a "one-to-one" multi-class classification strategy. Preset cell categories, training There are 1 binary classification support vector machine. Each binary classification support vector machine uses a radial basis function kernel function, the expression of which is: in: and For two training joint feature vectors, The kernel function parameters were optimized on the training dataset using five-fold cross-validation. The range of values ​​is to Step size is The soft margin penalty parameter for each binary support vector machine. Similarly optimized using five-fold cross-validation, the value range is... to Step size is During training, the sequence minimum optimization algorithm is used to solve the dual problem of the support vector machine, obtaining the support vectors, Lagrange multipliers, and bias terms for each binary classifier.

[0041] When inputting a joint feature vector to be classified At that time, each binary support vector machine calculates a decision function value, which reflects... The distance relative to the positive and negative classes of the binary classifier. For each preset cell class. All binary classifier decision function values ​​related to this category are voted on, and the number of votes received by this category is counted. The number of votes is divided by the total number of votes (i.e., ...). ), to obtain the classification confidence score corresponding to this category. The score ranges from 0 to 1. The support vector machine classifier outputs a joint feature vector with respect to the classification confidence score for each preset cell category, totaling... Each score.

[0042] Select the preset cell category with the highest classification confidence score as the cell category label output. If there are multiple categories with the highest scores, select the category with the smallest preset category number as the output label.

[0043] Example 5: In specific implementation, the movement velocity of the cells to be sorted within the microfluidic channel is acquired in real time, and the delay time is calculated based on the movement velocity and the distance the cells travel from the imaging area to the location of the sorting actuator. The movement velocity of the cells within the microfluidic channel is acquired in real time by placing a pair of optical detectors upstream or downstream of the imaging area of ​​the microfluidic channel, with the two optical detectors spaced a fixed distance apart along the channel axis. When the cells to be sorted pass through two detectors, they generate pulse signals, and the time difference between the two pulse signals is recorded. Then the speed of motion The distance from the imaging area to the location of the sorting actuator. Delay time is obtained by pre-calibrating the geometry of the microfluidic channel. The unit is seconds. During the calculation, the velocity value is updated in real time, and the latest velocity value is used to calculate the delay time each time a cell category label is output.

[0044] The timing begins from the moment the cell category label is output. A high-precision clock with a resolution of at least 1 microsecond is used. When the timer reaches the delay time, a trigger pulse signal is sent to the sorting actuator. The pulse width of the trigger pulse signal is greater than the action response time of the sorting actuator. The sorting actuator uses piezoelectric ceramics to drive the nozzle or a solenoid valve to drive the deflection plate, with an action response time in the microsecond to millisecond range. The pulse width of the trigger pulse signal is set to twice the action response time of the sorting actuator; for example, if the response time is 500 microseconds, the pulse width is set to 1000 microseconds. The trigger pulse signal is generated by a programmable logic controller and connected to the drive circuit of the sorting actuator through a digital output port.

[0045] The two-dimensional reconstructed image of the cells to be sorted is input into an auxiliary autoencoder. The auxiliary autoencoder has the same encoder structure as the deep convolutional reconstruction network, but its decoder parameters differ. The encoder structure of the deep convolutional reconstruction network is as follows: the input layer receives a one-dimensional temporal light intensity signal sequence, which passes through three one-dimensional convolutional layers. Each one-dimensional convolutional layer uses a kernel size of 2, a stride of 1, and zero padding. Each layer is followed by a batch normalization layer and a ReLU activation function, with output channels of 64, 128, and 256 respectively. The encoder structure of the auxiliary autoencoder is exactly the same as above, but its weight parameters are initialized and optimized during separate training. The decoder structure of the auxiliary autoencoder differs from that of the deep convolutional reconstruction network: the decoder of the deep convolutional reconstruction network contains three two-dimensional transposed convolutional layers with a kernel size of 3, a stride of 2, and output channels of 128, 64, and 1 respectively. Each layer is followed by a batch normalization layer and a ReLU activation function, with the last layer using a Sigmoid activation function. The decoder of the assisted autoencoder also contains three 2D transposed convolutional layers, but the kernel size is set to 3, the stride to 1, and the number of output channels to 64, 32, and 1 respectively. Each layer is followed by a batch normalization layer and a LeakyReLU activation function (negative slope 0.2), and the last layer uses a Sigmoid activation function. The input image size of the assisted autoencoder is the same as the size of the 2D reconstructed image of the cells to be sorted. Consistent, the output is an auxiliary reconstructed image, with the same size. .

[0046] Calculate the structural similarity index between the two-dimensional reconstructed image of the cells to be sorted and the auxiliary reconstructed image. The formula for calculating the structural similarity index is: in: and These represent the two-dimensional reconstructed image of the cells to be sorted and the auxiliary reconstructed image, respectively. and Representing images respectively and images The average grayscale value of the pixels; and Representing images respectively and images The pixel grayscale variance; Representing an image With images covariance; , ,in The dynamic range of pixel grayscale values, for an 8-bit grayscale image. The structural similarity index ranges from 0 to 1, with a higher value indicating greater similarity between the two images.

[0047] When the structural similarity index is lower than a preset similarity threshold, the 2D reconstructed image of the cell to be sorted is marked as a low-confidence image, triggering a re-acquisition of the actual 1D temporal intensity signal sequence. The preset similarity threshold is obtained through statistical analysis of positive samples in the training dataset. Specifically, using all labeled 2D reconstructed cell images and their corresponding auxiliary autoencoder outputs in the training dataset, the structural similarity index of each pair of images is calculated, and 0.8 times the mean of all structural similarity indices is taken as the preset similarity threshold. For example, if the mean is 0.95, the preset similarity threshold is 0.76. When the re-acquisition operation is triggered, a re-acquisition command is sent to the data acquisition system, causing the moving cell to re-acquire the 1D temporal intensity signal sequence when it passes through the imaging area again, and the subsequent reconstruction, feature extraction, and classification processes are re-executed until the structural similarity index is not lower than the preset similarity threshold or the maximum number of re-acquisitions (e.g., 3 times) is reached. If the index is still lower than the threshold after reaching the maximum number of re-acquisitions, the cell is marked as unclassifiable and the sorting is skipped.

[0048] The foregoing has provided a detailed description of one embodiment of the present invention, but this description is merely a preferred embodiment and should not be construed as limiting the scope of the invention. All equivalent variations and modifications made within the scope of the claims of this invention should still fall within the patent coverage of this invention.

Claims

1. A scalable imaging flow cytometer based on self-supervised deep learning, characterized in that, Includes a self-supervised learning module, which is used to perform: The one-dimensional temporal light intensity signal sequence obtained by detecting the moving cells in the microfluidic channel after the mask spatial domain-temporal domain transformation is input into the initialized deep convolutional reconstruction network, and the deep convolutional reconstruction network outputs the reconstructed two-dimensional cell image. The reconstructed two-dimensional cell image is convolved with the mask in the direction of cell movement to generate a predicted one-dimensional temporal light intensity signal sequence. The difference between the one-dimensional temporal light intensity signal sequence and the predicted one-dimensional temporal light intensity signal sequence is calculated, and the network parameters of the deep convolutional reconstruction network are updated by backpropagation using the difference value until the difference value converges to a preset threshold and the update stops, thus obtaining a fully trained deep convolutional reconstruction network. The actual one-dimensional temporal light intensity signal sequence of the cells to be sorted is input into the well-trained deep convolutional reconstruction network to generate the corresponding two-dimensional reconstruction image of the cells to be sorted. Image feature extraction is performed on the two-dimensional reconstructed image of the cells to be sorted to obtain cell morphology feature vectors and cell fluorescence intensity distribution feature vectors; The cell morphology feature vector and the cell fluorescence intensity distribution feature vector are input into the classifier, which outputs cell category labels and triggers the sorting actuator downstream of the microfluidic channel based on the cell category labels.

2. The scalable, imaging flow cytometer based on self-supervised deep learning of claim 1, wherein, Before inputting the one-dimensional temporal light intensity signal sequence obtained by detecting the moving cells in the microfluidic channel after mask spatial-temporal transformation into the initialized deep convolutional reconstruction network, the following operations are performed: A continuous laser beam is reflected by a first bipolar mirror and then irradiated into the cell flow area within the microfluidic channel, exciting the cells to produce fluorescence signals. The fluorescence signal is collected by the objective lens and imaged onto a mask placed on the mirror surface of the objective lens. The mask is provided with a randomly distributed array of light-transmitting holes. The pattern of the light-transmitting hole array is convolved with the two-dimensional image of the cell in the direction of cell movement, thereby converting the spatial distribution of the two-dimensional image of the cell into a one-dimensional temporal light intensity distribution. The one-dimensional temporal light intensity distribution is split into multiple single-photon detector arrays according to wavelength by a second bipolar mirror, and the one-dimensional temporal light intensity signal sequence is obtained by the multiple single-photon detector arrays synchronously.

3. The scalable, imaging flow cytometer based on self-supervised deep learning of claim 2, wherein, The array of light-transmitting holes on the mask is a modified uniform random array distribution, and the projection distances of any two light-transmitting holes in the direction of cell movement are different.

4. The scalable, imaging flow cytometer based on self-supervised deep learning of claim 2, wherein, Before irradiating the cell flow area within the microfluidic channel with the continuous laser beam after reflection by the first bipolar mirror, the following operations are performed: A third laser source is added to the transmission path of the first bipolar mirror. The third laser beam emitted by the third laser source is transmitted through the first bipolar mirror and then combined with the continuous laser beam to irradiate the area through which the cell flows, thereby achieving multi-wavelength excitation.

5. The scalable, self-supervised deep learning-based imaging flow cytometer of claim 1, wherein, Before performing a convolution operation between the reconstructed two-dimensional cell image and the mask in the direction of cell motion, perform the following operations: Obtain a binary matrix representation of the light-transmitting hole pattern of the mask. The number of rows in the binary matrix corresponds to the number of pixels in length of the mask along the cell movement direction, and the number of columns in the binary matrix corresponds to the number of pixels in width of the mask perpendicular to the cell movement direction. The element value at the position corresponding to the light-transmitting hole in the binary matrix is ​​one, and the element value at the position corresponding to the non-light-transmitting hole is zero. The reconstructed two-dimensional cell image is represented as a two-dimensional grayscale matrix. The number of rows in the two-dimensional grayscale matrix corresponds to the number of height pixels of the image along the direction of cell movement, and the number of columns in the two-dimensional grayscale matrix corresponds to the number of width pixels of the image perpendicular to the direction of cell movement. Using the binary matrix as the convolution kernel, a discrete convolution operation with a stride of one is performed along the row direction of the two-dimensional grayscale matrix to obtain the predicted one-dimensional temporal light intensity signal sequence.

6. The scalable, self-supervised deep learning-based imaging flow cytometer of claim 1, wherein, The step of calculating the difference between the one-dimensional temporal light intensity signal sequence and the predicted one-dimensional temporal light intensity signal sequence, and using the difference to backpropagate and update the network parameters of the deep convolutional reconstruction network, includes: The difference value is obtained by subtracting the light intensity values ​​at the same time index position in the one-dimensional time-series light intensity signal sequence from the light intensity values ​​at the same time index position in the predicted one-dimensional time-series light intensity signal sequence point by point, squaring the results and then summing them. The gradient of the difference value with respect to the output layer feature map of the deep convolutional reconstruction network is obtained to obtain the output layer gradient tensor. The gradient tensor of the output layer is backpropagated layer by layer to the input layer of the deep convolutional reconstruction network. During the propagation process of each layer, the gradient of the layer parameters is calculated and the layer parameters are updated according to the gradient descent direction.

7. The scalable, self-supervised deep learning-based imaging flow cytometer of claim 1, wherein, Image feature extraction is performed on the two-dimensional reconstructed image of the cells to be sorted to obtain cell morphology feature vectors and cell fluorescence intensity distribution feature vectors, including: A threshold segmentation operation is performed on the two-dimensional reconstructed image of the cells to be sorted, the foreground pixel set of the cell region is extracted, and the equivalent diameter, eccentricity, and ratio of the convex hull area to the foreground area of ​​the foreground pixel set are calculated. The equivalent diameter, the eccentricity, and the ratio are then concatenated in order to form the cell morphology feature vector. The gray values ​​of each pixel in the two-dimensional reconstructed image of the cells to be sorted are normalized to obtain a normalized gray value matrix. The normalized gray value matrix is ​​then expanded into a one-dimensional vector by rows, which serves as the feature vector of the cell fluorescence intensity distribution.

8. The scalable, imaging flow cytometer based on self-supervised deep learning of claim 7, wherein, After performing threshold segmentation on the two-dimensional reconstructed image of the cells to be sorted and extracting the foreground pixel set of the cell region, the following operations are also performed: Eight-neighborhood tracking is performed on the boundary of the foreground pixel set to obtain the pixel coordinate sequence of the cell outline; Calculate the Fourier descriptor of the pixel coordinate sequence of the cell outline, and extract the first few low-frequency Fourier coefficients to form the Fourier feature vector. The Fourier feature vector is appended to the end of the cell morphology feature vector to form an expanded cell morphology feature vector.

9. The scalable imaging flow cytometer based on self-supervised deep learning according to claim 1, characterized in that, The cell morphology feature vector and the cell fluorescence intensity distribution feature vector are input into a classifier, which outputs cell category labels, including: The cell morphology feature vector and the cell fluorescence intensity distribution feature vector are concatenated along the feature dimension to obtain a joint feature vector; The joint feature vector is input into a support vector machine classifier, which is trained using the joint feature vectors of multiple pre-collected labeled cell categories. The support vector machine classifier outputs the classification confidence score of the joint feature vector relative to each preset cell category. The preset cell category with the highest classification confidence score is selected as the cell category label output.

10. The scalable imaging flow cytometry cell sorter based on self-supervised deep learning according to claim 1, characterized in that, The sorting actuator downstream of the microfluidic channel is triggered based on the cell category label, including: The movement speed of the cells to be sorted within the microfluidic channel is acquired in real time, and the delay time is calculated based on the movement speed and the distance the cells to be sorted travel from the imaging area to the location of the sorting actuator. The timing starts from the moment the cell category label is output. When the timing reaches the delay time, a trigger pulse signal is sent to the sorting execution mechanism. The pulse width of the trigger pulse signal is greater than the action response time of the sorting execution mechanism.