Machine learning based flow cytometry cell classification method, system and terminal

CN116343204BActive Publication Date: 2026-08-21SHANGHAI HUANYI BIOLOGICAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211677238.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-26
Publication Date
2026-08-21
Estimated Expiration
2042-12-26

AI Technical Summary

Technical Problem

该方法的缺点是:分界线的选取主观性大;劳动强度大且效率低

Benefits of technology

[0016] As described above, this invention is a cell classification method, system, and terminal based on machine learning flow cytometry, which has the following beneficial effects: This invention obtains corresponding corrected training samples by normalizing the features and calculating the intra-sample density features of each cell feature in each training sample. Then, batch effects are removed from the corrected test samples and each corrected training sample to obtain multiple final training samples. The machine learning model is trained using these final training samples to obtain a cell classification model. Finally, the trained model is used to predict the test samples to obtain the cell classification result for each cell in the sample. This invention, by introducing sample correction and batch effect removal, alleviates the distribution difference between the new sample to be predicted and the training samples, thereby greatly improving the accuracy of cell classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116343204B_ABST
    Figure CN116343204B_ABST
Patent Text Reader

Abstract

The machine learning-based flow cytometry cell classification method, system and terminal of the present application obtain corresponding correction training samples by normalizing features and calculating intra-sample density features of each cell feature of each training sample, remove batch effects from the correction to-be-detected sample and each correction training sample to obtain a plurality of final training samples, train a machine learning model using each final training sample to obtain a cell classification model, and finally use the trained model for to-be-detected sample prediction to obtain the cell classification result of each cell in the sample. The present application alleviates the distribution difference between the to-be-predicted new sample and the training sample by introducing the sample correction and batch effect removal link, thereby greatly improving the accuracy of cell classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of flow cytometry technology, and in particular to a method, system, and terminal for cell classification using flow cytometry technology based on machine learning. Background Technology

[0002] Currently, the analysis of immune cells using flow cytometry is mainly divided into two types: manual analysis and model-assisted analysis.

[0003] The manual analysis method involves combining machine-generated features in a single dimension or in pairs to create a scatter plot of all cells in a single sample. Appropriate boundaries are then selected to classify the cells into different types. This process is repeated, using different feature pairs, to sequentially identify each cell subpopulation. The drawbacks of this method are: the selection of boundaries is highly subjective; it is labor-intensive and inefficient.

[0004] Model-assisted methods involve training a machine learning model using a subset of manually labeled samples to classify unlabeled samples. A drawback of this method is that when there are distributional differences between samples, the machine learning model, which models these differences at the single-cell level through gene expression, cannot identify them, leading to lower classification accuracy. Summary of the Invention

[0005] In view of the shortcomings of the prior art described above, the purpose of this invention is to provide a cell classification method, system and terminal based on machine learning flow cytometry technology to solve the above-mentioned technical problems in the prior art.

[0006] To achieve the above and other related objectives, this invention provides a cell classification method based on machine learning flow cytometry. The method includes: acquiring multiple training samples, each containing multiple cell types; wherein each training sample includes: original feature data of each cell feature sequentially arranged within the cell; wherein each cell has the same number of cell features; calculating normalized features and intra-sample density features for each cell feature of each training sample to obtain corresponding corrected training samples; wherein each corrected training sample includes: corrected feature data corresponding to each cell feature; removing batch effects from the corrected test samples obtained by normalizing features and calculating intra-sample density features from the test sample, as well as from each corrected training sample, to obtain multiple final training samples; training a machine learning model based on each final training sample to obtain a cell classification model; and obtaining the cell classification result for each cell in the test sample based on the cell classification model.

[0007] In one embodiment of the present invention, the step of calculating normalized features and intra-sample density features for each cell feature of each training sample to obtain corresponding corrected training samples includes: calculating normalized feature data corresponding to each cell feature of each training sample based on the original feature data of each cell feature of each training sample; normalizing the original feature data of each cell feature of each training sample to obtain intra-sample density feature data of each cell feature of each training sample; and concatenating the normalized feature data and intra-sample density features corresponding to each cell feature of each training sample to obtain corrected feature data corresponding to each cell feature of each training sample; wherein the corrected feature data includes: original feature data, normalized feature data, and intra-sample density feature data.

[0008] In one embodiment of the present invention, the calculation of normalized feature data corresponding to each cell feature of each training sample based on the original feature data of each cell feature of each training sample includes: calculating normalized feature data corresponding to each cell feature of each training sample based on the original feature data of each cell feature of each training sample using a normalized feature calculation method; wherein, the normalized feature calculation method includes: obtaining normalized feature data corresponding to the current cell feature based on the mean and standard deviation calculated based on the original feature data of the current cell feature; and wherein, the method of calculating the mean and standard deviation includes: calculating the mean and standard deviation of the original feature data of all cell features in the current training sample that are arranged in the same position within the cell as the current cell feature.

[0009] In one embodiment of the present invention, the step of normalizing the original feature data of each cell feature of each training sample and obtaining the intra-sample density feature data of each cell feature of each training sample includes: normalizing the original feature data of each cell feature of each training sample to obtain the normalized value of each cell feature corresponding to each training sample; setting a corresponding normalization value range based on the normalized value of each cell feature of each training sample; and calculating the intra-sample density feature data corresponding to each cell feature of each training sample based on the normalized value range of each cell feature of each sample and the normalized value of each cell feature of each training sample using a density feature calculation method.

[0010] In one embodiment of the present invention, the step of setting a normalized value interval corresponding to the normalized value of each cell feature of each training sample includes: setting a normalized value interval symmetrical about the normalized value for each cell feature based on the normalized value of each cell feature of each training sample and a set normalized interval offset threshold.

[0011] In one embodiment of the present invention, the density feature calculation method includes: obtaining the number of cells in the current training sample whose normalized values ​​fall within the normalized value range corresponding to the cell feature, having the same intracellular arrangement position as the current cell feature; and calculating the intra-sample density feature data corresponding to the cell feature based on the number of cells and the number of cells in the current training sample.

[0012] In one embodiment of the present invention, normalizing the original feature data of each cell feature of each training sample to obtain the normalized value of each cell feature corresponding to each training sample includes: normalizing the original feature data of each cell feature of each training sample based on a normalization calculation method to obtain the normalized value of each cell feature corresponding to each training sample; wherein, the normalization calculation method includes: calculating the normalized value corresponding to the current cell feature based on the minimum and maximum values ​​of the original feature data of all cell features that are in the same cell position as the current cell feature in the current training sample.

[0013] In one embodiment of the present invention, removing batch effects from the corrected test samples and each corrected training sample obtained by calculating the normalized features and intra-sample density features of the test sample to obtain multiple final training samples includes: calculating the normalized features and intra-sample density features of each cell feature of the test sample to obtain the corrected test sample; wherein, the corrected test sample includes: corrected feature data corresponding to each cell feature; merging the corrected test samples and each corrected training sample into a sample set, and removing batch effects from the corrected feature data corresponding to each cell feature of each sample in the sample set to obtain feature processing data for each cell feature of the corresponding corrected test sample and each corrected training sample; concatenating the corrected feature data corresponding to each cell feature in the corrected test sample and each corrected training sample with the feature processing data to obtain multiple final training samples; wherein, the feature data corresponding to each cell feature of each final training sample includes corrected feature data and feature processing data.

[0014] To achieve the above and other related objectives, this invention provides a cell classification system based on machine learning flow cytometry. The system includes: a training sample acquisition module for acquiring multiple training samples, each containing multiple cell types; wherein each training sample includes: original feature data of each cell feature sequentially arranged within the cell; wherein each cell has the same number of cell features; and a sample correction module connected to the training sample acquisition module for performing normalized features and in-sample density feature calculations on each cell feature of each training sample to obtain corresponding corrected training samples; wherein each corrected training sample... The training samples include: corrected feature data corresponding to each cell feature; a batch effect removal module, connected to the sample correction module, used to remove batch effects from the corrected test samples obtained by calculating the normalized features and intra-sample density features of the test samples, as well as each corrected training sample, to obtain multiple final training samples; a cell classification model training module, connected to the batch effect removal module, used to train the machine learning model based on each final training sample to obtain a cell classification model; and a cell classification module, connected to the cell classification model training module, used to obtain the cell classification result of each cell in the test sample based on the cell classification model.

[0015] To achieve the above and other related objectives, the present invention provides a machine learning-based flow cytometry cell classification terminal, comprising: one or more memory units and one or more processor units; the one or more memory units are used to store a computer program; the one or more processor units are connected to the memory units and are used to run the computer program to execute the machine learning-based flow cytometry cell classification method.

[0016] As described above, this invention is a cell classification method, system, and terminal based on machine learning flow cytometry, which has the following beneficial effects: This invention obtains corresponding corrected training samples by normalizing the features and calculating the intra-sample density features of each cell feature in each training sample. Then, batch effects are removed from the corrected test samples and each corrected training sample to obtain multiple final training samples. The machine learning model is trained using these final training samples to obtain a cell classification model. Finally, the trained model is used to predict the test samples to obtain the cell classification result for each cell in the sample. This invention, by introducing sample correction and batch effect removal, alleviates the distribution difference between the new sample to be predicted and the training samples, thereby greatly improving the accuracy of cell classification. Attached Figure Description

[0017] Figure 1 The diagram shows a flowchart of a machine learning-based flow cytometry cell classification method according to an embodiment of the present invention.

[0018] Figure 2 The diagram shows a flowchart of a traditional supervised machine learning classification method according to an embodiment of the present invention.

[0019] Figure 3 The diagram shows a flowchart of a machine learning-based flow cytometry cell classification method according to an embodiment of the present invention.

[0020] Figure 4 The diagram shown is a schematic representation of a machine learning-based flow cytometry cell classification system according to an embodiment of the present invention.

[0021] Figure 5 The diagram shown is a structural schematic of a machine learning-based flow cytometry cell classification terminal according to an embodiment of the present invention. Detailed Implementation

[0022] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, unless otherwise specified, the following embodiments and features described therein can be combined with each other.

[0023] It should be noted that in the following description, reference is made to the accompanying drawings, which illustrate several embodiments of the invention. It should be understood that other embodiments may also be used, and changes in mechanical composition, structure, electrical system, and operation may be made without departing from the spirit and scope of the invention. The following detailed description should not be considered limiting, and the scope of the embodiments of the invention is defined only by the claims of the published patents. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. Spatially related terms, such as “upper,” “lower,” “left,” “right,” “below,” “below,” “lower part,” “above,” “upper part,” etc., may be used herein to illustrate the relationship between one element or feature shown in the figures and another element or feature.

[0024] Throughout this specification, when it is said that a part is "connected" to another part, this includes not only "direct connection" but also "indirect connection" by placing other elements in between. Furthermore, when it is said that a part "includes" a certain constituent element, unless otherwise stated otherwise, this does not exclude other constituent elements, but rather means that other constituent elements may also be included.

[0025] The terms "first," "second," and "third," etc., used herein are for the purpose of describing various parts, components, regions, layers, and / or segments, but are not limiting. These terms are used only to distinguish one part, component, region, layer, or segment from others. Therefore, the "first part," "component," "region," "layer," or "segment" described below may refer to a "second part," "component," "region," "layer," or "segment" without departing from the scope of this invention.

[0026] Furthermore, as used herein, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context indicates otherwise. It should be further understood that the terms “comprising,” “including,” indicate the presence of the stated feature, operation, element, component, item, kind, and / or group, but do not preclude the presence, occurrence, or addition of one or more other features, operations, elements, components, items, kinds, and / or groups. The terms “or” and “and / or” as used herein are interpreted as inclusive, or mean any one or any combination thereof. Thus, “A, B, or C” or “A, B, and / or C” means “any one of: A; B; C; A and B; A and C; B and C; A, B, and C.” Exceptions to this definition arise only when combinations of elements, functions, or operations are inherently mutually exclusive in some manner.

[0027] This invention provides a cell classification method, system, and terminal based on machine learning flow cytometry. It obtains corresponding corrected training samples by normalizing the features of each cell in each training sample and calculating the intra-sample density features. Then, batch effects are removed from the corrected test samples and all corrected training samples to obtain multiple final training samples. These final training samples are used to train a machine learning model to obtain a cell classification model. Finally, the trained model is used to predict the test samples to obtain the cell classification result for each cell in the sample. This invention, by introducing sample correction and batch effect removal, alleviates the distribution difference between the new samples to be predicted and the training samples, thereby significantly improving the accuracy of cell classification.

[0028] Flow cytometry (FCM) is a single-cell quantitative analysis and sorting technique using a flow cytometer. In flow cytometry, cell markers are used as fluorescent labels to identify and separate cell populations and subpopulations. The molecular fluorescence spectral data obtained by detecting the fluorescently labeled cells are then used as cell characteristics in machine learning models.

[0029] This invention is a cell classification method based on flow cytometry technology, and the samples used are all cell collections produced by a flow cytometer.

[0030] The present invention will now be described in detail with reference to the accompanying drawings, so that those skilled in the art can readily implement it. The present invention can be embodied in many different forms and is not limited to the embodiments described herein.

[0031] like Figure 1 This is a flowchart illustrating a flow cytometry-based cell classification method based on machine learning, as shown in an embodiment of the present invention.

[0032] The method includes:

[0033] Step S1: Obtain multiple training samples, each containing cells of multiple cell types.

[0034] In detail, each training sample is a collection of cells produced by flow cytometry. In the cell collection, each cell corresponds to a cell type and can obtain a set of feature values ​​from the flow cytometry, which contains the feature values ​​of each cell's characteristics.

[0035] Specifically, each training sample includes: the original feature data of each cell feature arranged sequentially within the cell; wherein each cell has the same number of cell features. That is, the feature data of each cell feature corresponds to the feature data of the cell type of that cell.

[0036] Step S2: Calculate the normalized features and intra-sample density features for each cell feature of each training sample to obtain the corresponding corrected training samples.

[0037] In detail, each calibration training sample includes: calibration feature data corresponding to each cell feature.

[0038] In one embodiment, step S2 includes:

[0039] Based on the original feature data of each cell feature of each training sample, calculate the normalized feature data corresponding to each cell feature of each training sample.

[0040] The raw feature data of each cell feature in each training sample is normalized to obtain the intra-sample density feature data of each cell feature in each training sample.

[0041] The normalized feature data and intra-sample density features corresponding to each cell feature of each training sample are concatenated to obtain the corrected feature data corresponding to each cell feature of each training sample; wherein, the corrected feature data includes: original feature data, normalized feature data, and intra-sample density feature data. That is, if the original feature data of each cell feature was originally K-dimensional, it becomes 3K-dimensional after correction;

[0042] In sample calibration, this invention calculates normalized features and intra-sample density features, so that each cell obtains the distribution information of the sample, which helps to reduce the distribution differences between samples.

[0043] In one specific embodiment, the calculation of normalized feature data corresponding to each cell feature of each training sample based on the original feature data of each cell feature of each training sample includes:

[0044] Based on the normalized feature calculation method, the normalized feature data corresponding to each cell feature of each training sample is calculated according to the original feature data of each cell feature of each training sample.

[0045] The normalized feature calculation method includes:

[0046] The mean and standard deviation of the original feature data of the current cell features are calculated, and the normalized feature data corresponding to the current cell features are obtained based on the original feature data of the current cell features.

[0047] Furthermore, the calculation of the mean and standard deviation includes: calculating the mean and standard deviation of the original feature data of all cell features in the current training sample that are in the same position within the cell as the current cell feature. For example, if a sample has cell 1, cell 2, and cell 3, and each has three arrangements of cell features; the mean of the original feature data corresponding to cell feature 1 of cell 1 is calculated as the mean of the original feature data of cell feature 1 of cell 2 and cell 3 and cell feature 1 of cell 1; similarly, the standard interpolation is also calculated based on the original feature data of the above three cell features.

[0048] Specifically, let P be the sample size and Q be the number of cells contained in the sample. (p) (p = 1, 2, ..., P), the cells have a total of K features. Multiple training sample data are denoted as a matrix. The goal of this patent is to classify each cell.

[0049] First, calculate the Z-score for each cell feature within each sample. Let the mean of cell feature k (k = 1, 2, ..., K) in sample p be... Standard deviation is If the original feature data for a certain cell is x, then the Z-score is... Therefore, the cell features within each sample are normalized, resulting in the original feature data having a mean of 0 and a standard deviation of 1. This transformed feature is called Z-score normalized feature data.

[0050] In one specific embodiment, the normalization of the original feature data of each cell feature of each training sample and the acquisition of the intra-sample density feature data of each cell feature of each training sample includes:

[0051] The original feature data of each cell feature in each training sample is normalized to obtain the normalized value of each cell feature corresponding to each training sample.

[0052] A normalized value range is set based on the normalized value of each cell feature of each training sample.

[0053] Based on the density feature calculation method, the intra-sample density feature data corresponding to each cell feature of each training sample is calculated according to the normalized value range of each cell feature of each sample and the normalized value of each cell feature of each training sample.

[0054] In one specific embodiment, normalizing the original feature data of each cell feature for each training sample to obtain the normalized value of each cell feature corresponding to each training sample includes:

[0055] Based on the normalization calculation method, the original feature data of each cell feature of each training sample is normalized to obtain the normalized value of each cell feature corresponding to each training sample.

[0056] The normalization calculation method includes:

[0057] Based on the minimum and maximum values ​​of the original feature data of all cell features that are in the same position within the cell as the current cell feature in the current training sample, the normalized value corresponding to the current cell feature is calculated. For example, if a sample has cell 1, cell 2, and cell 3, and each has three arrangements of cell features; the maximum and minimum values ​​corresponding to the original feature data of cell feature 1 of cell 1 are calculated as the maximum and minimum values ​​of the original feature data of cell feature 1 of cell 2 and cell 3 and cell feature 1 of cell 1.

[0058] Specifically, within each sample, each cell feature is normalized to a value between 0 and 1. Specifically, let the minimum value of cell feature k within sample p be... The maximum value is If the value of this feature for a certain cell is x, then the normalized value is

[0059] In one embodiment, setting a normalization value range corresponding to the normalized value of each cell feature based on each training sample includes:

[0060] Based on the normalized values ​​of each cell feature of each training sample and the set normalization interval offset threshold, a normalization value interval symmetrical about the normalized value is set for each cell feature.

[0061] Specifically, if the normalized value of a cell feature is x”, and a normalization interval offset threshold s is set, then the normalization value interval is [x”-s, x”+s]. Preferably, s is generally a natural number less than 1 and greater than 0.

[0062] In one embodiment, the density feature calculation method includes:

[0063] Obtain the number of cells in the current training sample whose normalized values ​​fall within the normalized value range of the cell feature that has the same intracellular arrangement position as the current cell feature.

[0064] The in-sample density feature data corresponding to the cell feature is calculated based on the number of cells and the number of cells in the current training sample.

[0065] Specifically, for each x”, calculate the density of that value within sample p. Let the normalized value of feature k, which has t cells in sample p, fall within the interval [x”-s, x”+s]. Then the density of x” is x″′=t / Q (p) Q (p) Let x be the total number of cells in the sample. This completes the transformation from x to x″′. The feature resulting from this transformation is called the in-sample density feature.

[0066] Step S3: Remove batch effects from the corrected test samples and each corrected training sample obtained by normalizing the test samples and calculating the intra-sample density features to obtain multiple final training samples.

[0067] In one embodiment, step S3 may include:

[0068] The normalized feature and intra-sample density feature of each cell feature of the sample to be tested are calculated to obtain the corrected sample to be tested; wherein, the corrected sample to be tested includes: the corrected feature data corresponding to each cell feature;

[0069] The calibration test samples and each calibration training sample are merged into a sample set. Batch effects are then removed from the calibration feature data corresponding to each cell feature of each sample in the sample set to obtain the feature processing data for each cell feature of the calibration test samples and each calibration training sample. It should be noted that batch effect removal methods include, but are not limited to, Harmony, BERMUDA, DESC, and iMAP. Furthermore, the dimension of the generated feature processing data can be set according to requirements and can be any dimension.

[0070] The calibration feature data corresponding to each cell feature in the calibration test sample and each calibration training sample are concatenated with the feature processing data to obtain multiple final training samples. Each cell feature in each final training sample includes both calibration feature data and feature processing data. If the original feature data of each cell feature was K-dimensional before calibration, it becomes 3K+2K-dimensional after calibration; the generated feature processing data has a dimension of 2K.

[0071] This invention uses training samples obtained after removing batch effects to alleviate the distribution differences between the new samples to be predicted and the training samples, thereby improving the accuracy of cell classification.

[0072] Step S4: Train the machine learning model based on each final training sample to obtain a cell classification model.

[0073] In one embodiment, the machine learning model includes, but is not limited to, supervised models such as decision trees, support vector machines, Bayesian classifiers, GBDT algorithms and their variants (such as XGBoost, LightGBM, Histogram-based Gradient BoostingClassification Tree).

[0074] Step S5: Based on the cell classification model, obtain the cell classification result of each cell in the sample according to the sample to be tested.

[0075] It should be noted that the test sample used in this invention is in the same form as the training sample, that is, the sample includes the original feature data of each cell feature that is arranged sequentially in the cell, corresponding to each cell in the sample; wherein, each cell has the same number of cell features.

[0076] To better illustrate the differences between this invention and the prior art, the following description is provided:

[0077] like Figure 2 As shown, this is a traditional supervised machine learning method, which involves directly training the machine learning model using training samples; then inputting the samples to be predicted into the machine learning model to output cell classification results.

[0078] The classification method of this invention is as follows: Figure 3As shown, the process begins with sample calibration using training samples. This involves normalizing the features of each cell in each training sample and calculating the intra-sample density features to obtain corresponding calibrated training samples. The samples to be tested then undergo the same calibration process. Batch effects are removed from both the calibrated samples and the calibrated training samples. The batch-effect-removed sample data is then concatenated with the calibrated samples to obtain the final training samples. These final training samples are then used to train a supervised learning model. Finally, the samples to be tested are input into the trained machine learning model to obtain cell classification results.

[0079] This patent has the following advantages compared to existing technologies:

[0080] 1. Calculate normalized features and intra-sample density features during sample calibration so that each cell obtains the distribution information of the sample, which helps to reduce the distribution differences between samples.

[0081] 2. Breaking away from the traditional machine learning modeling process, the training samples and the samples to be predicted are merged to remove batch effects. The training samples are then used to model and predict on the new samples, further mitigating the model prediction error caused by the distribution differences between samples.

[0082] Similar to the principles of the above embodiments, the present invention provides a cell classification system based on machine learning flow cytometry technology.

[0083] The following specific embodiments are provided in conjunction with the accompanying drawings:

[0084] like Figure 4 This invention presents a schematic diagram of the structure of a flow cytometry cell classification system based on machine learning, as described in an embodiment of the present invention.

[0085] The system includes:

[0086] The training sample acquisition module 41 is used to acquire multiple training samples containing multiple cell types; wherein each training sample includes: the original feature data of each cell feature arranged sequentially within the cell corresponding to each cell in the sample; wherein each cell has the same number of cell features;

[0087] The sample correction module 42 is connected to the training sample acquisition module 41 and is used to calculate the normalization feature and intra-sample density feature for each cell feature of each training sample to obtain the corresponding corrected training samples; wherein, each corrected training sample includes: the corrected feature data corresponding to each cell feature respectively.

[0088] Batch effect removal module 43 is connected to the sample correction module 42 and is used to remove batch effects from the corrected test sample and each corrected training sample obtained by normalizing features and intra-sample density features of the test sample to obtain multiple final training samples.

[0089] The cell classification model training module 44 is connected to the batch effect removal module 43 and is used to train the machine learning model based on each final training sample to obtain the cell classification model.

[0090] The cell classification module 45 is connected to the cell classification model training module 44 and is used to obtain the cell classification result of each cell in the sample based on the cell classification model and the sample to be detected.

[0091] It should be noted that, as should be understood Figure 4 The division of modules in the system embodiment is merely a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, these units can be implemented entirely in software through processing element calls; they can be implemented entirely in hardware; or some units can be implemented by processing element calls to software, while others are implemented in hardware.

[0092] Since the implementation principle of this machine learning-based flow cytometry cell classification system has been described in the foregoing embodiments, it will not be repeated here.

[0093] Optionally, the step of calculating normalized features and intra-sample density features for each cell feature of each training sample to obtain corresponding corrected training samples includes: calculating normalized feature data corresponding to each cell feature of each training sample based on the original feature data of each cell feature of each training sample; normalizing the original feature data of each cell feature of each training sample to obtain intra-sample density feature data of each cell feature of each training sample; and concatenating the normalized feature data and intra-sample density features corresponding to each cell feature of each training sample to obtain corrected feature data corresponding to each cell feature of each training sample; wherein the corrected feature data includes: original feature data, normalized feature data, and intra-sample density feature data.

[0094] Optionally, the step of calculating the normalized feature data corresponding to each cell feature of each training sample based on the original feature data of each cell feature of each training sample includes: calculating the normalized feature data corresponding to each cell feature of each training sample based on the original feature data of each cell feature of each training sample using a normalized feature calculation method; wherein, the normalized feature calculation method includes: obtaining the normalized feature data corresponding to the current cell feature based on the mean and standard deviation calculated based on the original feature data of the current cell feature; and wherein, the method of calculating the mean and standard deviation includes: calculating the mean and standard deviation of the original feature data of all cell features in the current training sample that are arranged in the same position within the cell as the current cell feature.

[0095] Optionally, the step of normalizing the original feature data of each cell feature of each training sample and obtaining the intra-sample density feature data of each cell feature of each training sample includes: normalizing the original feature data of each cell feature of each training sample to obtain the normalized value of each cell feature corresponding to each training sample; setting a corresponding normalization value range based on the normalized value of each cell feature of each training sample; and calculating the intra-sample density feature data corresponding to each cell feature of each training sample based on the normalized value range of each cell feature of each sample and the normalized value of each cell feature of each training sample using a density feature calculation method.

[0096] Optionally, setting a normalization value range based on the normalized value of each cell feature of each training sample includes: setting a normalization value range symmetrical about the normalized value for each cell feature based on the normalized value of each cell feature of each training sample and a set normalization range offset threshold.

[0097] Optionally, the density feature calculation method includes: obtaining the number of cells in the current training sample whose normalized values ​​fall within the normalized value range corresponding to the current cell feature, having the same intracellular arrangement position as the current cell feature; and calculating the intra-sample density feature data corresponding to the cell feature based on the number of cells and the number of cells in the current training sample.

[0098] Optionally, normalizing the original feature data of each cell feature in each training sample to obtain the normalized value of each cell feature corresponding to each training sample includes: normalizing the original feature data of each cell feature in each training sample based on a normalization calculation method to obtain the normalized value of each cell feature corresponding to each training sample; wherein, the normalization calculation method includes: calculating the normalized value corresponding to the current cell feature based on the minimum and maximum values ​​among the original feature data of all cell features that are in the same cell position as the current cell feature in the current training sample.

[0099] Optionally, the step of removing batch effects from the corrected test samples and each corrected training sample obtained by calculating the normalized features and intra-sample density features of the test sample to obtain multiple final training samples includes: calculating the normalized features and intra-sample density features of each cell feature of the test sample to obtain the corrected test sample; wherein, the corrected test sample includes: corrected feature data corresponding to each cell feature; merging the corrected test samples and each corrected training sample into a sample set, and removing batch effects from the corrected feature data corresponding to each cell feature of each sample in the sample set to obtain feature processing data for each cell feature of the corresponding corrected test sample and each corrected training sample; concatenating the corrected feature data corresponding to each cell feature in the corrected test sample and each corrected training sample with the feature processing data to obtain multiple final training samples; wherein, the feature data corresponding to each cell feature of each final training sample includes corrected feature data and feature processing data.

[0100] like Figure 5 A schematic diagram of the structure of the cell classification terminal 10 based on machine learning in an embodiment of the present invention is shown.

[0101] The machine learning-based flow cytometry cell classification terminal 50 includes: a memory 51 and a processor 52. The memory 51 stores computer programs; the processor 52 runs the computer programs to implement, for example... Figure 1 The aforementioned cell classification method based on machine learning flow cytometry technology.

[0102] Optionally, the number of memories 51 can be one or more, and the number of processors 52 can be one or more. Figure 5 Each example is taken as an instance.

[0103] Optionally, the processor 52 in the machine learning-based flow cytometry cell classification terminal 50 will follow the instructions as follows: Figure 1The steps described involve loading one or more instructions corresponding to the process of an application into memory 51, and having the processor 52 run the application stored in the first memory 51, thereby achieving the following: Figure 1 Various functions in the machine learning-based flow cytometry cell classification method.

[0104] Optionally, the memory 51 may include, but is not limited to, high-speed random access memory and non-volatile memory. For example, one or more disk storage devices, flash memory devices, or other non-volatile solid-state storage devices; the processor 52 may include, but is not limited to, a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0105] Optionally, the processor 52 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0106] The present invention also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed, implements as follows: Figure 1The illustrated method is a machine learning-based flow cytometry-based cell classification method. The computer-readable storage medium may include, but is not limited to, floppy disks, optical disks, CD-ROMs (Read-Only Optical Disk Memory), magneto-optical disks, ROMs (Read-Only Memory), RAMs (Random Access Memory), EPROMs (Erasable Programmable Read-Only Memory), EEPROMs (Electrically Erasable Programmable Read-Only Memory), magnetic cards or optical cards, flash memory, or other types of media / machine-readable media suitable for storing machine-executable instructions. The computer-readable storage medium may be a product not connected to a computer device or a component used with a computer device.

[0107] In summary, the machine learning-based flow cytometry cell classification system of this invention obtains corresponding corrected training samples by normalizing the features of each cell in each training sample and calculating the intra-sample density features. Then, batch effects are removed from the corrected test samples and each corrected training sample to obtain multiple final training samples. These final training samples are used to train a machine learning model to obtain a cell classification model. Finally, the trained model is used to predict the test samples to obtain the cell classification result for each cell in the sample. This invention, by introducing sample correction and batch effect removal, alleviates the distribution difference between the new samples to be predicted and the training samples, thereby greatly improving the accuracy of cell classification. Therefore, this invention effectively overcomes the various shortcomings of existing technologies and has high industrial application value.

[0108] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.

Claims

1. A cell classification method based on machine learning flow cytometry, characterized in that, The method includes: Obtain multiple training samples containing multiple cell types; each training sample includes: the original feature data of each cell feature arranged sequentially within the cell; each cell has the same number of cell features. For each cell feature of each training sample, normalization features and intra-sample density features are calculated to obtain corresponding corrected training samples; wherein, each corrected training sample includes: corrected feature data corresponding to each cell feature; The batch effect is removed from the corrected test samples and each corrected training sample obtained by normalizing the test samples and calculating the intra-sample density features, so as to obtain multiple final training samples. The machine learning model is trained based on each final training sample to obtain a cell classification model; Based on the cell classification model, the cell classification result of each cell in the sample is obtained according to the sample to be tested; The step of calculating the normalized feature and intra-sample density feature for each cell feature of each training sample to obtain the corresponding corrected training samples includes: Based on the original feature data of each cell feature of each training sample, calculate the normalized feature data corresponding to each cell feature of each training sample. The raw feature data of each cell feature in each training sample is normalized to obtain the intra-sample density feature data of each cell feature in each training sample. The normalized feature data and intra-sample density feature data corresponding to each cell feature of each training sample are concatenated to obtain the corrected feature data corresponding to each cell feature of each training sample; wherein, the corrected feature data includes: original feature data, normalized feature data and intra-sample density feature data. The process of normalizing the original feature data of each cell feature in each training sample and obtaining the intra-sample density feature data of each cell feature in each training sample includes: normalizing the original feature data of each cell feature in each training sample to obtain the normalized value of each cell feature corresponding to each training sample; setting a normalized value range symmetrical about the normalized value for each cell feature based on the normalized value of each cell feature in each training sample and a set normalization range offset threshold; obtaining the number of cells in the current training sample whose normalized value falls within the normalized value range corresponding to the current cell feature, among the cell features that are arranged in the same position within the cell as the current cell feature; and calculating the intra-sample density feature data corresponding to the cell feature based on the number of cells and the number of cells in the current training sample.

2. The cell classification method based on machine learning flow cytometry as described in claim 1, characterized in that, The calculation of normalized feature data corresponding to each cell feature of each training sample based on the original feature data of each cell of each training sample includes: Based on the normalized feature calculation method, the normalized feature data corresponding to each cell feature of each training sample is calculated according to the original feature data of each cell feature of each training sample. The normalized feature calculation method includes: The mean and standard deviation of the original feature data of the current cell features are calculated, and the normalized feature data corresponding to the current cell features are obtained by normalizing the original feature data of the current cell features. Furthermore, the method for calculating the mean and standard deviation includes: calculating the mean and standard deviation of the original feature data of all cell features that are in the same cell arrangement position as the current cell feature in the current training sample.

3. The cell classification method based on machine learning flow cytometry as described in claim 1, characterized in that, The process of normalizing the original feature data of each cell feature in each training sample to obtain the normalized value of each cell feature corresponding to each training sample includes: Based on the normalization calculation method, the original feature data of each cell feature of each training sample is normalized to obtain the normalized value of each cell feature corresponding to each training sample. The normalization calculation method includes: Based on the minimum and maximum values ​​of the original feature data of all cell features that are in the same position within the cell as the current cell feature in the current training sample, the normalized value corresponding to the current cell feature is calculated.

4. The cell classification method based on machine learning flow cytometry as described in claim 1, characterized in that, The process of removing batch effects from the corrected test samples (obtained by normalizing the test samples and calculating the intra-sample density features) and each corrected training sample to obtain multiple final training samples includes: The normalized feature and intra-sample density feature of each cell feature of the sample to be tested are calculated to obtain the corrected sample to be tested; wherein, the corrected sample to be tested includes: the corrected feature data corresponding to each cell feature; The calibration test samples and each calibration training sample are merged into a sample set, and the batch effect is removed from the calibration feature data corresponding to each cell feature of each sample in the sample set to obtain the feature processing data of each cell feature of the calibration test sample and each calibration training sample. The calibration feature data and feature processing data corresponding to each cell feature in the calibration test sample and each calibration training sample are concatenated to obtain multiple final training samples; wherein, the feature data corresponding to each cell feature in each final training sample includes calibration feature data and feature processing data.

5. A cell classification system based on machine learning flow cytometry, characterized in that, The system includes: The training sample acquisition module is used to acquire multiple training samples containing multiple cell types; each training sample includes: the original feature data of each cell feature arranged sequentially within the cell; wherein each cell has the same number of cell features. The sample correction module, connected to the training sample acquisition module, is used to calculate the normalized feature and intra-sample density feature for each cell feature of each training sample to obtain the corresponding corrected training samples; wherein, each corrected training sample includes: corrected feature data corresponding to each cell feature; The batch effect removal module is connected to the sample correction module and is used to remove batch effects from the corrected test sample and each corrected training sample obtained by normalizing the test sample and calculating the intra-sample density features, so as to obtain multiple final training samples. The cell classification model training module, connected to the batch effect removal module, is used to train the machine learning model based on each final training sample to obtain the cell classification model. A cell classification module, connected to the cell classification model training module, is used to obtain the cell classification result of each cell in the sample based on the cell classification model and the sample to be detected. The step of calculating the normalized feature and intra-sample density feature for each cell feature of each training sample to obtain the corresponding corrected training samples includes: Based on the original feature data of each cell feature of each training sample, calculate the normalized feature data corresponding to each cell feature of each training sample. The raw feature data of each cell feature in each training sample is normalized to obtain the intra-sample density feature data of each cell feature in each training sample. The normalized feature data and intra-sample density feature data corresponding to each cell feature of each training sample are concatenated to obtain the corrected feature data corresponding to each cell feature of each training sample; wherein, the corrected feature data includes: original feature data, normalized feature data and intra-sample density feature data. The process of normalizing the original feature data of each cell feature in each training sample and obtaining the intra-sample density feature data of each cell feature in each training sample includes: normalizing the original feature data of each cell feature in each training sample to obtain the normalized value of each cell feature corresponding to each training sample; setting a normalized value range symmetrical about the normalized value for each cell feature based on the normalized value of each cell feature in each training sample and a set normalization range offset threshold; obtaining the number of cells in the current training sample whose normalized value falls within the normalized value range corresponding to the current cell feature, among the cell features that are arranged in the same position within the cell as the current cell feature; and calculating the intra-sample density feature data corresponding to the cell feature based on the number of cells and the number of cells in the current training sample.

6. A cell classification terminal based on machine learning flow cytometry technology, characterized in that, include: One or more memories and one or more processors; The one or more memories are used to store computer programs; The one or more processors are connected to the memory and are used to run the computer program to perform the method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Large-scale single cell typing method and system and storage medium

    CN112908414A

  • Data correction and classification method and storage medium

    CN113270191A