Method and device for screening, setting gate and grouping based on multiple protein flow data

By employing an automated gating and clustering method for multiple protein flow cytometry data, impurities and off-target data are removed, and gating convex hull polygons are used for clustering. This solves the problem of complex data processing in multifactor flow cytometry detection and improves detection efficiency and accuracy.

CN116646014BActive Publication Date: 2026-03-17SHANGHAI XINDIAN BIOTECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-17
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

In existing flow cytometry multifactor detection technologies, the uncertainty of data distribution leads to severe cell debris and adhesion phenomena. The manual removal of invalid data is complicated, time-consuming, and labor-intensive, which affects the detection efficiency.

Method used

An automatic gating and clustering method based on multiple protein flow cytometry data is adopted. By removing impurity and off-target data, automatic clustering is performed using gating convex hull polygons. Combined with the optimized 3σ principle and the standard deviation calculation of the median replacing the mean, fully automatic gating and clustering is achieved.

Benefits of technology

It achieves automated data preprocessing and filtering, accurately retains valid data, reduces manual operations, improves detection efficiency and result accuracy, and expands the scope of data applicability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116646014B_ABST
    Figure CN116646014B_ABST
Patent Text Reader

Abstract

The application discloses a method and device for screening and gating and grouping based on multiple protein flow data. The method comprises the following steps: reading sample flow data files and corresponding kit configuration information to obtain original data and gating and grouping parameters; the original data is a matrix of all cell sampling events of the sample files; obtaining original binary data under the gating channel coordinates, removing invalid data after twice screening to obtain screened gating channel data; obtaining a gating convex polygon and grouping data; obtaining binary data of a report fluorescence channel ReportX and a classification fluorescence channel ClassifyY after gating and grouping, and grouping to obtain factor grouping of the sample files; and performing statistical calculation on the factor grouping data of the sample files to obtain factor microsphere numbers and average fluorescence intensities. The application can automatically pre-process and screen multiple protein flow data, remove impurity data, more accurately retain effective data, automatically gate and group the screened data, and obtain results through statistical analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of flow cytometry multifactor detection technology, and particularly relates to a gating and clustering method and apparatus based on multiple protein flow cytometry data. Background Technology

[0002] In recent years, in vitro diagnostic reagents have developed rapidly, and the types of flow cytometry in vitro detection kits are constantly evolving. However, the distribution pattern of flow cytometry data (FSC-A and SSC-A coordinate data) is uncertain, and it is mostly skewed. For various reasons, it may be due to the presence of a lot of cell debris, which may expand the overall data range; or cell adhesion may occur, resulting in a lot of invalid data.

[0003] Currently, flow cytometry multifactor detection data all require manual removal of some invalid data, followed by automatic or manual data clustering to extract valid data. This results in a complicated, time-consuming, and labor-intensive process for interpreting flow cytometry multifactor detection results. Summary of the Invention

[0004] To address the shortcomings described in the prior art, this invention provides a method and apparatus for gating and clustering multiplex protein flow cytometry data. This method can automatically exclude impurity data in the sample and perform fully automated gating and clustering to obtain the original binary data clusters. The fluorescence channel is automatically gated to accurately classify and extract sample factor data.

[0005] The technical solution adopted in this invention is as follows:

[0006] A gating and clustering method based on multiplex protein flow cytometry data screening includes the following steps:

[0007] Step 1: Read the sample flow cytometry data file and the corresponding reagent kit configuration information to obtain the raw data and gating / grouping parameters;

[0008] The original data is a matrix of all cell sampling events in the sample file;

[0009] The gate clustering parameters include gate channel coordinates, number of gates, factor classification channel coordinates, and number of sample factor clusters; the gate channels include forward channel FSC and lateral channel SSC.

[0010] The factor classification channels include the report fluorescence channel ReportX and the classification fluorescence channel ClassifyY.

[0011] Step 2: Obtain the original binary data under the coordinates of the set gate channel, and remove invalid data after two screenings to obtain the screened set gate channel data;

[0012] Step 21: Based on the gate channel coordinates, obtain the original binary data of the forward channel FSC and the original binary data of the lateral channel SSC from the original data, which is the data to be filtered;

[0013] Step 22: Calculate the data range, average, median, and ratio of the forward channel FSC and the lateral channel SSC of the data to be filtered; the ratio is the ratio of the average to the median.

[0014] Step 23: Determine the difference between the number of data points in the forward channel FSC and the set number of FSCs. If the number of data points in the forward channel FSC is less than the set number of FSCs, proceed to step 3. If the number of data points in the forward channel FSC is greater than or equal to the set number of FSCs, remove impurity data from both the forward channel FSC data and the side channel SSC data to obtain the preprocessed data for the forward channel FSC and the preprocessed data for the side channel SSC.

[0015] The rules for removing impurity data are based on the data range and ratio relationship between the mean and median, identifying and removing impurity data to obtain preprocessed data to be screened. The number of data points refers to the number of rows in the cell sampling event matrix under the given gate channel.

[0016] The specific steps for removing impurity data are as follows:

[0017] xAvg is the average value of the forward channel FSC data;

[0018] The xMid is the median of the forward channel FSC data;

[0019] The xRange is the forward channel FSC data range;

[0020] The xRatio is the ratio of the mean to the median of the forward channel FSC data;

[0021] The yAvg is the average value of the lateral channel SSC data;

[0022] The yMid is the median of the lateral channel SSC data;

[0023] The yRange is the lateral channel SSC data range;

[0024] The yRatio is the ratio of the mean to the median of the lateral channel SSC data;

[0025] Step 231: If the condition yAvg < 20% * yRange && (7.5 < yRatio < 36) || (3.3 < yRatio && 3.3 > xRatio) || (2.3 < yRatio && 3.0 > xRatio > 2.0) or 15% * yRange < yAvg < 20% * yRange && yMid < 10% * yRange && yRatio > 2 && xRatio > 0.98 holds, remove the lateral channel SSC data less than 2.5% * yRange and the corresponding forward channel FSC data, obtain new data to be screened, and return to Step 22 to start execution;

[0026] Step 232: If the condition yAvg < 23% * yRange && yMid < 10% * yRange && yRatio > 1.2) && yRatio < 2.6 or yAvg < 20% * yRange && 1.2 < yRatio < 1.85 && 1.0 < xRatio < 2.3 holds, remove the forward channel SSC data greater than 80% * yRange and the corresponding forward channel FSC data, obtain new data to be screened, and return to Step 22 to start execution;

[0027] Step 233: If the condition (yAvg > 22% * yRange || (yAvg < 22% * yRange && (xRatio < 0.98) || xRatio > 2.0))) && yMid < 16.5% * yRange && yRatio > 1.2 holds, remove the lateral channel SSC data less than 2.5% * yRange and the corresponding forward channel FSC data, obtain new data to be screened, and return to Step 22 to start execution;

[0028] Step 234: If the condition yAvg < 28% * yRange && yMid < 16% * yRange && yRatio >= 1.0) && 1.0 < xRatio < 2.0) or yAvg < 20% * yRange && yRatio > 1.0 or yMid < 20% * yRange && yRatio < 1.0 holds, remove the lateral channel SSC data greater than 80% * yRange and the corresponding forward channel FSC data, obtain new data to be screened, and return to Step 22 to start execution;

[0029] Step 235: If the condition xAvg < 12.5% * xRange && xRatio > 3.5 holds, remove the forward channel FSC data less than 3.33% * xRange and the corresponding lateral channel SSC data, obtain new data to be screened, and return to Step 22 to start execution;

[0030] Step 236: If the condition xMid < 20% * xRange && xRatio < 1.0 or xMid < 33.3% * xRange && xRatio < 1.0 && yRatio < 1.0 is met, remove the forward channel FSC data and the corresponding lateral channel SSC data that are greater than 80% * xRange, obtain new data to be filtered, and return to step 22 to start execution;

[0031] Step 237: If the conditions xRatio>1.1&&yRatio>1.42 are met, discard the lateral channel SSC data and the corresponding forward channel FSC data that are less than 2.5%*yRange, and discard the forward channel FSC data and the corresponding lateral channel SSC data that are less than 3.33%*xRange, to obtain new data to be filtered and return to step 22 to start execution;

[0032] Step 238: If the conditions xAvg < 20% * xRange && xRatio > = 1.0, or xAvg < 33.3% * xRange && xMid > 25% * xRange && xRatio > 1.0, or xMid < 42% * xRange && 0.8 <= xRatio < 0.91 are met, remove the forward channel FSC data and the corresponding lateral channel SSC data that are greater than 80% * xRange, obtain new data to be filtered, and return to step 22 to start execution;

[0033] Step 239: If yMid>24%*yRange&&yAvg>62.5%*yRange&&yRatio>1.0 or yMid<33.3%*yRange&&yAvg>25%*yRange&&yRatio<1.0 is true, remove the lateral channel SSC data and the corresponding forward channel FSC data that are greater than 80%*yRange, obtain new data to be filtered, and return to step 22 to start execution;

[0034] Step 2310: If the conditions xMid>28.5%*xRange&&xAvg>50%*xRange&&xRatio>1.0 or xMid<25%*xRange&&xAvg<67%*xRange&&xAvg>50%*xRange&&xRatio>=1.1 or xAvg>50%*xRange&&xAvg<67%*xRange&&xMid<25%*xRange&&xRatio>1.0 are met, remove the forward channel FSC data and the corresponding lateral channel SSC data that are less than 3.33%*xRange, obtain new data to be filtered, and return to step 22 to start execution;

[0035] Step 24: Determine the difference between the number of data points in the lateral channel SSC and the set number of SSCs. If the number of data points in the lateral channel SSC is less than the set number of SSCs, proceed to step 3. If the number of data points in the lateral channel SSC is greater than or equal to the set number of SSCs, remove the deviation data from the preprocessed data of the forward channel FSC and the preprocessed binary data of the lateral channel SSC according to the optimized 3σ principle to obtain the screening data of the forward channel FSC and the screening data of the lateral channel SSC.

[0036] The number of data points refers to the number of rows in the cell sampling event matrix under the gate channel;

[0037] The steps for removing off-target data are as follows:

[0038] Step 241: Obtain intermediate data for the forward channel FSC and the lateral channel SSC;

[0039] Remove the maximum and minimum values ​​from the preprocessed data of the forward channel FSC to obtain the intermediate data of the forward channel FSC; remove the maximum and minimum values ​​from the preprocessed data of the side channel SSC to obtain the intermediate data of the side channel SSC.

[0040] Step 242: Calculate the median of the forward channel FSC intermediate data and the median of the lateral channel SSC intermediate data respectively:

[0041] Step 243: Calculate the standard deviation of the intermediate data for the forward channel FSC and the lateral channel SSC respectively, replacing the mean with the median:

[0042] Step 244: Based on the 3σ principle, remove the deviation data from the intermediate data of the forward channel FSC and the deviation data from the intermediate data of the lateral channel SSC respectively, to obtain the screening data of the forward channel FSC and the screening data of the lateral channel SSC:

[0043] The exclusion criteria are to exclude data that are greater than the median + 3 * standard deviation and data that are less than the median - 3 * standard deviation.

[0044] Step 3: Obtain the convex hull polygon and cluster data of the gate;

[0045] Step 31: Obtain the gate convex hull polygon of the original binary data after screening;

[0046] Step 311: Standardize the raw binary data within the limited range of the screening data of the forward channel FSC and the screening data of the lateral channel SSC according to the compression ratio n, and divide the standardized data into n groups.

[0047] Step 312: Obtain a binary Boolean array of length n:

[0048] Compare the forward FSC mean of each standardized data set with the number of data points * n of the forward FSC screening data. 2 *The relationship between the threshold values;

[0049] If the mean of the forward channel FSC of the standardized data set is greater than or equal to the number of data points in the forward channel FSC screening data set * n 2 *The threshold is true; if the mean of the forward channel FSC of the standardized data is less than the number of data points in the forward channel FSC screening data * n 2 If the threshold is not met, the result is false, resulting in a binary boolean array of length n.

[0050] The threshold is not less than 2 and not greater than 60. The larger the threshold, the smaller the area of ​​the convex hull polygon of the gate.

[0051] Step 313: Obtain the gate convex hull polygon of the original binary data after filtering;

[0052] Step 3131: Select the binary Boolean points in the binary Boolean array whose top, bottom, left, and right sides are all false, and set them as the boundary points of the gate's convex hull;

[0053] Step 3132: Obtain the gate convex hull polygon of the original binary data after filtering by inverse normalization according to the compression ratio n;

[0054] Step 32: Perform gate grouping based on the gate convex hull polygon and the filtered original binary data;

[0055] Step 321: Obtain the filtered original binary data within the numerical range defined by the convex hull polygon of the gate, and use the original binary data as the gate clustering data to be set, and perform gate clustering on the gate clustering data to be set.

[0056] Step 322: When the number of gates is greater than 1, filter the final binary data of the gate grouping based on the number of data in the binary data of the gate grouping and the number of data in the original binary data or the area relationship between the gate convex hull polygons of the gate grouping.

[0057] When using the number of data points as a filtering criterion, the number of data points refers to the number of rows in the cell sampling event matrix under the gate channel:

[0058] If the sum of the number of data points in each gate group of binary data is greater than the number of data points in the original binary data multiplied by 0.9; or if the number of data points in a single gate group of binary data is greater than the number of data points in the original binary data multiplied by 0.85, then adjust the threshold and return to step 312 to obtain the gate convex hull polygon again; otherwise, the final gate group is obtained.

[0059] When using the area of ​​the convex hull polygon of a gate as a screening criterion:

[0060] If the maximum difference in the forward channel FSC of the convex hull polygon of one gate group is greater than the maximum difference in the forward channel FSC of the convex hull polygon of the other gate group multiplied by 3, then adjust the threshold and return to step 312 to obtain the gated convex hull polygon again; otherwise, it is the final gate group.

[0061] Step 4: Obtain the binary data of the ReportX and ClassifyY fluorescence channels after gating and clustering, and perform clustering to obtain the factor clustering of the sample file;

[0062] Step 41: Based on the final gating and grouping binary data, the original data, and the reagent kit configuration information, obtain the binary data of the ReportX and ClassifyY fluorescent channels.

[0063] Step 42: Calculate the peaks and valleys of the ClassifyY data in the classification fluorescence channel, and take the peaks and valleys as the ClassifyY values ​​of the binary data of the classification convex hull polygon;

[0064] Step 43: Cluster the binary data of ReportX and ClassifyY fluorescent channels based on the ClassifyY value of the classification convex hull polygon binary data;

[0065] Step 44: Apply the optimized 3σ principle to the binary data of each cluster to remove off-target data, and obtain the final factor clusters of the sample file;

[0066] Step 5: Perform statistical calculations on the factor cluster data of the sample file to obtain the number of factor microspheres and the average fluorescence intensity.

[0067] This invention also provides a gating and clustering device based on multiplex protein flow cytometry data screening, including an experimental panel creation unit, a data screening unit, a gating unit, and a statistical analysis unit; the experimental panel creation unit includes a sample data reading module and a reagent kit information extraction module; the sample data reading module is used to read sample flow cytometry data files to obtain the original cell sampling event matrix; the reagent kit information extraction module is used to extract reagent kit configuration information to obtain gating and clustering parameters;

[0068] The data filtering unit includes a data preprocessing module and an optimized 3σ filtering module;

[0069] The data preprocessing module removes some invalid data based on the relationship between the mean and median and the data distribution pattern, making the data more likely to be normally distributed.

[0070] The optimized 3σ screening module removes the maximum and minimum values ​​from the data to be processed and then replaces the average value of the traditional 3σ principle with the median for calculation, thus eliminating deviation data.

[0071] The data is categorized into gating units, and the filtered data is then grouped into gating units. The optimal gating units are determined by evaluating the grouped data. Finally, factor grouping is performed to obtain the factor grouping of the sample files.

[0072] The statistical analysis unit includes a statistical unit used to statistically calculate the number of factor microspheres and the average fluorescence intensity.

[0073] As a preferred embodiment of the present invention, the experimental panel creation unit further includes a standard curve setting module, which sets information such as the concentration and dilution factor of the standard sample for the flow cytometry experiment.

[0074] As a preferred embodiment of the present invention, the statistical analysis unit further includes a standard curve fitting module, which uses four parameters of the logarithm of the data, but is not limited to four parameters of the logarithm of the data.

[0075] This invention can automatically preprocess and screen multiplex protein flow cytometry data, remove impurity data, and more accurately retain valid data. It can also automatically gating and cluster the screened data and perform statistical analysis to obtain the results. Based on the data preprocessing and screening functions of the device of this invention, data containing a lot of impurities can be screened and processed, and finally gating can be used to obtain valid results, thus expanding the data applicability of the method of this invention. Moreover, the screening, gating and clustering of data and analysis can all be completed automatically with one click based on the method of this invention, reducing manual operation and improving detection efficiency. Attached Figure Description

[0076] Figure 1 This is a flowchart of the screening and grouping method of the present invention.

[0077] Figure 2 This is a schematic diagram of the screening and grouping device of the present invention.

[0078] Figure 3 This is a scatter plot of gate data from Example 1 of the present invention without preprocessing and filtering.

[0079] Figure 4 This is a scatter plot of the gated data after screening, as shown in Example 1 of the present invention.

[0080] Figure 5 This is a scatter plot of the data after gate grouping in Example 1 of the present invention.

[0081] Figure 6 This is a scatter plot of the data after factor clustering in Example 1 of the present invention. Detailed Implementation

[0082] This invention provides an embodiment of a gating and clustering method based on multiplex protein flow cytometry data screening, comprising the following steps: Figure 1 As shown,

[0083] Step 1: Read the sample flow cytometry data file and the corresponding reagent kit configuration information to obtain the raw data and gating / grouping parameters;

[0084] The original data is a matrix of all cell sampling events in the sample file;

[0085] The gate clustering parameters include gate channel coordinates, number of gates, factor classification channel coordinates, and number of sample factor clusters; the gate channels include forward channel FSC and lateral channel SSC.

[0086] The factor classification channels include the report fluorescence channel ReportX and the classification fluorescence channel ClassifyY.

[0087] Step 2: Obtain the original binary data under the coordinates of the set gate channel, and remove invalid data after two screenings to obtain the screened set gate channel data;

[0088] Step 21: Based on the gate channel coordinates, obtain the original binary data of the forward channel FSC and the original binary data of the lateral channel SSC from the original data, which is the data to be filtered;

[0089] Step 22: Calculate the data range, average, median, and ratio of the forward channel FSC and the lateral channel SSC of the data to be filtered; the ratio is the ratio of the average to the median.

[0090] Step 23: Determine the size of the number of data points in the forward channel FSC compared to the set number of FSCs. In this embodiment, the set number of FSCs is 1000. If the number of data points in the forward channel FSC is less than the set number of FSCs, proceed to step 3. If the number of data points in the forward channel FSC is greater than or equal to the set number of FSCs, remove impurity data from both the forward channel FSC data and the side channel SSC data to obtain preprocessed data for the forward channel FSC and preprocessed data for the side channel SSC. The number of data points refers to the number of rows in the cell sampling event matrix under the gate channel.

[0091] The rules for removing impurity data are based on the data range and ratio relationship between the mean and median, identifying and removing impurity data to obtain preprocessed data to be screened; the specific steps are as follows:

[0092] xAvg is the average value of the forward channel FSC data;

[0093] The xMid is the median of the forward channel FSC data;

[0094] The xRange is the forward channel FSC data range;

[0095] The xRatio is the ratio of the average value to the median value of the forward channel FSC data;

[0096] The yAvg is the average value of the sideward channel SSC data;

[0097] The yMid is the median value of the sideward channel SSC data;

[0098] The yRange is the sideward channel SSC data range;

[0099] The yRatio is the ratio of the average value to the median value of the sideward channel SSC data;

[0100] Step 231: If the condition yAvg < 20% * yRange && (7.5 < yRatio < 36) || (3.3 < yRatio && 3.3 > xRatio) || (2.3 < yRatio && 3.0 > xRatio > 2.0) or 15% * yRange < yAvg < 20% * yRange && yMid < 10% * yRange && yRatio > 2 && xRatio > 0.98 holds, filter out the sideward channel SSC data less than 2.5% * yRange and the corresponding forward channel FSC data, obtain new data to be screened, and return to Step 22 to start execution;

[0101] Step 232: If the condition yAvg < 23% * yRange && yMid < 10% * yRange && yRatio > 1.2) && yRatio < 2.6 or yAvg < 20% * yRange && 1.2 < yRatio < 1.85 && 1.0 < xRatio < 2.3 holds, filter out the sideward channel SSC data greater than 80% * yRange and the corresponding forward channel FSC data, obtain new data to be screened, and return to Step 22 to start execution;

[0102] Step 233: If the condition (yAvg > 22% * yRange || (yAvg < 22% * yRange && (xRatio < 0.98) || xRatio > 2.0))) && yMid < 16.5% * yRange && yRatio > 1.2 holds, filter out the sideward channel SSC data less than 2.5% * yRange and the corresponding forward channel FSC data, obtain new data to be screened, and return to Step 22 to start execution;

[0103] Step 234: If the conditions yAvg < 28% * yRange && yMid < 16% * yRange && yRatio >= 1.0) && 1.0 < xRatio < 2.0) or yAvg < 20% * yRange && yRatio > 1.0 or yMid < 20% * yRange && yRatio < 1.0 are satisfied, eliminate the lateral channel SSC data greater than 80% * yRange and the corresponding forward channel FSC data, obtain new data to be screened, and return to Step 22 to start execution;

[0104] Step 235: If the conditions xAvg < 12.5% * xRange && xRatio > 3.5 are satisfied, eliminate the forward channel FSC data less than 3.33% * xRange and the corresponding lateral channel SSC data, obtain new data to be screened, and return to Step 22 to start execution;

[0105] Step 236: If the conditions xMid < 20% * xRange && xRatio < 1.0 or xMid < 33.3% * xRange && xRatio < 1.0) && yRatio < 1.0 are satisfied, eliminate the forward channel FSC data greater than 80% * xRange and the corresponding lateral channel SSC data, obtain new data to be screened, and return to Step 22 to start execution;

[0106] Step 237: If the conditions xRatio > 1.1 && yRatio > 1.42 are satisfied, eliminate the lateral channel SSC data less than 2.5% * yRange and the corresponding forward channel FSC data and eliminate the forward channel FSC data less than 3.33% * xRange and the corresponding lateral channel SSC data, obtain new data to be screened, and return to Step 22 to start execution;

[0107] Step 238: If the conditions xAvg < 20% * xRange && xRatio >= 1.0 or xAvg < 33.3% * xRange && xMid > 25% * xRange && xRatio > 1.0 or xMid < 42% * xRange && 0.8 <= xRatio < 0.91 are satisfied, eliminate the forward channel FSC data greater than 80% * xRange and the corresponding lateral channel SSC data, obtain new data to be screened, and return to Step 22 to start execution;

[0108] Step 239: If yMid>24%*yRange&&yAvg>62.5%*yRange&&yRatio>1.0 or yMid<33.3%*yRange&&yAvg>25%*yRange&&yRatio<1.0 is true, remove the lateral channel SSC data and the corresponding forward channel FSC data that are greater than 80%*yRange, obtain new data to be filtered, and return to step 22 to start execution;

[0109] Step 2310: If the conditions xMid>28.5%*xRange&&xAvg>50%*xRange&&xRatio>1.0 or xMid<25%*xRange&&xAvg<67%*xRange&&xAvg>50%*xRange&&xRatio>=1.1 or xAvg>50%*xRange&&xAvg<67%*xRange&&xMid<25%*xRange&&xRatio>1.0 are met, remove the forward channel FSC data and the corresponding lateral channel SSC data that are less than 3.33%*xRange, obtain new data to be filtered, and return to step 22 to start execution.

[0110] After preprocessing, data that is too small due to cell adhesion and too large due to cell debris are removed, narrowing the data range. This step replaces the manual adjustment of the gating data range used in existing gating techniques, shortening data processing time and improving efficiency.

[0111] Step 24: Determine the difference between the number of data points in the lateral channel SSC and the set number of SSCs. In this embodiment, the set number of SSCs is 400. If the number of data points in the lateral channel SSC is less than the set number of SSCs, proceed to step 3. If the number of data points in the lateral channel SSC is greater than or equal to the set number of SSCs, remove the deviation data from the preprocessed data of the forward channel FSC and the preprocessed data of the lateral channel SSC according to the optimized 3σ principle to obtain the screening data of the forward channel FSC and the screening data of the lateral channel SSC.

[0112] The number of data points refers to the number of rows in the cell sampling event matrix under the gate channel;

[0113] The steps for removing off-target data are as follows:

[0114] Step 241: Obtain intermediate data for the forward channel FSC and the lateral channel SSC:

[0115] Remove the maximum and minimum values ​​from the preprocessed data of the forward channel FSC to obtain the intermediate data of the forward channel FSC; remove the maximum and minimum values ​​from the preprocessed data of the side channel SSC to obtain the intermediate data of the side channel SSC.

[0116] Step 242: Calculate the median of the forward channel FSC intermediate data and the median of the lateral channel SSC intermediate data respectively:

[0117] Step 243: Calculate the standard deviation of the intermediate data for the forward channel FSC and the lateral channel SSC respectively, replacing the mean with the median:

[0118] Step 244: Based on the 3σ principle, remove the deviation data from the intermediate data of the forward channel FSC and the deviation data from the intermediate data of the lateral channel SSC respectively, to obtain the screening data of the forward channel FSC and the screening data of the lateral channel SSC:

[0119] The exclusion criteria are to exclude data that are greater than the median + 3 * standard deviation and data that are less than the median - 3 * standard deviation.

[0120] The traditional 3σ principle uses the mean to calculate the standard deviation, with the selection range being the mean + 3 * standard deviation and less than the mean - 3 * standard deviation. For multimodal data with varying peak values, this can lead to accidental deletion of data. However, this application targets streaming data, which is characterized by a multimodal distribution with varying peak values. Therefore, the traditional 3σ principle is not applicable to this application.

[0121] However, regardless of the data distribution, the median is closer to the data peak than the mean. Therefore, the median is used to replace the mean in the traditional 3σ principle to eliminate discrete data.

[0122] Furthermore, in order to maximize the removal of deviation data, this application first removes the maximum and minimum values ​​of the data, and then calculates the median and standard deviation.

[0123] Step 3: Obtain the convex hull polygon and cluster data of the gate;

[0124] Step 31: Obtain the gate convex hull polygon of the original binary data after screening;

[0125] Step 311: Standardize the original binary data within the limited range of the screening data of the forward channel FSC and the screening data of the lateral channel SSC according to the compression ratio n, and divide the standardized data into n groups; in this embodiment, the compression ratio n is 100.

[0126] Step 312: Obtain a binary Boolean array of length n:

[0127] Compare the forward FSC mean of each standardized data set with the number of data points * n of the forward FSC screening data. 2 *The relationship between the threshold values;

[0128] If the mean of the forward channel FSC of the standardized data set is greater than or equal to the number of data points in the forward channel FSC screening data set * n 2 *The threshold is true; if the mean of the forward channel FSC of the standardized data is less than the number of data points in the forward channel FSC screening data * n 2 If the threshold is not met, the result is false, resulting in a binary boolean array of length n.

[0129] The threshold is not less than 2 and not greater than 60. The larger the threshold, the smaller the area of ​​the convex hull polygon of the gate.

[0130] Step 313: Obtain the gate convex hull polygon of the original binary data after filtering;

[0131] Step 3131: Select the binary Boolean points in the binary Boolean array whose top, bottom, left, and right sides are all false, and set them as the boundary points of the gate's convex hull;

[0132] Step 3132: Obtain the gate convex hull polygon of the original binary data after filtering by inverse normalization according to the compression ratio n;

[0133] Step 32: Perform gate grouping based on the gate convex hull polygon and the filtered original binary data;

[0134] Step 321: Obtain the filtered original binary data within the numerical range defined by the convex hull polygon of the gate, and use the original binary data as the gate clustering data to be set, and perform gate clustering on the gate clustering data to be set.

[0135] Step 322: When the number of gates is greater than 1, filter the final binary data of the gate grouping based on the number of data in the binary data of the gate grouping and the number of data in the original binary data or the area relationship between the gate convex hull polygons of the gate grouping.

[0136] When using the number of data points as a filtering criterion, the number of data points refers to the number of rows in the cell sampling event matrix under the gate channel:

[0137] If the sum of the number of data points in each gate group of binary data is greater than the number of data points in the original binary data multiplied by 0.9; or if the number of data points in a single gate group of binary data is greater than the number of data points in the original binary data multiplied by 0.85, then adjust the threshold and return to step 312 to obtain the gate convex hull polygon again. In this embodiment, adjusting the threshold means adding 2 to the threshold; otherwise, the final gate group is obtained.

[0138] When using the area of ​​the convex hull polygon of a gate as a screening criterion:

[0139] If the maximum difference in the forward channel FSC of the convex hull polygon of one gate group is greater than the maximum difference in the forward channel FSC of the convex hull polygon of the other gate group multiplied by 3, then adjust the threshold and return to step 312 to obtain the gate convex hull polygon again. In this embodiment, adjusting the threshold is to add 2 to the threshold; otherwise, it is the final gate group.

[0140] Step 4: Obtain the binary data of the ReportX and ClassifyY fluorescence channels after gating and clustering, and perform clustering to obtain the factor clustering of the sample file;

[0141] Step 41: Based on the final gating and grouping binary data, the original data, and the reagent kit configuration information, obtain the binary data of the ReportX and ClassifyY fluorescent channels.

[0142] Step 42: Calculate the peaks and valleys of the ClassifyY data in the classification fluorescence channel, and take the peaks and valleys as the ClassifyY values ​​of the binary data of the classification convex hull polygon;

[0143] Step 43: Cluster the binary data of ReportX and ClassifyY fluorescent channels based on the ClassifyY value of the classification convex hull polygon binary data;

[0144] Step 44: Apply the optimized 3σ principle to the binary data of each cluster to remove off-target data, and obtain the final factor clusters of the sample file;

[0145] Step 5: Perform statistical calculations on the factor cluster data of the sample file to obtain the number of factor microspheres and the average fluorescence intensity; the number of microspheres is the number of binary data in the gate channel or factor cluster channel; the average fluorescence intensity refers to the sum of the binary data of the factor cluster divided by the number of data.

[0146] To further explain the principles of this invention, a specific example will be used to illustrate the above-described method steps.

[0147] In this example, the method is implemented using C# programming.

[0148] This example uses sample data files from a 13-factor flow cytometry kit for screening and grouping.

[0149] The cell sampling event matrix within the file was read, containing a total of 2684 cell events;

[0150] Read the reagent kit configuration information to obtain the following gating and grouping parameters:

[0151] Let the door passages be FSC-A (forward passage) and SSC-A (side passage);

[0152] Assume the number of doors is 2;

[0153] Factor classification channel coordinates: FL4-H (reporter fluorescence channel), FL-H9 (classification fluorescence channel);

[0154] The number of sample factors is 13;

[0155] The average number of microspheres within each factor cluster is shown in the table below:

[0156]

[0157] The original data sets the gate channel data range as follows: FSC-A: (100129, 22396939), SSC-A: (14475, 19464175). The scatter plot is shown as follows. Figure 3 ;

[0158] After preprocessing and filtering the gated data to exclude some adherent cells in the lower left corner and cell debris in the upper right corner, the gated channel data ranges are: FSC-A: (111255.4, 2534858), SSC-A: (16084.1, 2906948.5). The scatter plot is shown as follows. Figure 4 ;

[0159]

[0160] The filtered data is gated to obtain gated convex hull polygons and gated clusters. Based on the relationship between the number of binary data points in each gated cluster and the number of original binary data points, and the area relationship of the convex hull polygons, the smaller clusters in the lower left and upper right corners are excluded, resulting in the final gated cluster binary data. In the gated convex hull polygons, microsphere A accounts for 39% and microsphere B accounts for 48%. The gated cluster scatter plot is shown below. Figure 5 As shown;

[0161] The ReportX and ClassifyY fluorescence channels were then clustered to obtain the data factor clusters of the sample files. The clustering is shown in the scatter plot. Figure 6 , Figure 6 In the text, FL4-H represents ReportX, and FL9-H represents ClassifyY;

[0162] The final number of factor cluster microspheres is shown in the table below:

[0163] 1 2 3 4 5 6 7 8 9 10 11 12 13 163 195 170 165 160 173 162 180 182 196 158 182 175

[0164] Furthermore, the effectiveness of the gating method in this case can be evaluated by comparing the number of microspheres in the data factor cluster of the sample file with the average number of microspheres in the kit configuration information factor * [0.5, 1.5].

[0165] Based on the analysis, it can be determined that the 100% factor clustering results in this case fall within the range of the average number of microspheres per factor in the kit configuration information * [0.5, 1.5], indicating that the method of this application is effective.

[0166] The present invention also provides an embodiment of a gating and clustering device based on multiplex protein flow cytometry data screening, such as... Figure 2 As shown, it includes an experimental panel creation unit, a data filtering unit, a gating unit, and a statistical analysis unit; the experimental panel creation unit includes a sample data reading module and a reagent kit information extraction module; the sample data reading module is used to read sample streaming data files to obtain the original cell sampling event matrix; the reagent kit information extraction module is used to extract reagent kit configuration information to obtain gating and clustering parameters;

[0167] The data filtering unit includes a data preprocessing module and an optimized 3σ filtering module;

[0168] The data preprocessing module removes some invalid data based on the relationship between the mean and median and the data distribution pattern, making the data more likely to be normally distributed.

[0169] The optimized 3σ screening module removes the maximum and minimum values ​​from the data to be processed and then replaces the average value of the traditional 3σ principle with the median for calculation, thus eliminating deviation data.

[0170] The data is categorized into gating units, and the filtered data is then grouped into gating units. The optimal gating units are determined by evaluating the grouped data. Finally, factor grouping is performed to obtain the factor grouping of the sample files.

[0171] The statistical analysis unit includes a statistical unit used to statistically calculate the number of factor microspheres and the average fluorescence intensity.

[0172] In order to directly output the curve, this embodiment also includes a standard curve setting module in the experimental panel unit. The standard curve setting module sets information such as the concentration and dilution factor of the standard sample for the flow cytometry experiment.

[0173] Furthermore, the statistical analysis unit also includes a standard curve fitting module. This module uses four parameters, but is not limited to, the logarithm of the data for fitting, and directly outputs the fitted curve for researchers to view intuitively.

[0174] All or part of the steps of the various methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include: read-only memory, random access memory, disk or optical disk, etc.

[0175] The present application has been described above by way of example, but the present application is not limited to the specific embodiments described above. Any modifications or variations made based on the present application shall fall within the scope of protection claimed by the present application.

Claims

1. A method for screening, gating and clustering based on multiplexed protein flow data, characterized in that, The method comprises the following steps: step 1, reading sample stream data files and corresponding kit configuration information to obtain original data and gating cluster parameters; step 2, obtaining original binary data under gating channel coordinates and removing invalid data to obtain gating channel data after twice screening; step 3, obtaining gating convex polygon and cluster data; step 4, obtaining binary data of report fluorescence channel ReportX and classification fluorescence channel ClassifyY after gating cluster and performing cluster to obtain factor cluster of sample files; and step 5, performing statistical calculation on factor cluster data of sample files to obtain factor microsphere number and average fluorescence intensity. The original data is a matrix of all cell sampling events of sample files. The gating channel includes a forward channel FSC and a side channel SSC. The factor classification channel includes a report fluorescence channel ReportX and a classification fluorescence channel ClassifyY. The step 2 comprises the following steps: Step 21, obtaining original binary data of the forward channel FSC and the side channel SSC from the original data according to the gating channel coordinates, that is, the data to be screened; Step 22, calculating the data range, average value, median value and ratio of the forward channel FSC and the side channel SSC of the data to be screened respectively; the ratio is the ratio of the average value to the median value; Step 23, judging the size of the data number of the forward channel FSC and the FSC set number, if the data number of the forward channel FSC is less than the FSC set number, jumping to step 3, if the data number of the forward channel FSC is greater than or equal to the FSC set number, removing impurity data from the data of the forward channel FSC and the data of the side channel SSC respectively to obtain preprocessed data of the forward channel FSC and the side channel SSC; The removal rule of the impurity data is to judge and remove the impurity data according to the data range and the ratio of the average value to the median value to obtain preprocessed data to be screened; The specific steps of removing the impurity data are as follows: xAvg is the average value of the forward channel FSC data; xMid is the median value of the forward channel FSC data; xRange is the data range of the forward channel FSC data; xRatio is the ratio of the average value to the median value of the forward channel FSC data; yAvg is the average value of the side channel SSC data; yMid is the median value of the side channel SSC data; yRange is the data range of the side channel SSC data; yRatio is the ratio of the average value to the median value of the side channel SSC data; ​ ​ ​ Step 24: judging the size of the data number of the lateral channel SSC and the SSC setting number, if the data number of the lateral channel SSC < the SSC setting number, jumping to step 3; if the data number of the lateral channel SSC ≥ the SSC setting number, removing the deviated data from the pretreated data of the forward channel FSC and the pretreated binary data of the lateral channel SSC according to the optimized 3σ principle, to obtain the screening data of the forward channel FSC and the screening data of the lateral channel SSC; The specific steps of the step 3 are: Step 31: obtaining the set door convex polygon of the original binary data in the range; Step 311: standardizing the original binary data in the range of the screening data of the forward channel FSC and the screening data of the lateral channel SSC according to the compression ratio n, and dividing the standardized data into n groups; Step 312: obtaining a binary Boolean array with a length of n: Comparing the forward channel FSC mean of the normalized data to the number of data *n of the forward channel FSC fractionated data for each group 2 * The size relationship of the threshold value; If the forward channel FSC mean of the group normalized data >= the data number *n of the forward channel FSC screening data 2 * threshold, then true; if the forward channel FSC mean of the group normalized data < the data number *n of the forward channel FSC screening data 2 * threshold, then false, to obtain a binary Boolean array with a length of n; Step 313: obtaining the set door convex polygon of the screened original binary data according to the binary Boolean array; Step 32: setting door grouping according to the set door convex polygon and the screened original binary data; Step 321: obtaining the screened original binary data in the range defined by the set door convex polygon, and taking the original binary data as the to-be-set-door-grouping data, and setting door grouping for the to-be-set-door-grouping data; Step 322: when the number of set doors is greater than 1, screening the final set door grouping binary data according to the area relationship between the data number of the set door grouping binary data and the data number of the original binary data or the set door convex polygon of the set door grouping.

2. The method of claim 1, wherein the method is based on multiplexed protein flow data screening gating. The step 23 includes: Step 231: if the conditions yAvg < 20% * yRange && (7.5 < yRatio < 36) || (3.3 < yRatio && 3.3 < xRatio) || (2.3 < yRatio && 3.0 < xRatio < 2.0) or 15% * yRange < yAvg < 20% * yRange && yMid < 10% * yRange && yRatio > 2 && xRatio > 0.98 are met, removing the lateral channel SSC data and the corresponding forward channel FSC data less than 2.5% * yRange, to obtain new to-be-screened data and returning to step 22 for execution; Step 232: if the conditions yAvg < 23% * yRange && yMid < 10% * yRange && yRatio > 1.2) && yRatio < 2.6 or yAvg < 20% * yRange && 1.2 < yRatio < 1.85 && 1.0 < xRatio < 2.3 are met, removing the lateral channel SSC data and the corresponding forward channel FSC data greater than 80% * yRange, to obtain new to-be-screened data and returning to step 22 for execution; Step 233: If the condition (yAvg>22%*yRange||(yAvg<22%*yRange&&(xRatio<0.98)||xRatio>2.0)))&&yMid<16.5%*yRange&&yRatio>1.2 is true, the lateral channel SSC data and corresponding forward channel FSC data less than 2.5%*yRange are removed, the new data to be screened is obtained, and the step 22 is returned to start execution; Step 234: If the condition yAvg<28%*yRange&&yMid<16%*yRange&&yRatio>=1.0)&&1.0<xRatio<2.0) or yAvg<20%*yRange&&yRatio>1.0 or yMid<20%*yRange&&yRatio<1.0 is true, the lateral channel SSC data and corresponding forward channel FSC data greater than 80%*yRange are removed, the new data to be screened is obtained, and the step 22 is returned to start execution; Step 235: If the condition xAvg<12.5%*xRange&&xRatio>3.5 is true, the forward channel FSC data and corresponding lateral channel SSC data less than 3.33%*xRange are removed, the new data to be screened is obtained, and the step 22 is returned to start execution; Step 236: If the condition xMid<20%*xRange&&xRatio<1.0 or xMid<33.3%*xRange&&xRatio<1.0)&&yRatio<1.0 is true, the forward channel FSC data and corresponding lateral channel SSC data greater than 80%*xRange are removed, the new data to be screened is obtained, and the step 22 is returned to start execution; Step 237: If the condition xRatio>1.1&&yRatio>1.42 is true, the lateral channel SSC data and corresponding forward channel FSC data less than 2.5%*yRange are removed, and the forward channel FSC data and corresponding lateral channel SSC data less than 3.33%*xRange are removed, the new data to be screened is obtained, and the step 22 is returned to start execution; Step 238: If the condition xAvg<20%*xRange&&xRatio>=1.0 or xAvg<33.3%*xRange&&xMid>25%*xRange&&xRatio>1.0 or xMid<42%*xRange&&0.8<=xRatio<0.91 is true, the forward channel FSC data and corresponding lateral channel SSC data greater than 80%*xRange are removed, the new data to be screened is obtained, and the step 22 is returned to start execution; Step 239: If the condition yMid>24%*yRange&&yAvg>62.5%*yRange&&yRatio>1.0 or yMid<33.3%*yRange&&yAvg>25%*yRange&&yRatio<1.0 is met, remove the side scatter SSC data and the corresponding forward scatter FSC data greater than 80%*yRange, obtain new data to be screened and return to step 22 for execution; Step 2310: If the condition xMid>28.5%*xRange&&xAvg>50%*xRange&&xRatio>1.0 or xMid<25%*xRange&&xAvg<67%*xRange&&xAvg>50%*xRange&&xRatio>=1.1 or xAvg>50%*xRange&&xAvg<67%*xRange&&xMid<25%*xRange&&xRatio>1.0 is met, remove the forward scatter FSC data and the corresponding side scatter SSC data less than 3.33%*xRange, obtain new data to be screened and return to step 22 for execution.

3. The method according to claim 1, wherein the step of removing the deviated data comprises: the step of removing the deviated data comprises: Step 241: obtaining the median data of the forward scatter FSC and the median data of the side scatter SSC; Step 242: calculating the median of the median data of the forward scatter FSC and the median of the median data of the side scatter SSC, respectively; Step 243: calculating the standard deviation of the median data of the forward scatter FSC and the standard deviation of the median data of the side scatter SSC, respectively, by using the median to replace the mean; Step 244: removing the deviated data of the median data of the forward scatter FSC and the deviated data of the median data of the side scatter SSC according to the removal condition of 3σ principle, to obtain the screening data of the forward scatter FSC and the screening data of the side scatter SSC; the removal condition is to remove the data greater than the median+3*standard deviation and the data less than the median-3*standard deviation. the step of obtaining the gated convex polygon comprises:

4. The method of claim 1, wherein the method is based on multiplexed protein flow data screening gating. Step 3131: selecting the binary Boolean value as the gated convex boundary point when the binary Boolean values of the upper, lower, left and right of the binary Boolean array are all false; Step 3132: obtaining the gated convex polygon of the screened original binary data according to the inverse standardization of the compression ratio n. in step 322, when the number of data is used as the screening standard, 5. The method of claim 1, wherein the method is based on multiplexed protein flow data screening gating and clustering. ​ If the sum of the data quantity of each gated cluster binary data is greater than 0.9 times of the data quantity of the original binary data, or the data quantity of a single gated cluster binary data is greater than 0.85 times of the data quantity of the original binary data, the threshold is adjusted and the step 312 is returned to obtain the gated convex polygon again, otherwise the final gated cluster is obtained. When the area of the gated convex polygon is used as the screening standard: If the maximum difference of the forward scatter channel FSC in the convex polygon of one of the gated clusters is greater than 3 times of the maximum difference of the forward scatter channel FSC in the convex polygon of another gated cluster, the threshold is adjusted and the step 312 is returned to obtain the gated convex polygon again, otherwise the final gated cluster is obtained.

6. The method of claim 5, wherein the method is based on multiplexed protein flow data screening gating. The threshold is adjusted by adding 2 to the threshold, and the threshold is not less than 2 and not greater than 60.

7. The method of claim 4 or 5 or 6, wherein the method is based on multiplexed protein flow data screening gating. The specific steps of the step 4 are: Step 41: obtaining the binary data of the report fluorescence channel ReportX and the classification fluorescence channel ClassifyY according to the binary data of the final gated cluster, the original data and the report fluorescence channel ReportX and the classification fluorescence channel ClassifyY of the kit configuration information; Step 42: calculating the peak and valley of the classification fluorescence channel ClassifyY data, and taking the peak and valley as the ClassifyY value of the classification convex polygon binary data; Step 43: clustering the binary data of the report fluorescence channel ReportX and the classification fluorescence channel ClassifyY according to the ClassifyY value of the classification convex polygon binary data; Step 44: removing the deviating data of each cluster binary data by using the optimized 3σ principle to obtain the final factor cluster of the sample file.

8. A device for screening, gating and grouping based on multiplexed protein flow data, characterized by: The device executes the method of any one of claims 1-7, and includes a experimental panel creating unit, a data screening unit, a gating unit and a statistical analysis unit; The experimental panel creating unit includes a sample data reading module and a kit information extracting module; the sample data reading module is used to read the sample flow data file to obtain the original cell sampling event matrix; and the kit information extracting module is used to extract the kit configuration information to obtain the gated cluster parameters; The data screening unit includes a data preprocessing module and an optimized 3σ screening module; The data preprocessing module removes part of invalid data according to the relationship between the average value and the median and the data distribution form, so that the data tends to be normal distribution; The optimized 3σ screening module removes the maximum value and the minimum value of the data to be processed, replaces the average value of the traditional 3σ principle with the median, and removes the deviating data; The gating unit performs gated cluster on the screened data, judges the clustered data to obtain the optimized gated cluster, and then performs factor cluster to obtain the factor cluster of the sample file; The statistical analysis unit includes a statistical unit, which is used to statistically calculate the number of factor microspheres and the average fluorescence intensity.

9. The apparatus for screening, gating and grouping based on multiplexed protein flow data of claim 8, wherein: The experimental panel creating unit further includes a standard curve setting module, which sets the concentration, dilution multiple and other information of the standard sample of the flow experiment. And / or the statistical analysis unit further comprises a standard curve fitting module, which adopts a four-parameter data logarithm fitting, but is not limited to a four-parameter data logarithm fitting.

Citation Information

Patent Citations

  • Flow cytometry cell data fast automatic grouping and circling method

    CN106548205A

  • Multi-factor cell factor automatic analysis method

    CN113188981A