A data analysis method, system and storage medium based on a bioanalyzer

By employing a two-stage progressive analysis logic and standard curve fitting, the data analysis method of the bioanalyst solves the problems of insufficient automation and poor threshold matching, realizing automated batch parsing of streaming data and improving the stability and accuracy of test results, which is suitable for large-sample clinical testing.

CN122332843APending Publication Date: 2026-07-03ANHUI HUIZHONG TONGDA CANCER EARLY SCREENING RESEARCH INSTITUTE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ANHUI HUIZHONG TONGDA CANCER EARLY SCREENING RESEARCH INSTITUTE CO LTD
Filing Date
2026-06-05
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing ultrasensitive detection technologies based on bioanalyzers suffer from insufficient automation, poor threshold matching, and limited detection dynamic range, resulting in poor consistency and accuracy of detection results, especially with large errors when detecting high and low concentration samples.

Method used

A data analysis method based on a bioanalyzer is adopted, which uses a two-stage progressive analysis logic of standard modeling and batch analysis of samples to be analyzed. Combined with sample peak shape classification processing, F-value interval segmented quantification and automatic standard curve fitting, it realizes automated batch parsing of flow cytometry data and adapts to samples with different concentration distributions.

Benefits of technology

It enables automated batch parsing of streaming data, improves the stability and accuracy of test results, reduces hardware and consumable investment, and is suitable for routine clinical large-sample testing scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122332843A_ABST
    Figure CN122332843A_ABST
Patent Text Reader

Abstract

A data analysis method, system, and storage medium based on a bioanalyzer are disclosed. The samples to be analyzed include the test sample and standards. The method includes: acquiring the target channel signal measured by the bioanalyzer and determining whether the sample to be analyzed is a single-peak or non-single-peak sample based on the histogram of the signal; for non-single-peak samples, calculating the proportion of positive microspheres using the signal intensity corresponding to the deepest trough of the histogram as a first threshold; for single-peak samples, calculating the proportion of positive microspheres based on a second threshold, which is determined based on the first threshold of multiple non-single-peak samples of the standards; and calculating the average target capture amount of a single positive microsphere in the sample to be analyzed based on the interval where the proportion of positive microspheres is located and / or the signal intensity information of the positive microspheres. This method effectively extends the detection dynamic range and achieves automated and standardized data analysis. This application also discloses a corresponding system and computer storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of biological detection data analysis technology, and in particular to a digital detection data full-process analysis method, system and corresponding storage medium based on a biological analyzer. Background Technology

[0002] In clinical disease testing, the concentration of analytes in plasma and other detection systems is often low, sometimes even reaching the level of picograms per milliliter. Traditional detection methods require a 10 to 100-fold increase in concentration to detect them effectively. At the same time, there is also high-abundance protein interference in the peripheral environment of the detection system, which is not easy to solve through preprocessing methods such as concentration. Therefore, ultrasensitive digital detection technology has been developed, which can capture low concentrations of analytes and amplify the signal to a detectable level, providing support for early disease diagnosis.

[0003] Digital detection, with its single-molecule-level capture and counting capabilities, achieves detection sensitivity at the fg / mL level, an improvement of 3-5 orders of magnitude compared to traditional immunoassays. It can accurately capture extremely low concentrations of biomarkers, making it suitable for early disease screening and analysis of trace body fluid samples. Taking microsphere-based digital detection as an example, it captures target substances through microspheres and amplifies and detects signals using individual microspheres as the basic unit. The signal of a single microsphere is interpreted as "0 / 1" for digital output. By establishing a relationship with concentration using Poisson distribution, it effectively avoids background noise and signal fluctuation interference, resulting in strong stability and good repeatability. At the same time, combined with the high-throughput and multi-parameter parallel characteristics of instruments such as flow cytometry or fluorescence microscopy, it enables parallel analysis of multiple samples, significantly shortening the detection cycle and reducing instrument and reagent costs. It has the comprehensive advantages of ultra-high sensitivity, high stability, low cost, and easy automation integration, providing efficient and reliable technical support for clinical ultrasensitive detection.

[0004] Bioanalysts are a type of detection equipment used to perform qualitative, quantitative, or semi-quantitative analysis of biomarkers such as cells, proteins, nucleic acids, and small molecules in biological samples. They are widely used in clinical diagnosis, life science research, drug development, and environmental monitoring, and are core instruments for bioanalysis and medical testing.

[0005] Common types of bioanalysts include: flow cytometers, digital single-molecule array analyzers (SiMoA), electrochemiluminescence analyzers, chemiluminescence analyzers, enzyme-linked immunosorbent assay (ELISA) analyzers, colloidal gold rapid assay analyzers, polymerase chain reaction (PCR) analyzers, gene sequencers, and biosensors. Among these, flow cytometers can perform high-speed, multi-parameter analysis of single cells or microspheres; digital single-molecule array analyzers possess ultra-high sensitivity detection capabilities at the single-molecule level. Both are typical representatives of high-sensitivity biological detection technologies.

[0006] However, while existing SiMoA and traditional flow cytometry digital detection methods can achieve ultrasensitive quantification, existing SiMoA methods rely on high-precision microporous structures, resulting in expensive equipment and limited availability. Current flow cytometry-based digital detection data analysis methods suffer from poor quantitative accuracy and batch stability, model failure under high positive rates, and inconsistent positive / negative clustering thresholds. Therefore, there is a need in this field for a versatile and automated flow cytometry ultrasensitive data analysis method. Summary of the Invention

[0007] The following provides a brief overview of one or more aspects to offer a basic understanding of them. This overview is not an exhaustive summary of all conceived aspects, nor is it intended to identify the key or decisive elements of all aspects, nor to define the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed descriptions that follow.

[0008] As mentioned above, existing ultrasensitive detection technologies based on bioanalyzers currently have many shortcomings: First, insufficient automation: For example, the subsequent analysis steps of FCS format data output by flow cytometer are complex and require a lot of manual intervention. The analysis efficiency is low in large clinical sample scenarios, and the differences in operation by different analysts can lead to poor consistency of results. Second, the threshold division has poor adaptability: existing threshold division methods (such as, but not limited to, clustering algorithms, negative peak normal fitting, etc.) have poor adaptability to extremely low / extremely high concentration samples: low concentration samples have a very high proportion of negative peaks and no obvious positive peaks, while high concentration samples have almost no negative peaks. Both are prone to threshold division deviation, resulting in F value calculation errors exceeding 10%. Third, the detection dynamic range is limited: Pure digital quantification relies solely on the proportion of positive microspheres, F. When F exceeds 70%, most microspheres have already captured a target analyte, and the statistical assumption of Poisson distribution no longer applies, failing to accurately reflect the actual concentration. The dynamic range is only 3 to 4 orders of magnitude. On the other hand, the method of directly taking the average signal of all microspheres is easily affected by signal fluctuations, with low-value samples having an error of over 20%, making it impossible to balance the detection accuracy of high and low concentration samples.

[0009] To this end, this application proposes a digital end-to-end analysis scheme for detection data based on a bioanalyzer. This scheme can classify and analyze samples with different concentration distributions, and perform overall collaborative correction and unified calibration by combining all samples in the same batch. Unlike high-cost single-molecule array detection technologies, this scheme can achieve ultrasensitive quantification using a conventional flow cytometry platform, significantly reducing hardware and consumable investment. Simultaneously, through sample peak shape classification, F-value interval segmented quantification, and automatic standard curve fitting, it achieves automated batch analysis of flow cytometry data, balancing analytical efficiency, result stability, and cost advantages, making it suitable for routine large-sample clinical testing scenarios.

[0010] Specifically, according to the first aspect of this application, a data analysis method based on a bioanalyzer is proposed, wherein the sample to be analyzed includes a test sample and a standard, and the method includes the following steps: The target channel signal measured by the bioanalyzer is acquired, and the histogram of the target channel signal is used to determine whether the sample to be analyzed is a single-peaked sample or a non-single-peaked sample. For the non-single-peak sample, the signal intensity corresponding to the deepest valley of its histogram is used as the first division threshold to calculate the corresponding proportion of positive microspheres. For the single-peak sample, the corresponding percentage of positive microspheres is calculated based on a second partitioning threshold, wherein the second partitioning threshold is determined based on a first partitioning threshold of multiple non-single-peak samples of the standard; and Based on the range of positive microsphere percentage in the sample to be analyzed and / or the signal intensity information of positive microspheres in the sample to be analyzed, the average target capture amount of a single positive microsphere in the sample to be analyzed is calculated.

[0011] According to one embodiment of this application, determining whether a sample to be analyzed is a unimodal sample or a non-unimodal sample based on the histogram includes: extracting the deepest valley in the histogram, obtaining the valley parameter of the deepest valley, the valley parameter including the valley horizontal coordinate, the valley vertical coordinate and the valley depth; and determining whether the sample to be analyzed is a unimodal sample or a non-unimodal sample based on conditions related to the valley parameter.

[0012] According to one embodiment of this application, if at least one of the following conditions is met, the sample to be analyzed is determined to be a unimodal sample; otherwise, the sample to be analyzed is determined to be a non-unimodal sample: The horizontal coordinate of the trough is less than a first threshold; or the horizontal coordinate of the trough is greater than a second threshold; or the vertical coordinate of the trough is less than a third threshold and the depth of the trough is less than a fourth threshold.

[0013] According to one embodiment of this application, the second division threshold is calculated by taking the mean or median after removing outliers from the first division threshold of non-single-peak samples in the first interval where the proportion of positive microspheres is in the first interval.

[0014] According to one embodiment of this application, the first interval is 0.1 to 0.3.

[0015] According to one embodiment of this application, the step of calculating the average target substance capture amount A of a single positive microsphere in the sample to be analyzed based on the interval in which the proportion of positive microspheres in the sample to be analyzed is located and / or the signal intensity information of positive microspheres in the sample to be analyzed includes: screening samples where the proportion of positive microspheres F is in a second interval, determining the mean signal intensity Ib of all positive microspheres in each screened sample and the number of positive microspheres n; using the number of positive microspheres n as the weight, performing a weighted average of the single-molecule signal intensity Isi of the positive microspheres that captured a single target substance in the screened samples to obtain the mean signal intensity Is of the positive microspheres that captured a single target substance in the sample to be analyzed, wherein Isi is based on the mean signal intensity Ib and the proportion of positive microspheres F; and calculating the average target substance capture amount A of a single positive microsphere in the sample to be analyzed based on the interval in which the proportion of positive microspheres F in the sample to be analyzed is located, combined with the mean signal intensity Is of the positive microspheres that captured a single target substance and the mean signal intensity I of the positive microspheres in the sample to be analyzed.

[0016] According to one embodiment of this application, calculating the average target capture amount A of a single positive microsphere in the sample to be analyzed, based on the interval in which the proportion F of positive microspheres in the sample to be analyzed falls, includes: when the proportion F of positive microspheres is less than a first preset percentage threshold, according to...

[0017] Calculate A = -ln(1-F); and

[0018] When the proportion F of the positive microspheres is greater than or equal to the first preset percentage threshold, according to

[0019] Calculate A = (F × I) / Is.

[0020] According to one embodiment of this application, the method further includes: obtaining the average target analyte capture amount corresponding to the standard of known concentration, and fitting a standard curve; and converting the average target analyte capture amount of the test sample into the concentration of the test sample based on the standard curve.

[0021] According to a second aspect of this application, a data analysis system based on a bioanalyte is provided, wherein the sample to be analyzed includes a test sample and a standard, the data analysis system comprising: The sample classification module is used to acquire the target channel signal measured by the bioanalyst and determine whether the sample to be analyzed is a single-peaked sample or a non-single-peaked sample based on the histogram of the target channel signal. The non-single-peak sample processing module is used to calculate the proportion of positive microspheres for non-single-peak samples by using the signal intensity corresponding to the deepest valley of its histogram as the first division threshold. A single-peak sample processing module is used to calculate the proportion of positive microspheres corresponding to a single-peak sample based on a second partitioning threshold, wherein the second partitioning threshold is determined based on a first partitioning threshold of multiple non-single-peak samples of the standard; and The quantitative module is used to calculate the average target capture amount of a single positive microsphere in the sample based on the range of the proportion of positive microspheres in the sample to be analyzed, and in combination with the signal intensity information of the positive microspheres in the sample to be analyzed.

[0022] According to a third aspect of this application, a computer-readable storage medium is provided, the storage medium storing a computer program that, when executed by a processor, implements the data analysis method described above.

[0023] To achieve the foregoing and related objectives, these one or more aspects include the features fully described below and specifically pointed out in the appended claims. Certain illustrative features of these one or more aspects are set forth in detail in the following description and accompanying drawings. However, these features merely indicate a few of the various ways in which the principles of these aspects may be employed, and this description is intended to cover all such aspects and their equivalents. Attached Figure Description

[0024] To gain a detailed understanding of the features described above, a more specific description of the above-briefly summarized aspects can be obtained by referring to the accompanying drawings, some of which are illustrated in the figures. However, it should be noted that the drawings only illustrate certain typical aspects of this application and should not be considered as limiting its scope, as other equivalent aspects are permissible in this description.

[0025] Figure 1 This is a flowchart of the data analysis method based on a bioanalyzer according to this application, which illustrates the various specific steps; Figure 2 This is an example of a signal channel histogram, which shows the smooth curve, trough locations, and division thresholds. Figure 3 This is a schematic diagram of the standard curve analysis interface described in this application, showing functional areas such as standard curve fitting result display and parameter input; Figure 4 The result table in the sample analysis interface described in this application contains intermediate data such as the F value, A value, concentration, and mean fluorescence intensity I of each sample to be analyzed, for users to review and export. Figure 5 This is a flowchart of the core data processing algorithm of the data analysis method based on a biological analyzer described in this application, which shows the execution logic of histogram preprocessing, optimal valley detection, signal type judgment, and threshold template generation and application. Figure 6This application presents a flowchart illustrating the overall architecture of a data analysis system based on a bioanalyst, showcasing the entire process from FCS file reading and data gating to key analysis and standard curve plotting. Figure 7 This is a structural block diagram of the data analysis system based on a bioanalyzer according to this application, which shows the modules used to perform the various steps of the above method. Detailed Implementation

[0026] The specific embodiments of this application will now be described in detail with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit the scope of this application. Those skilled in the art will understand that the embodiments described below are only some embodiments of this application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0027] As mentioned earlier, existing ultrasensitive detection technologies still have many shortcomings. Taking flow cytometry-based analysis as an example: the overall level of automation is low, the subsequent analysis process of flow cytometry FCS data is cumbersome and highly dependent on manual intervention, making it difficult to adapt to the needs of large-scale clinical testing. Moreover, differences in human operation can easily lead to poor consistency of test results. At the same time, traditional threshold division methods are not well adapted to samples with extreme high and low concentrations, and are prone to threshold shifts, resulting in large errors in calculating the proportion of positive microspheres (F). In addition, the detection dynamic range is limited. Pure digital quantification methods no longer meet the statistical premise of Poisson distribution when the proportion of positive microspheres is too high, and quantification fails. Signal mean analysis has weak anti-interference ability and significant errors in low-value samples. Neither of these methods can accurately detect samples with high and low concentrations.

[0028] Therefore, this application proposes a data analysis scheme based on a biological analyzer. The implementation process of this application will be described in detail below with reference to the accompanying drawings.

[0029] The data analysis scheme based on the bioanalyst described in this application adopts a two-stage progressive analysis logic of "standard modeling → batch analysis of samples to be analyzed". It is divided into two major functional units: a preliminary processing module and a classification batch processing optimization module. This solves the problems of poor consistency of results and inaccurate quantification of extreme concentration samples caused by the reliance on manual setting of thresholds for each sample in traditional methods.

[0030] The preliminary processing module is a common pre-processing procedure for all samples to be analyzed (standards and test samples), used to locate the optimal threshold candidate point from the noisy raw fluorescence signal; the classification batch processing optimization module is responsible for intelligent signal type judgment, standardized threshold template generation and batch application, and quantitative data correction.

[0031] The first stage is the standard data analysis stage (executed only once for the first test or when changing batches): After the system preprocesses the standards of all concentration gradients, it automatically generates a standardized threshold template applicable to the entire batch, and constructs a standard curve of "positive microsphere percentage - analyte concentration" in combination with the known concentration of the standards, thus completing the establishment of the quantitative benchmark.

[0032] The second stage is the batch analysis stage of the samples to be analyzed: the system directly applies the established general threshold template and standard curve to all samples to be analyzed, without recalculating the threshold for each sample. After unified data correction, the final concentration result is automatically output, and dozens to hundreds of clinical samples can be processed at the same time.

[0033] First refer to Figure 1 , Figure 1 This describes the specific implementation process of the data analysis method 10 based on a bioanalyzer according to this application.

[0034] This application uses a commonly used flow cytometer as an example to describe the specific implementation of the application. However, as those skilled in the art will understand, various types of bioanalysts known in the art can utilize the data analysis methods described in this application with the aid of a matching data reading interface.

[0035] Before formally executing step S1, the data analysis method 10 may also include a preprocessing step for the sample to be analyzed. Specifically, the system first reads the FCS file output by the flow cytometer.

[0036] Flow cytometer (FCM) is an instrument for high-speed, multi-parameter quantitative analysis (and sorting) of single cells / microspheres in liquids. Its core technology is "single cell + laser + fluorescence + signal clustering". It can measure thousands to tens of thousands of particles per second and is widely used in medicine, biology and ultrasensitive detection.

[0037] FCS (Flow Cytometry Standard) files are a standard format for flow cytometry data (with the .fcs extension) developed by the International Society for Cell Analysis (ISAC). It is used for the standardized storage and exchange of raw data from multi-parameter measurements of single particles / cells. The current mainstream version is FCS 3.1 / 3.2. Of course, other versions of the FCS format are also applicable to the scheme described in this application.

[0038] Data obtained from a single flow cytometry test is packaged by the instrument itself and output as a single FCS (Flow Cytometry System) file. Currently, flow cytometry manufacturers have established universal and standardized FCS file formats (such as FCS 3.0 / 3.1 / 3.2 standards), with uniformly defined encoding formats and data locations. As mentioned above, other FCS file formats are also applicable to this application.

[0039] The system, based on its established structure and common naming conventions, calls the basic functions for reading FCS files to extract the data required for subsequent analysis. When performing high-throughput detection of individual microspheres in a suspension, the flow cytometer simultaneously acquires multi-dimensional detection signals, including forward scattered light (FSC) signals, side scattered light (SSC) signals, and at least one target detection channel signal.

[0040] Among them, the forward scattered light (FSC) signal is the light signal scattered at a small angle along the incident light direction after the laser irradiates the microspheres. Its intensity is positively correlated with the overall particle size and geometric dimensions of the microspheres. In this application, it is used to initially screen microsphere particles that meet the preset size range and exclude fragments, impurities and multi-microsphere aggregates with abnormal sizes. The side scattered light (SSC) signal is the light signal scattered perpendicular to the incident light direction after the laser irradiates the microspheres. Its intensity is related to the internal structural complexity, surface roughness and refractive index of the microspheres. In this application, it is used to further distinguish the microspheres from other non-target particles and achieve precise gate formation of the target microsphere group.

[0041] The target detection channel signal is used to characterize the amount of target analyte bound to the microsphere surface, and its signal intensity is positively correlated with the amount of target analyte bound to the microsphere surface. The target detection channel signal includes, but is not limited to, fluorescent channel signals and non-fluorescent channel signals. Fluorescent channel signals include traditional monochromatic fluorescence channel signals (such as green fluorescence channel, orange fluorescence channel, red fluorescence channel, far-infrared fluorescence channel, near-infrared fluorescence channel, and ultraviolet-excited fluorescence channel, which can specifically correspond to the detection channels of dyes such as phycoerythrin (PE), allophycocyanin (APC), fluorescein isothiocyanate (FITC), and polydinoflavin chlorophyll protein (PerCP), as well as novel fluorescence detection channel signals such as spectral fluorescence channel signals, time-resolved fluorescence channel signals, phosphorescence channel signals, and fluorescence lifetime channel signals; Non-fluorescent channel signals include light absorption channel signals, Raman scattering channel signals, impedance channel signals, mass spectrometry channel signals, chemiluminescence channel signals, etc., which can be used for the detection of quantitative target analytes.

[0042] Those skilled in the art should understand that the specific types of target detection channel signals described above are merely illustrative examples, and any detection channel signal capable of quantitatively characterizing the amount of target analyte bound to the microsphere surface falls within the protection scope of this application.

[0043] In one embodiment of this application, the system plots an FSC-SSC scatter plot and filters out effective microsphere populations by gating (e.g., rectangular or polygonal gates), excluding debris, aggregates, and clustered microspheres. In a specific implementation, the gating process is as follows: The raw data acquired by the flow cytometer contains all detected particle events, which may include target magnetic microspheres, debris, microsphere aggregates, and non-specific particles. To ensure the accuracy of the analysis results, this method first selects the target microsphere population on the FSC-SSC scatter plot using rectangular or polygonal gates, excluding debris (usually due to low FSC signals) and aggregates (usually due to high FSC and SSC signals). The first selected position is saved as a gating template, which can be directly applied to other samples in the same batch without repeated manual operation. For the selected effective microsphere population, the system extracts the signal intensity data of its target fluorescence channels and plots a signal channel histogram with a fixed number of intervals (e.g., 200, 500, 1000 intervals, etc.; 1000 intervals are used in the following example).

[0044] This application presents an interactive dedicated data analysis software (version 1.5.3.3) developed based on the MATLAB platform, supporting file reading of all FCS versions 3.0 / 3.1 / 3.2. The software provides a visual interface, allowing users to drag and drop to perform gate operations on FSC-SSC scatter plots and save the gate parameters as templates, which can then be applied to the same batch of samples for analysis with a single click, eliminating the need for repeated manual adjustments.

[0045] Preferably, the distribution curve of the histogram is smoothed and denoised. Various smoothing algorithms can be used, such as MATLAB's `smooth` function (using the locally weighted regression scatter smoothing method `LOWESS`), or other smoothing methods such as moving average (`movmean` function), low-pass filtering (`lowpass` function), or Butterworth filtering (`butter` function). Those skilled in the art will understand that the above smoothing algorithms are merely examples, and any known algorithm capable of smoothing and denoising histogram curves can be used in this application.

[0046] Figure 2 An example of a smoothed signal channel histogram is shown, with the smoothing curve, trough locations, and segmentation thresholds marked. This figure uses a flow cytometer with a specific channel signal as an example; however, as those skilled in the art will understand, this is merely a specific example of this application and does not limit the scope of its use. Other types of target channel signals from other bioanalysts can also be used to implement the scheme of this application.

[0047] As shown in the figure, this figure is a single-parameter histogram of the APC-H (allophycocyanin channel height) of the microsphere immunoassay sample collected by flow cytometry, which is used to show the fluorescence intensity distribution characteristics of the microspheres.

[0048] Where the horizontal axis represents the fluorescence intensity of the APC-H channel (unit: 10). x The vertical axis represents the amount of APC-labeled target analyte bound to the surface of a single microsphere; the vertical axis represents counts (number of events), which indicates the number of microsphere particles with corresponding fluorescence intensities.

[0049] The figure shows a typical bimodal distribution: the low-intensity peak on the left is the negative microsphere peak, and the peak value is approximately at 10. 3 The area represents microspheres that have not bound the target analyte; the high-intensity peak on the right represents positive microspheres, with a peak value approximately at 10. 4.4 The area represents a microsphere cluster that has been bound to the target analyte; the thick solid line is a bimodal fitting curve, clearly outlining the distribution contours of the two microsphere clusters. The vertical dashed line represents the positive / negative threshold line, set in the valley region between the two peaks (approximately 10). 3.5 (The location) is used to distinguish between negative and positive microspheres, and to count the proportion of positive microspheres F.

[0050] The figure shows that the negative and positive microspheres in this detection system have good separation and excellent signal-to-noise ratio, providing a reliable basis for subsequent accurate quantitative analysis.

[0051] Preferably, step S1 is continued to be executed to obtain the target channel signal measured by the flow cytometer, and the histogram of the target channel signal is used to determine whether the sample to be analyzed is a single-peak sample or a non-single-peak sample.

[0052] It should be noted that the samples to be analyzed in this application include standards and test samples, wherein the concentration of the standards is known. There are two optional strategies for the analysis process: one is to first perform a complete analysis of all standards and determine each threshold parameter, and then directly calculate all subsequent test sample data based on the standard parameters; the other strategy is to include both standards and test samples in the same batch of analysis each time. The advantage of the first strategy is that the analysis process is simple, results can be obtained quickly when analyzing test samples, improving analysis efficiency, and this modular distinction facilitates the docking of standards with test samples in the same batch. The advantage of the second strategy is that it avoids the threshold deviating from the actual signal range of the test sample due to using only standards, increases the sample size, and improves the robustness of Is estimation. This embodiment adopts the first optional strategy to increase analysis efficiency. However, this does not mean that the other strategy cannot be used in this application; it is also an optional embodiment.

[0053] Figure 2The smoothed distribution curve is shown in the figure (represented by a thick solid line). Based on the smoothed curve, the system uses a peak-finding function (e.g., MATLAB's findpeaks function) to find all valleys on the smoothed curve, recording the x-coordinate (corresponding signal strength), y-coordinate (corresponding microsphere count), and valley depth (the relative drop of the valley to its two adjacent peaks, which can be defined as the difference between the smaller value of the heights of the two adjacent peaks and the valley height, or the difference between the average height of the peaks on both sides of the valley and the valley height), and selects the deepest valley from them.

[0054] Then, the system determines whether the sample is a "single-peak sample" or a "non-single-peak sample" based on a multi-dimensional threshold judgment rule. Among them, single-peak samples include samples with an extremely low proportion of positive microspheres (with negative peaks predominating) or an extremely high proportion of positive microspheres (with positive peaks predominating), resulting in no obvious bimodal separation in the histogram.

[0055] The specific determination rule is as follows: if the deepest trough meets any of the following conditions, the histogram of the sample is considered to lack clear bimodal separation characteristics and is determined to be a single-peak sample; otherwise, it is determined to be a non-single-peak sample.

[0056] The conditions include: 1. The horizontal coordinate of the trough (i.e., signal strength) is less than the first threshold; or 2. The horizontal coordinate of the trough (i.e., signal strength) is greater than the second threshold; or 3. The vertical coordinate of the trough (i.e., the microsphere count corresponding to the signal strength) is less than the third preset threshold and the trough depth is less than the fourth threshold.

[0057] The first, second, third, and fourth thresholds are predetermined through statistical analysis of trough parameters from a large number of historical samples. For example, in one specific embodiment, the first threshold is 2.5, the second threshold is 3.5, the third threshold is 1.4, and the fourth threshold is 2.2. It should be noted that the above threshold values ​​(2.5, 3.5, 1.4, 2.2) are determined statistically based on specific experimental systems and reagent batches. In practical applications, they can be adaptively adjusted according to different detection systems and instruments; for example, the depth threshold can be selected within the range of 1.9 to 2.2.

[0058] If the deepest trough meets any of the above conditions, it is determined to be a "single-peak sample", that is, the histogram does not show a clear bimodal shape.

[0059] In step S2, for the identified non-single-peak samples, the signal intensity corresponding to the deepest valley of its signal channel histogram is used as the first division threshold to calculate the corresponding positive microsphere ratio F.

[0060] For the data that was identified as "non-single-peak sample" in the previous step, the system directly uses the horizontal coordinate position of the deepest valley as the first division threshold of the sample, then counts the number of microspheres with signal intensity greater than the threshold, calculates the proportion of the microspheres to the total number of effective microspheres in the sample, and obtains the intermediate parameter F (proportion of positive microspheres). The calculation formula is: F = number of microspheres with signal intensity greater than the threshold / total number of effective microspheres.

[0061] In this paper, "total effective microsphere count" refers to the total number of effective microspheres that are morphologically intact, uniform in size, and usable for subsequent signal analysis after screening using forward scattered (FSC) and side scattered (SSC) scatter plots. Specifically, the raw data acquired by flow cytometry includes all detected particle events, which may include target magnetic microspheres, debris, microsphere aggregates, and non-specific particles. To ensure the accuracy of the analysis results, this method first selects the target microsphere population on the FSC-SSC scatter plot using rectangular or polygonal gates, excluding debris (usually due to low FSC signals) and aggregates (usually due to high FSC and SSC signals). The microspheres retained after screening are the "effective microspheres," and their total number is the "total effective microsphere count."

[0062] In step S3, for samples determined to be single-peak samples, the corresponding proportion of positive microspheres is calculated based on a second division threshold, wherein the second division threshold is determined based on the first division threshold of multiple non-single-peak samples of the standard.

[0063] In one embodiment, samples with a positive microsphere percentage F falling within a preset first interval are selected from all non-single-peak samples. The histograms of these samples exhibit clear bimodal separation characteristics, with their deepest troughs being the most reliable. A first dividing threshold (i.e., the signal intensity corresponding to its deepest trough) for each such non-single-peak sample is added to a threshold template array.

[0064] According to a preferred embodiment of this application, the preset first interval is preferably 0.1 to 0.3. When the F value is between 0.1 and 0.3, it is considered that the peaks of the sample are clearly distinguishable, and the location of its deepest trough is most reliable.

[0065] Next, the system integrates the values ​​in the threshold template array: first, outlier removal is performed on the partitioning thresholds in the threshold template array. Methods for removing outliers include, but are not limited to, the interquartile range (IQR method, typically using 1.5 times the IQR as the threshold), Grubbs' test, the 3σ principle (Laida criterion), and the modified Z-score method. As those skilled in the art will understand, the above methods for removing outliers are merely examples, and this application is not limited thereto; any statistical method capable of identifying and removing outliers can be selected based on the actual data characteristics. The interquartile range (IQR method) is preferred for outlier removal because it does not rely on the assumption of a normal distribution of the data and has better robustness for the small sample size (typically less than 20) of the threshold template array in this application. After removing outliers, the mean or median of the remaining thresholds is calculated to obtain the second partitioning threshold.

[0066] If the number of non-single-peak samples obtained after screening is zero or the threshold template array is empty, the user will be prompted that there are no reference bimodal samples in the current batch, and the experiment needs to be re-checked or the threshold needs to be manually set. Alternatively, other solutions can be adopted for this situation. For example, a value can be set in advance as the threshold, which is based on empirical data obtained from statistics of data from the same test source, to avoid the inability of users to manually set an appropriate threshold.

[0067] For the data previously labeled as "single-peak samples," these samples could not find an effective threshold on their own. Therefore, their unstable troughs were no longer used. Instead, a second dividing threshold was uniformly applied, and the proportion of positive microspheres (F) of these samples was calculated using the same method (the proportion of microspheres with statistical signal strength greater than the second dividing threshold). Thus, each sample (including standards and test samples) obtained a stable and consistent F value.

[0068] In step S4, based on the range of positive microsphere percentage in the sample to be analyzed and / or the signal intensity information of positive microspheres in the sample to be analyzed, the average target capture amount of a single positive microsphere in the sample to be analyzed is calculated. This step specifically includes the following sub-steps.

[0069] Sub-step 4.1: Screen samples where the percentage F of positive microspheres falls within the second interval, and determine the mean signal intensity Ib of all positive microspheres in each screened sample and the number of positive microspheres n. According to a preferred embodiment of this application, the second interval is preset to be less than 0.2. Under extremely low concentration conditions of F < 0.2, most positive microspheres capture only one target molecule, therefore these samples are an ideal source for calculating the signal intensity of a single target molecule.

[0070] Specifically, the system selects samples whose F values ​​fall within this range from all samples to be analyzed (including standards and test samples). For each sample i that satisfies F < 0.2, the system calculates the mean fluorescence intensity Ib of all positive microspheres in the sample (i.e., the signal intensity is greater than the corresponding division threshold for the sample, or greater than the second division threshold for single-peak samples), and counts the number of positive microspheres n.

[0071] Ib represents the mean signal intensity of all positive microspheres in a sample. The specific calculation method is as follows: For each sample i satisfying F < 0.2, firstly, based on the corresponding dividing threshold (for non-single-peak samples, it is the x-coordinate of its deepest trough; for single-peak samples, it is the second dividing threshold), all microspheres with signal intensities greater than this threshold are selected and considered as positive microspheres. Then, the signal intensity values ​​of these positive microspheres in the target fluorescence channel (such as the APC channel) are extracted, and the arithmetic mean of these signal intensities is calculated, which is Ib. Simultaneously, the total number of these positive microspheres is counted, denoted as n. Mathematically, if the set of signal intensities of positive microspheres in this sample is {I1, I2, …, I…}, then… n}, then Ib = (I1 + I2 + …+ I n ) / n.

[0072] Sub-step 4.2: Using the number of positive microspheres n as the weight, perform a weighted average of the single-molecule signal intensity Isi of the positive microspheres that captured a single target in the screened sample to obtain the mean signal intensity Is of the positive microspheres that captured a single target in the sample to be analyzed, wherein Isi is based on the mean signal intensity Ib and the proportion F of the positive microspheres.

[0073] Specifically, for each sample i that satisfies F < 0.2, first follow the formula: Isi = F Ib / (-ln(1-F)) The single-molecule signal intensity of the positive microspheres that captured a single target analyte for this sample is calculated. Then, the system performs a weighted average of all Isi, using the number of positive microspheres n as the weight, to calculate the final Is: Is = Σ(Isi × n) / Σ(n).

[0074] The Is value represents the average signal intensity of a single target microsphere captured under the current batch of experiments, reagent batch number, and instrument conditions. It will be used for quantitative calculations of high-concentration samples.

[0075] Sub-step 4.3: Based on the range of the positive microsphere percentage F in the sample to be analyzed, and combining the mean signal intensity Is of the positive microspheres that captured a single target analyte and the mean signal intensity I of the positive microspheres in the sample to be analyzed (hereinafter referred to as "mean signal intensity I of positive microspheres"), calculate the average target analyte capture amount A of a single positive microsphere in the sample to be analyzed. Here, I is the mean signal intensity of all positive microspheres in the sample itself, which has the same meaning as Ib used to calculate Is in sub-step 4.2, but the source sample is different.

[0076] Specifically, for each sample (each standard and each test sample), the system determines the interval in which its F-value falls. In a preferred embodiment, the interval is divided by 0.7.

[0077] If F < 0.7, the system determines that the sample is in a low concentration range, and the distribution of the target substance on the microspheres conforms to the Poisson distribution assumption. Therefore, the Poisson distribution transformation formula is used: A = -ln(1-F).

[0078] The derivation of this formula is based on the following: According to the Poisson distribution, the probability that a single microsphere fails to capture any target object is e. -A Therefore, the proportion of positive microspheres is F = 1-e -A The above formula is obtained through mathematical transformation.

[0079] When F ≥ 0.7, the system determines that the sample is in a high concentration range. At this point, most microspheres have captured at least one target analyte, and the Poisson distribution assumption no longer applies. Therefore, a simulated signal conversion formula based on fluorescence intensity is used: A = (F × I) / Is, Where I is the mean fluorescence intensity of all positive microspheres (signal intensity greater than the corresponding segmentation threshold) in the current sample to be analyzed.

[0080] The segmented processing described above utilizes the advantages of both digital counting and analog fluorescence intensity. High precision is achieved in the low-concentration region through Poisson statistics, while the fluorescence intensity ratio is used to break through the upper limit of counting in the high-concentration region, thereby extending the detection dynamic range to more than 5 orders of magnitude.

[0081] After completing the above steps S1 to S4, the scheme of this application can also perform concentration quantification.

[0082] Specifically, the system first combines standards from the same batch with known concentrations and their corresponding average target analyte capture amount A to obtain a standard curve. In such cases... Figure 3The standard curve analysis interface pop-up window shows the actual concentration value corresponding to each standard. Since standard tests are usually set up with multiple replicates (e.g., 3 replicates), the system will automatically merge data of the same concentration group (e.g., using MATLAB's unique function) and calculate the mean and standard deviation of the A value for each concentration group. The mean is used for subsequent fitting, and the standard deviation is used for the error bars in the standard curve plot.

[0083] Preferably, a standard curve is fitted using a specific fitting algorithm (e.g., four-parameter logistic fitting (4-PL) or weighted four-parameter fitting, specifically by calling MATLAB's `fit` function and selecting the logistic4 model), with the known concentration as the x-axis and the A value as the y-axis. This yields a formula for converting the A value to concentration. This formula is applicable to the analysis of this batch of samples. The fitted curve, formula, and goodness of fit (e.g., R²) will be displayed in the standard curve analysis interface.

[0084] After the analysis is completed, all parameters (including the gating template, the first dividing threshold, the second dividing threshold, the mean signal intensity Is of the positive microspheres that capture a single target, the standard curve formula, the average target capture amount A, the concentration, etc.) can be exported as a working file for quick access and viewing in subsequent analyses without the need to repeat the analysis of the standard.

[0085] Then, based on the standard curve, the system converts the average target analyte capture amount A of the test sample into the corresponding sample concentration. The specific conversion method is as follows: substitute the A value of each test sample into the standard curve formula (such as the function obtained by fitting a four-parameter logistic curve) to calculate the corresponding concentration value.

[0086] Specific Analysis Examples

[0087] This embodiment uses a set of protein biomarker hypersensitive detection scenarios as an example. The instruments, reagents, and core judgment parameters used are as follows: Instrument: Flow cytometer supporting FCS3.2 format data output; Reagents: Encoded magnetic microspheres, signal-labeled antibodies, and APC fluorescence signal amplification reagent.

[0088] Valley determination parameters: Valley horizontal axis threshold range 2.5~3.5, vertical axis threshold <1.4, depth threshold <2.2; F-value intervals are defined as follows: the first interval used to generate the general threshold (i.e., the second threshold) is 0.1 to 0.3; the screening interval for calculating Is is F < 0.2; and the threshold for distinguishing between high and low concentration intervals is 0.7.

[0089] Specifically, the peak-splitting algorithm thresholds (first to fourth preset thresholds) of this application can be predetermined by the following statistical methods.

[0090] A large amount of standard sample data was collected (65 data files were used in this example, for example). A trough detection algorithm was used to extract the x-coordinate, y-coordinate, and depth of the deepest trough for each sample. For the trough x-coordinate, outliers were removed using the quartile method, and the mean and standard deviation were calculated. A threshold range for the x-coordinate (e.g., 2.5–3.5) was determined based on the mean ± 3 times the standard deviation, appropriately expanded by about 20%. For the trough y-coordinate and depth, the samples were manually labeled into unimodal and bimodal groups. After removing outliers using the quartile method, the mean of the highest value in the unimodal group and the lowest value in the bimodal group was used as the dividing threshold, resulting in a trough y-coordinate threshold (e.g., 1.4) and a trough depth threshold (e.g., 1.9, which can be adjusted to 2.2 in practical applications depending on the experimental system). Using these thresholds, another group of 50 data files was automatically classified as unimodal / bimodal. The results were completely consistent with the manual classification, with no misclassifications.

[0091] The determination of the "first interval" (F-value range) used to generate the universal threshold (i.e., the second dividing threshold) is based on the following: For multiple sets of standard sample data, bimodal samples were screened with different upper limits such as F<0.7, F<0.3, and F<0.15, and the coefficient of variation (CV) of the universal threshold under each condition was calculated. The results show that the CV decreases as the upper limit of F decreases (threshold stability improves), but in order to ensure that a sufficient number of samples participate in the statistics to avoid extreme errors, 0.1 to 0.3 was finally selected as the first interval.

[0092] The sample configuration used in this embodiment is as follows: 8 concentration gradients of the standard were set, with a concentration range of 0.026~10 pg / mL, and 2 parallel replicates were set for each concentration, for a total of 16 samples; 7 samples were used for testing.

[0093] According to the method of this application, the standard data is first analyzed: the standard file is imported, the FSC-SSC scatter plot is manually selected and gated, and saved as a template. The software automatically performs trough detection and single-peak / non-single-peak determination. For non-single-peak samples, those with F values ​​between 0.1 and 0.3 are collected, outliers are removed, and the mean is taken to obtain a general threshold C (approximately 2.94). Then, C is applied to all single-peak standard samples to calculate the F and A values ​​of each standard. After inputting the known concentrations of each standard, the software automatically fits a standard curve. Subsequently, samples to be tested are imported in batches, and the software automatically applies the same gate template and general threshold C to quickly calculate the F, A, and concentration of each sample.

[0094] Some of the analysis results are shown in Table 1 below. In the "bimodal" column, "0" represents data judged as unimodal and "1" represents data judged as non-unimodal.

[0095] Table 1. Detailed analysis results of protein biomarker detection data for a certain batch, including standards.

[0096] Upon review, the automated analysis results of all samples were consistent with the manual analysis results, with a quantitative error of ≤5% and a detection dynamic range covering 5 orders of magnitude. It can stably adapt to detection scenarios of extremely low to extremely high concentrations of protein biomarkers.

[0097] After obtaining the concentration, the system will perform the following follow-up processing: First, in the sample analysis interface (such as... Figure 4 The table shown presents intermediate data for each sample, including the F-value, A-value, calculated concentration, and mean fluorescence intensity (I) of the positive microspheres, for user verification. For large-scale clinical testing, this intermediate data presentation aids in quality control and outlier tracing. Secondly, users can export all analysis results (including concentrations and intermediate parameters) into a software-recognizable working file format for direct subsequent access and analysis. Furthermore, users can export the data from the table as a .txt or .xls / .xlsx file, facilitating subsequent statistical analysis, report generation, or integration with other database systems.

[0098] Users first need to load the previously saved working file, which contains all analytical parameters, including the gating template, first and second thresholds, the mean signal intensity (Is) of positive microspheres capturing a single target analyte, and the standard curve formula. After loading, users can batch import the FCS files of the samples to be tested. The software automatically performs the aforementioned steps S1 to S4 for each sample, calculating the F-value and A-value for each sample, and converting the A-value into the analyte concentration using the standard curve formula. Simultaneously, the software will display intermediate data such as the F-value, A-value, concentration, and mean fluorescence intensity (I) of the positive microspheres for each sample in a table on the interface. To facilitate intuitive verification of the threshold division's rationality, clicking on any sample will display the signal channel histogram for that sample (e.g., ...). Figure 2 As shown in the figure, the smoothed distribution curve, the location of the deepest trough (i.e., the first dividing threshold for non-unimodal samples) or the second dividing threshold (for unimodal samples), and the specific values ​​of the dividing thresholds are marked. Users can quickly determine the reliability of the analysis results by observing the degree of separation between the positive and negative peaks in the histogram and the matching of the threshold positions. For large-sample clinical tests, the presentation of intermediate data and graphical verification functions help with quality control and outlier tracing. After all data analysis is completed, users can export the results to a working file format that the software can recognize, or export the data in the table as a .txt or .xls / .xlsx format as a universally editable text file, which is convenient for subsequent statistics, report generation, or integration with other database systems.

[0099] Figure 5The detailed algorithm flowchart of the data analysis process in the embodiments of this application is shown, which illustrates the core execution logic of this application, including histogram preprocessing, optimal valley detection, intelligent signal type judgment and threshold template batch processing, which are the core of this application to achieve automatic threshold division and batch standardized analysis.

[0100] The process begins with a preprocessing step common to all samples to be analyzed. First, the system receives valid microsphere data points filtered by a data gating procedure and generates a single-parameter fluorescence intensity histogram with fluorescence intensity as the x-axis and the number of microsphere events as the y-axis. Then, a sliding window algorithm averaging adjacent values ​​is used to smooth and denoise the original histogram to eliminate interference from flow cytometry electronic noise, individual microsphere differences, and signal acquisition fluctuations, avoiding misclassification of random noise as threshold candidate troughs. After smoothing, the entire smoothed curve is traversed by subtracting adjacent values ​​to identify and record all local minima, i.e., all potential positive / negative threshold candidate troughs. Next, the depth of each candidate trough (the height difference between the trough point and its two adjacent peak points) is calculated, and the trough with the largest depth is selected as the optimal threshold candidate point. The depth, x-axis (fluorescence intensity value), and y-axis (number of events) parameters of this trough are recorded.

[0101] After obtaining the optimal trough parameters, the method enters the signal type determination and branching stage. As mentioned above, the system determines whether a signal is a single-peak signal or a non-single-peak signal based on the comparison of the trough's horizontal and vertical coordinates, as well as the trough depth, with the corresponding threshold. If it is determined to be a single-peak signal (corresponding to extremely low concentration all-negative samples or extremely high concentration all-positive samples, where traditional manual thresholding methods are prone to serious bias), the system directly enters the data correction step, using a preset single-peak quantitative correction algorithm for calibration calculation. If it is determined to be a double-peak signal (corresponding to samples of normal concentration), the system automatically determines the current positive / negative threshold based on the horizontal coordinate of the currently selected deepest trough and further determines the standardized threshold template.

[0102] In the template judgment branch for bimodal data, if there is no batch processing template (usually when analyzing standards for the first time), a standardized threshold template is generated based on the existing non-unimodal data and the set threshold and stored in the system template library for use by all subsequent samples to be analyzed; if a batch processing template already exists, the system directly applies the established standardized threshold template to the current data to be corrected, without having to recalculate the threshold for each sample.

[0103] Regardless of whether a single-peak correction algorithm or a batch template threshold is used, all data ultimately undergo a unified data correction process. The final output is a standardized quantitative signal value, which is used for subsequent standard curve plotting and analyte concentration conversion.

[0104] The key analysis method described in this embodiment completely replaces the traditional manual thresholding operation for each sample by using automated trough detection and signal type judgment. At the same time, by generating and applying standardized threshold templates, it enables batch analysis of large clinical samples, improving the analysis efficiency by more than 10 times compared with traditional methods. The quantitative results of all samples to be analyzed have an error of less than 5%, and it has excellent adaptability to extreme samples with extremely low and extremely high concentrations.

[0105] The above is a detailed description of the data analysis method 10 described in this application. Accordingly, this application also provides a data analysis system based on flow cytometry.

[0106] Figure 6 This is a flowchart illustrating the overall architecture of the flow cytometer-based data analysis system described in this application embodiment. It shows the entire execution process from importing raw flow cytometer data to outputting quantitative results of analyte concentration, supporting integrated automated analysis of standards and test samples from the same batch.

[0107] The workflow begins with the import of raw flow cytometry data: the system uses an FCS file reading and transfer program to convert the raw FCS format data output by the instrument into a program-readable format and store it. This supports batch import and unified management of multiple sets of sample data from the database, providing a standardized data foundation for the collaborative analysis of standards and test samples from the same batch. The next step is the data gating program, which has a built-in user interface, allowing operators to flexibly adjust gating parameters according to the actual sample conditions. Gating parameters can be saved as templates and applied to other samples in the same batch without repeated manual adjustments. The program uses FSC-SSC scatter plot gating to separate uniformly shaped effective microsphere data points from the raw data, excluding non-target particles such as debris and microsphere aggregates, providing pure and effective data for subsequent analysis.

[0108] The selected effective microsphere data points are fed into the key analysis module shown in the dashed box. This module is the core functional unit of the system, enabling fully automated conversion from effective microsphere data to quantitative signal values. The quantitative signal values ​​output by this module are then fed into the standard curve plotting unit. Combined with the known concentrations of standards from the same batch, a four-parameter logistic fitting is used to construct a standard curve, establishing a quantitative correspondence between the signal values ​​and the concentration of the analyte, thus achieving automatic conversion of the sample concentration. Finally, the system outputs standardized quantitative results, completing the entire analysis process.

[0109] This architecture integrates batch data preprocessing, core algorithm analysis, standard curve construction, and automatic result output. It eliminates the need for manual gate setting, threshold setting, and standard curve fitting for each sample, achieving fully automated processing from raw data to quantitative results. This significantly improves the efficiency and stability of large clinical sample analysis and is suitable for routine clinical hypersensitive testing applications.

[0110] Figure 7 The module structure of the data analysis system 20 according to this application and the logical relationships between the modules are illustrated. The entire system operates in a fully automated manner, from FCS file input to final output, sequentially or on demand calling the following functional modules: First, the sample classification module 21 is responsible for acquiring the target channel signal measured by the flow cytometer, and determining whether the sample to be analyzed is a single-peaked sample or a non-single-peaked sample based on the histogram of the target channel signal.

[0111] Specifically, the sample classification module 21 extracts the deepest valley in the histogram and obtains the valley parameters of the deepest valley. The valley parameters include the signal intensity corresponding to the horizontal axis of the valley, the microsphere count corresponding to the vertical axis of the valley, and the valley depth. Based on the comparison of the valley parameters with a preset threshold range, it determines whether the sample is a single-peak sample or a non-single-peak sample.

[0112] For samples identified as non-single-peak samples, the non-single-peak sample processing module 22 uses the signal intensity corresponding to the deepest valley of its histogram as the first division threshold to calculate the corresponding positive microsphere proportion F.

[0113] For single-peak samples, the single-peak sample processing module 23 calculates the corresponding percentage of positive microspheres based on a second division threshold, wherein the second division threshold is determined based on the first division threshold of multiple non-single-peak samples of the standard.

[0114] In other words, for data that were originally labeled as single-peak samples, the single-peak sample processing module 23 no longer relies on its own unstable troughs, but directly uses the aforementioned second division threshold as the division standard, and calculates the proportion F of positive microspheres for each single-peak sample accordingly. Thus, all samples to be analyzed have obtained a uniform and comparable F value.

[0115] The quantitative module 24 calculates the average target capture amount of a single positive microsphere in the sample based on the range of the proportion of positive microspheres in the sample to be analyzed, combined with the signal intensity information of the positive microspheres in the sample to be analyzed. The specific process is as described in the specific steps of the method above, and will not be repeated here.

[0116] Optionally, the data analysis system 20 may also include a preprocessing module 25 and a concentration conversion module 26.

[0117] The preprocessing module 25 is responsible for reading the FCS file exported from the flow cytometer, using forward scatter and side scatter scatter plots to screen out effective microsphere populations with complete morphology, and plotting the target fluorescence channel signals of the population into a histogram. At the same time, the histogram curves are smoothed and denoised to provide a clean data foundation for subsequent analysis.

[0118] The concentration conversion module 26 uses the known concentrations of the same batch of standards and their corresponding A values ​​to fit a standard curve (e.g., using a four-parameter logistic fitting). Then, the A value of the sample to be tested is substituted into this curve to calculate the final concentration of the analyte. Users can export all intermediate parameters and concentration results through the interface for quality control and subsequent analysis.

[0119] Figure 7 The modules shown work together to fully automate the analysis process from raw data to quantitative concentration. Figure 1 Flowcharts and Figure 6 The data processing flowcharts form a one-to-one correspondence.

[0120] This example uses nearly 400 clinical samples for data analysis practice. The analysis efficiency is significantly higher than manual analysis, effectively improving it by nearly 6 times. For some samples with low, medium, and high signal values, the coefficients of variation (CV) of the F values ​​obtained by different people analyzing the same sample reached 14%, 4%, and 2%, respectively. The automatic algorithm can stably reproduce the results multiple times, and the accuracy has been verified by professional analysts. The results obtained are consistent with the manual mean, ensuring accuracy and stability. For the detection biomarkers and detection methodologies in this example, the detection dynamic range covers 0.013 pg / mL to 3319 pg / mL (approximately 5.4 orders of magnitude), fully meeting the accuracy and efficiency requirements of clinical hypersensitive detection.

[0121] The table below shows a comparison of effects according to specific embodiments.

[0122]

[0123] Table 2 Comparison of Manual and Automatic Analysis

[0124] Furthermore, this application also provides a computer-readable storage medium storing a computer program. When the computer program is executed by one or more processors, it implements the data analysis method described in any of the preceding claims. The computer-readable storage medium may include, but is not limited to: magnetic storage devices (such as hard disks, floppy disks, magnetic stripes), optical storage devices (such as compact CDs, digital multi-disc DVDs), smart cards, flash memory devices (such as memory cards, memory sticks, key drives), random access memory (RAM), read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), registers, removable disks, and any other suitable medium for storing software and / or instructions accessible and readable by a computer. By running the program in this medium, users can conveniently perform batch, standardized, and automated analysis of digital ultrasensitive detection data on a conventional flow cytometer platform, significantly improving analysis efficiency and data consistency.

[0125] The above, combined with Figures 1 to 7 The preferred embodiments of the method, system, and storage medium of this application are described in detail, along with the corresponding operating procedures. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. Those skilled in the art can make various modifications and equivalent substitutions without departing from the spirit and principles of this application, and such modifications and equivalent substitutions should also be considered within the scope of protection of this application.

[0126] All aspects, elements, or any part of the elements, or any combination of the elements of this application, may be implemented using a "processing system" comprising one or more processors. Examples of processors include: microprocessors, microcontrollers, digital signal processors (DSPs), field-programmable gate arrays (FPGAs), programmable logic devices (PLDs), state machines, gated logic, discrete hardware circuitry, and other suitable hardware configured to perform the various functionalities described throughout this application. One or more processors in the processing system may execute software. Software should be broadly interpreted as instructions, instruction sets, code, code segments, program code, programs, subroutines, software modules, applications, software applications, software packages, routines, subroutines, objects, executables, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description languages, or other terms. Software may reside on a computer-readable medium. The computer-readable medium may be a non-transient computer-readable medium. As examples, non-transient computer-readable media include: magnetic storage devices (e.g., hard disks, floppy disks, magnetic stripes), optical disks (e.g., compact discs (CDs), digital multi-purpose discs (DVDs)), smart cards, flash memory devices (e.g., memory cards, memory sticks, key drives), random access memory (RAM), read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), registers, removable disks, and any other suitable medium for storing software and / or instructions accessible and readable by a computer. As examples, computer-readable media may also include carrier waves, transmission lines, and any other suitable medium for transmitting software and / or instructions accessible and readable by a computer. Computer-readable media may reside in a processing system, be external to a processing system, or be distributed across multiple entities including the processing system. Computer-readable media may be implemented in a computer program product. As an example, a computer program product may include computer-readable media within packaging material. Those skilled in the art will recognize how the functionality described throughout this application is best implemented depending on the specific application and the overall design constraints imposed on the system as a whole.

[0127] It should be understood that the specific order or hierarchy of the steps in the disclosed methods is an exemplary illustration of the process. Based on design preferences, it should be understood that the specific order or hierarchy of the steps in the methods or methodologies of the embodiments of this application may be rearranged. The appended method claims present the elements of various steps in a sample order and are not intended to be limited to the presented specific order or hierarchy, unless specifically stated herein.

[0128] The preceding description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the universal principles defined herein may be applied to other aspects. Therefore, the claims are not intended to be limited to the aspects shown herein, but should be granted the full scope consistent with the language of the claims, wherein references to the singular form of an element are not intended to mean “one and only one” (unless specifically stated otherwise) but “one or more”. Unless specifically stated otherwise, the term “some” refers to one or more. The phrase “at least one” referring to a list of items means any combination of these items, including a single member. As an example, “at least one of a, b, or c” is intended to cover: at least one a; at least one b; at least one c; at least one a and at least one b; at least one a and at least one c; at least one b and at least one c; and at least one a, at least one b, and at least one c. All structural and functional equivalents of the various aspects described throughout this application that are currently or hereafter known to a person skilled in the art are expressly incorporated herein by reference and are intended to be covered by the claims. Furthermore, nothing disclosed herein is intended to be made public, whether or not such disclosure is explicitly stated in the claims.

Claims

1. A data analysis method based on a bioanalyst, wherein the sample to be analyzed includes a test sample and a standard, characterized in that, The method includes the following steps: S1: Obtain the target channel signal measured by the bioanalyzer, and determine whether the sample to be analyzed is a single-peaked sample or a non-single-peaked sample based on the histogram of the target channel signal. S2: For the non-single-peak sample, the signal intensity corresponding to the deepest valley of its histogram is used as the first division threshold to calculate the corresponding positive microsphere ratio. S3: For the single-peak sample, calculate the corresponding percentage of positive microspheres based on a second partitioning threshold, wherein the second partitioning threshold is determined based on a first partitioning threshold of multiple non-single-peak samples of the standard; and S4: Based on the range of positive microsphere percentage in the sample to be analyzed and / or the signal intensity information of positive microspheres in the sample to be analyzed, calculate the average target capture amount of a single positive microsphere in the sample to be analyzed.

2. The data analysis method based on a biological analyzer as described in claim 1, characterized in that, Determining whether a sample to be analyzed is a unimodal or non-unimodal sample based on the histogram includes: Extract the deepest valley from the histogram and obtain the valley parameters of the deepest valley, which include the valley horizontal coordinate, the valley vertical coordinate, and the valley depth; Based on the conditions related to the trough parameter, it is determined whether the sample to be analyzed is a single-peak sample or a non-single-peak sample.

3. The data analysis method based on a biological analyzer as described in claim 2, characterized in that, The sample to be analyzed is determined to be a unimodal sample if at least one of the following conditions is met; otherwise, the sample to be analyzed is determined to be a non-unimodal sample: The horizontal coordinate of the trough is less than the first threshold; or The horizontal coordinate of the trough is greater than the second threshold; or The trough's ordinate is less than the third threshold and the trough's depth is less than the fourth threshold.

4. The data analysis method as described in claim 1, characterized in that, The second division threshold is calculated by taking the mean or median after removing outliers from the first division threshold of non-single-peak samples where the proportion of positive microspheres is in the first interval.

5. The data analysis method as described in claim 4, characterized in that, The first interval is 0.1 to 0.

3.

6. The data analysis method as described in claim 1, characterized in that, Based on the range of positive microsphere percentage in the sample to be analyzed and / or the signal intensity information of positive microspheres in the sample to be analyzed, the average target capture amount A of a single positive microsphere in the sample to be analyzed is calculated, including: Samples with a positive microsphere percentage F in the second interval were screened, and the mean signal intensity Ib and the number of positive microspheres n in each screened sample were determined. Using the number of positive microspheres n as the weight, the single-molecule signal intensity Isi of the positive microspheres that captured a single target in the screened sample is weighted and averaged to obtain the mean signal intensity Is of the positive microspheres that captured a single target in the sample to be analyzed, wherein Isi is based on the mean signal intensity Ib and the proportion of positive microspheres F. Based on the range of positive microsphere percentage F in the sample to be analyzed, and combined with the mean signal intensity Is of the positive microsphere that captures a single target and the mean signal intensity I of the positive microsphere in the sample to be analyzed, the average target capture amount A of a single positive microsphere in the sample to be analyzed is calculated.

7. The data analysis method as described in claim 6, characterized in that, Based on the range of the proportion F of positive microspheres in the sample to be analyzed, the average target capture amount A of a single positive microsphere in the sample to be analyzed is calculated, including: When the proportion F of the positive microspheres is less than the first preset percentage threshold, according to Calculate A = -ln(1-F); and When the proportion F of the positive microspheres is greater than or equal to the first preset percentage threshold, according to Calculate A = (F × I) / Is.

8. The data analysis method as described in claim 1, characterized in that, The method further includes: Obtain the average target analyte capture amount corresponding to the standard at known concentrations, and fit a standard curve to obtain the target analyte capture amount; and Based on the standard curve, the average target capture amount of the test sample is converted into the concentration of the test sample.

9. A data analysis system based on a biological analyzer, wherein the sample to be analyzed includes a test sample and a standard, characterized in that, The data analysis system includes: The sample classification module is used to acquire the target channel signal measured by the bioanalyst and determine whether the sample to be analyzed is a single-peaked sample or a non-single-peaked sample based on the histogram of the target channel signal. The non-single-peak sample processing module is used to calculate the proportion of positive microspheres for non-single-peak samples by using the signal intensity corresponding to the deepest valley of its histogram as the first division threshold. A single-peak sample processing module is used to calculate the proportion of positive microspheres corresponding to a single-peak sample based on a second division threshold, wherein the second division threshold is determined based on a first division threshold of multiple non-single-peak samples in the standard; and The quantitative module is used to calculate the average target capture amount of a single positive microsphere in the sample based on the range of the proportion of positive microspheres in the sample to be analyzed, and in combination with the signal intensity information of the positive microspheres in the sample to be analyzed.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the data analysis method as described in any one of claims 1 to 8.