Marker combination for evaluating erythrocyte pollution degree and application thereof
Quantitative detection of the degree of red blood cell contamination by marker combination solves the problem that red blood cell contamination cannot be accurately quantified in the prior art, and improves the accuracy of plasma proteomic analysis and the reliability of disease marker screening.
Patent Information
- Application Number
- CN202510492870.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-04-18
AI Technical Summary
The prior art cannot accurately quantify the degree of red blood cell contamination, resulting in large errors in plasma proteomic analysis results, affecting the accuracy of disease marker screening and clinical verification effect.
The marker combinations include PRPS1, ADD2, PA2G4, OLA1, ACLY, PIP4K2A, FLOT1, CFL1, ADD1, SNCA, STOM, RAP1B, PPIA, MPP1, EPB42, RAN, CD59, HSPA8, SLC4A1, HBD, EPB41, HBA1, BLVRB, GAPDH, PGK1, SPTB, HBB, SPTA1, CA1 and ANK1. Quantitative detection was performed by mass spectrometry, chromatography and other methods, and the pollution index was calculated and mapped to the number of red blood cells.
Accurate identification and quantification of the degree of red blood cell contamination is achieved, the accuracy of plasma proteomics analysis is improved, the false positive rate is reduced, and the reliability of disease marker screening is ensured.
Smart Images

Figure BDA0005366073040000091 
Figure BDA0005366073040000101 
Figure HDA0005366073050000011
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biomedical technologies, and particularly to a marker combination for evaluating the degree of red blood cell contamination and its application. Background Art
[0002] Existing methods for evaluating red blood cell contamination of samples rely on subjective experience or a single marker (such as HBA1) to evaluate contamination, and cannot quantify the degree of contamination or correct data deviation, resulting in the disconnection between the analysis results and the true biological effects. Moreover, as a core technology for non-invasive disease marker discovery, circulating blood proteomics has significantly improved the depth of plasma protein detection in recent years driven by mass spectrometry (MS) and nanoparticle (NP) enrichment methods. However, the evaluation of the contamination problem of NP-based plasma proteomics has not been disclosed in the prior art.
[0003] Moreover, in terms of clinical translation, the biomarkers screened in the prior art for evaluating diseases often have specific proteins related to red blood cells. If the proteins screened and related to red blood cell contamination are used as disease biomarkers, it will lead to an increase in incorrect results, specifically manifested as an increase in false positive rate, a decrease in specificity, and failure of clinical verification. For example, the hemoglobin subunits released by red blood cell residues (such as HBA1) may interfere with the analysis of inflammation or tumor-related pathways, resulting in incorrect association of disease mechanisms. Therefore, there is an urgent need for a marker that can accurately evaluate the degree of red blood cell contamination in plasma proteomics samples. Summary of the Invention
[0004] In view of the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a marker combination for evaluating the degree of red blood cell contamination to solve the problems in the prior art.
[0005] To achieve the above object and other related objects, the present invention is obtained through the following technical solutions.
[0006] In the first aspect of the present invention, there is provided a marker combination for evaluating the degree of red blood cell contamination, the marker combination including PRPS1, ADD2, PA2G4, OLA1, ACLY, PIP4K2A, FLOT1, CFL1, ADD1, SNCA, STOM, RAP1B, PPIA, MPP1, EPB42, RAN, CD59, HSPA8, SLC4A1, HBD, EPB41, HBA1, BLVRB, GAPDH, PGK1, SPTB, HBB, SPTA1, CA1 and ANK1.
[0007] The second aspect of the present invention provides the use of a reagent for quantitatively detecting a biomarker combination in a sample in at least one of the following, the combination including PRPS1, ADD2, PA2G4, OLA1, ACLY, PIP4K2A, FLOT1, CFL1, ADD1, SNCA, STOM, RAP1B, PPIA, MPP1, EPB42, RAN, CD59, HSPA8, SLC4A1, HBD, EPB41, HBA1, BLVRB, GAPDH, PGK1, SPTB, HBB, SPTA1, CA1, and ANK1;
[0008] (1) Evaluating the degree of red blood cell contamination;
[0009] (2) Preparing a product for evaluating the degree of red blood cell contamination;
[0010] (3) Calculating the number of red blood cells;
[0011] (4) Preparing a product for calculating the number of red blood cells;
[0012] (5) Screening disease-related biomarkers;
[0013] (6) Preparing a product for screening disease biomarkers.
[0014] In some embodiments of the present invention, the reagent performs quantitative detection by at least one method of mass spectrometry, chromatography, surface-enhanced Raman spectroscopy, Western blotting, flow cytometry, protein array, immunoprecipitation, immunoadsorption.
[0015] In some embodiments of the present invention, the reagent is selected from one or more of substances specific to the biomarker, biomarker-specific probes, and protein chips.
[0016] In some embodiments of the present invention, the substance specific to the biomarker includes antibodies, ligand proteins, polypeptides, non-protein compounds, and / or nucleic acid aptamers.
[0017] In some embodiments of the present invention, the sample is at least one of plasma, serum, tissue, and cell samples.
[0018] In some embodiments of the present invention, the sample includes an enriched sample obtained through pretreatment, and the pretreatment step includes or does not include a step of incubating with nanoparticles; that is, the biomarker combination of the present invention can be used to evaluate the degree of red blood cell contamination of any sample, and the degree of contamination can be accurately identified for samples with or without nanoparticle pretreatment.
[0019] In some embodiments of the present invention, the product includes a reagent, a kit, a test strip, a chip, a device, and a detection system.
[0020] In a third aspect of the present invention, there is provided a product, which comprises a reagent for detecting the above-mentioned biomarker combination.
[0021] In some embodiments of the present invention, the reagent is quantitatively detected by at least one of mass spectrometry, chromatography, surface-enhanced Raman spectroscopy, Western blotting, flow cytometry, protein array, immunoprecipitation, immunosorption.
[0022] In some embodiments of the present invention, the reagent is selected from one or more of a substance specific for a biomarker, a biomarker-specific probe, and a protein chip.
[0023] In some embodiments of the present invention, the substance specific for a biomarker includes an antibody, a ligand protein, a polypeptide, a non-protein compound, and / or a nucleic acid aptamer.
[0024] In some embodiments of the present invention, the product comprises a reagent, a kit, a test strip, a chip, a device, a detection system.
[0025] In some embodiments of the present invention, the product further comprises a reagent for assisting in detecting the protein expression level.
[0026] In some embodiments of the present invention, the test sample is at least one of plasma, serum, tissue, and cell sample.
[0027] In some embodiments of the present invention, the sample includes an enriched sample obtained by pre-treatment.
[0028] In a fourth aspect of the present invention, there is provided a method for evaluating the degree of red blood cell contamination, comprising the following steps:
[0029] 1) Collect a sample to be tested and determine the abundance of the biomarkers in the above-mentioned biomarker combination for evaluating the degree of red blood cell contamination;
[0030] 2) Calculate the contamination index and / or the number of red blood cells of the sample to be tested, and the calculation formula of the contamination index is as follows:
[0031] Contamination index = (abundance of PRPS1 + abundance of ADD2 + abundance of PA2G4 + abundance of OLA1 + abundance of ACLY + abundance of PIP4K2A + abundance of FLOT1 + abundance of CFL1 + abundance of ADD1 + abundance of SNCA + abundance of STOM + abundance of RAP1B
[0032] + Abundance of PPIA + Abundance of MPP1 + Abundance of EPB42 + Abundance of RAN + Abundance of CD59 + Abundance of HSPA8 + Abundance of SLC4A1 + Abundance of HBD + Abundance of EPB41 + Abundance of HBA1 + Abundance of BLVRB + Abundance of GAPDH + Abundance of PGK1
[0033] + (Abundance of SPTB + Abundance of HBB + Abundance of SPTA1 + Abundance of CA1 + Abundance of ANK1) / Total protein abundance;
[0034] The calculation method of the number of red blood cells is: mapping the contamination index to the actual cell count through a regression model.
[0035] 3) Predict the contamination situation of the sample based on the calculated contamination index and / or the number of red blood cells of the sample to be tested: the lower the contamination index and the number of red blood cells of the sample to be tested, the lower the contamination degree; the higher the contamination index and / or the number of red blood cells of the sample to be tested, the higher the contamination degree.
[0036] In some embodiments of the present invention, the regression model includes: Polynomial Regression, Support Vector Regression, Regression Tree, and Multivariate Adaptive Regression Splines or Spline Regression; preferably Spline Regression.
[0037] In some embodiments of the present invention, the contamination index is compared with a defined value. If it is higher than the defined value, the contamination degree is higher; if it is lower than the defined value, the contamination degree is lower.
[0038] In some embodiments of the present invention, the defined value is 0.04.
[0039] In some embodiments of the present invention, the sample is at least one of plasma, serum, tissue, and cell sample.
[0040] In some embodiments of the present invention, the sample includes an enriched sample obtained through pretreatment, and the pretreatment step includes or does not include a step of incubating with nanoparticles; that is, the foregoing marker combination can be used to evaluate the red blood cell contamination degree of any sample, and the contamination degree can be accurately identified for samples with or without nanoparticle pretreatment.
[0041] In some embodiments of the present invention, the pretreatment including the step of incubating with nanoparticles specifically comprises the following steps: incubating the sample with nanoparticles, washing, denaturing, and enzymatically digesting after the incubation is completed; in the pretreatment, the sample is incubated with nanoparticles, and the sample and the nanoparticles combine to form a protein corona (soft protein corona and hard protein corona), and then the soft protein corona is removed by washing, and the hard protein corona with high affinity is retained, and then the sample that can be used for proteomic analysis is obtained by denaturing and enzymatically digesting.
[0042] In some embodiments of the present invention, the temperature of the incubation is 20 - 40 °C; it can also be 20 - 25 °C, 25 - 30 °C, 30 - 35 °C, 35 - 40 °C, and it can also be 26 °C, 27 °C, 28 °C, 29 °C, 30 °C, 31 °C, 32 °C, 33 °C or 34 °C. In some embodiments of the present invention, the time of the incubation is 30 - 120 min; it can also be 30 - 120 min, 30 - 120 min, 30 - 120 min, 30 - 120 min, 30 - 120 min, 30 - 120 min or 30 - 120 min.
[0043] In some embodiments of the present invention, the rotation speed of the incubation is 100 - 500 revolutions per minute; it can also be 100 revolutions per minute, 200 revolutions per minute, 300 revolutions per minute, 400 revolutions per minute or 500 revolutions per minute.
[0044] In some embodiments of the present invention, the material of the nanoparticles includes but is not limited to silica, molecular sieve, Fe3O4, liposome, polymer nanoparticles, etc.
[0045] In some embodiments of the present invention, the type of the nanoparticles includes but is not limited to solid spherical, mesoporous structure, hollow mesoporous and hierarchical porous structure.
[0046] In some embodiments of the present invention, the particle size of the nanoparticles is 300 - 1000 nm, and it can also be 400 nm, 500 nm, 600 nm, 700 nm, 800 nm or 900 nm.
[0047] In some embodiments of the present invention, the mass - volume ratio of the nanoparticles to the sample is 3 - 15 mg / mL; it can also be 3 - 6 mg / mL, 6 - 9 mg / mL, 9 - 12 mg / mL or 12 - 15 mg / mL.
[0048] In some embodiments of the present invention, the sample is a diluted sample, specifically, the sample to be measured is diluted by a diluent.
[0049] In some embodiments of the present invention, the diluent includes an ionic surfactant and a buffer solution.
[0050] In some embodiments of the present invention, the ionic surfactant includes 3-[(3-cholamidopropyl)dimethylammonio]-1-propanesulfonate (CHAPS), 3-[(3-cholamidopropyl)dimethylammonio]-2-hydroxy-1-propanesulfonate (CHAPSO), cetyltrimethylammonium bromide (CTAB), sodium dodecyl sulfate (SDS), sodium lauroyl sarcosinate (sarkosyl) or / and dodecyltrimethylammonium bromide (DTAB); preferably 3-[(3-cholamidopropyl)dimethylammonio]-1-propanesulfonate (CHAPS).
[0051] In some embodiments of the present invention, the concentration of the ionic surfactant is 0.01 - 0.1 w / w%, and can also be 0.03 w / w%, 0.04 w / w%, 0.05 w / w%, 0.06 w / w% or 0.07 w / w%.
[0052] In some embodiments of the present invention, the buffer solution includes PBS buffer solution, Tris buffer solution, etc.
[0053] In some embodiments of the present invention, the diluent further includes a pH regulator, such as ammonia water, sodium hydroxide, sodium bicarbonate, sodium carbonate, etc.
[0054] In some embodiments of the present invention, the pH of the diluent is 7 - 11, and can also be 7 - 8, 8 - 9, 9 - 10 or 10 - 11, preferably 10 - 11.
[0055] In some embodiments of the present invention, the dilution factor is 2 - 8, where the dilution factor is the ratio of the volume after dilution to the volume before dilution; the dilution factor can also be 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5 or 8.
[0056] In some embodiments of the present invention, the washing includes first cleaning the soft corona layer with a buffer solution, and then centrifugally washing to remove the soft protein corona.
[0057] In some embodiments of the present invention, the number of centrifugation times is 2 - 5 times, and can also be 2 times, 3 times, 4 times or 5 times.
[0058] In some embodiments of the present invention, the conditions for single centrifugation are 2000 - 10000 g, 5 - 15 min; where the centrifugal force can also be 2000 - 4000 g, 4000 - 6000 g, 6000 - 8000 g or 8000 - 10000 g; the centrifugation time can also be 5 - 8 min, 8 - 12 min or 12 - 15 min.
[0059] In some embodiments of the present invention, the pretreatment further includes a purification step; the purification method includes but is not limited to methods such as using commercial purification kits; it can be column purification reagents, magnetic bead purification reagents, gel electrophoresis reagents, etc.
[0060] In a fifth aspect of the present invention, there is provided a system for evaluating the level of red blood cell contamination, including the following modules:
[0061] a) Data collection module: Collect a sample to be tested and measure the abundances of the markers in the above-mentioned marker combination for evaluating the degree of red blood cell contamination.
[0062] b) Model calculation module: Calculate the contamination index and / or the number of red blood cells of the sample to be tested. The calculation formula of the contamination index is as follows:
[0063] Contamination index = (abundance of PRPS1 + abundance of ADD2 + abundance of PA2G4 + abundance of OLA1 + abundance of ACLY + abundance of PIP4K2A + abundance of FLOT1 + abundance of CFL1 + abundance of ADD1 + abundance of SNCA + abundance of STOM + abundance of RAP1B + abundance of PPIA + abundance of MPP1 + abundance of EPB42 + abundance of RAN + abundance of CD59 + abundance of HSPA8 + abundance of SLC4A1 + abundance of HBD + abundance of EPB41 + abundance of HBA1 + abundance of BLVRB + abundance of GAPDH + abundance of PGK1 + abundance of SPTB + abundance of HBB + abundance of SPTA1 + abundance of CA1 + abundance of ANK1) / total protein abundance;
[0064] The calculation method of the number of red blood cells is: map the contamination index to the actual cell count through a regression model.
[0065] c) Output prediction module: Predict the contamination situation of the sample based on the calculated contamination index and / or the number of red blood cells of the sample to be tested: the lower the contamination index and / or the number of red blood cells of the sample to be tested, the lower the contamination degree; the higher the contamination index and / or the number of red blood cells of the sample to be tested, the higher the contamination degree.
[0066] In some embodiments of the present invention, the contamination index is compared with a defined value. If it is higher than the defined value, a higher contamination degree is output. If it is lower than the defined value, a lower contamination degree is output.
[0067] In a sixth aspect of the present invention, there is provided a computing device, including:
[0068] At least one processing unit; and
[0069] At least one memory, the memory is coupled to the processing unit and stores a program for execution by the processing unit. When the program is executed by the processor, the processor realizes the above-mentioned evaluation of the degree of red blood cell contamination.
[0070] In a seventh aspect of the present invention, there is provided a computer-readable storage medium storing a computer program, which when executed by a processor implements the steps of evaluating the degree of red blood cell contamination described above.
[0071] In an eighth aspect of the present invention, there is provided a method for constructing a model for differentiating benign pulmonary nodules and lung cancer, comprising:
[0072] 1) Obtaining imaging data and clinically relevant data of patients with benign pulmonary nodules and lung cancer to construct a data set;
[0073] 2) Screening out key features by machine learning methods to construct an algorithm model for differentiating benign pulmonary nodules and lung cancer.
[0074] In the method for constructing the model, when obtaining clinically relevant data such as proteomics, the method described in the fourth aspect of the present invention can be used to evaluate the contamination level of red blood cells and thus exclude contaminated samples, so as to prevent screening out proteins related to contamination as biomarkers of the disease.
[0075] In some embodiments of the present invention, the machine learning method includes logistic regression, K-nearest neighbor, support vector machine, random forest, gradient boosting, multi-layer perceptron, adaptive boosting or extremely randomized trees; preferably extremely randomized trees.
[0076] In some embodiments of the present invention, the key features include: tumor size, CA-125, IGHV4-4, IGLV9-49, TIMM44, GMDS, SELENOF, UGDH, FCGR3A, TNFSF13, PAPLN, IGLV1-51, SLC3A2, DBH, AZU1, ANOS1, COL15A1, HAMP, FGL1, SKIV2L, ATP6AP1, FSCN1, HID1, PLVAP, PDGFD, SPTBN5, IMPAD1, RRN3, VPS18 and IGLC7.
[0077] In a ninth aspect of the present invention, there is provided a marker combination for differentiating benign pulmonary nodules and lung cancer, comprising tumor size, CA-125, IGHV4-4, IGLV9-49, TIMM44, GMDS, SELENOF, UGDH, FCGR3A, TNFSF13, PAPLN, IGLV1-51, SLC3A2, DBH, AZU1, ANOS1, COL15A1, HAMP, FGL1, SKIV2L, ATP6AP1, FSCN1, HID1, PLVAP, PDGFD, SPTBN5, IMPAD1, RRN3, VPS18 and IGLC7.
[0078] In the tenth aspect of the present invention, there is provided an application of a substance for detecting the marker combination for differentiating benign pulmonary nodules and lung cancer in the preparation of a product for differentiating benign pulmonary nodules and lung cancer.
[0079] In some embodiments of the present invention, the substance for detecting the marker combination for differentiating benign pulmonary nodules and lung cancer includes substances for quantitatively detecting CA-125, IGHV4-4, IGLV9-49, TIMM44, GMDS, SELENOF, UGDH, FCGR3A, TNFSF13, PAPLN, IGLV1-51, SLC3A2, DBH, AZU1, ANOS1, COL15A1, HAMP, FGL1, SKIV2L, ATP6AP1, FSCN1, HID1, PLVAP, PDGFD, SPTBN5, IMPAD1, RRN3 and VPS18 by at least one of mass spectrometry, chromatography, surface-enhanced Raman spectroscopy, Western blotting, flow cytometry, protein array, immunoprecipitation, immunosorption; and also includes substances for detecting the tumor size by imaging and pathology.
[0080] In some embodiments of the present invention, the product includes a reagent, a kit, a test strip, a system, a device or a chip.
[0081] In the eleventh aspect of the present invention, there is provided a system for differentiating benign pulmonary nodules and lung cancer, which includes:
[0082] An acquisition module, configured to acquire the detection results of each marker in the marker combination of the sample to be tested; wherein, the marker combination is the marker combination described in the ninth aspect of the present invention;
[0083] A prediction module, configured to input the detection results of all markers into the model described in the eighth aspect of the present invention to obtain the differentiation result of the sample to be tested.
[0084] In the twelfth aspect of the present invention, there is provided a computer-readable storage medium storing a computer program, which when executed by a processor, can implement the functions of the system described in the eleventh aspect of the present invention.
[0085] In view of the frequently occurring sample contamination problem in proteomics, the present invention for the first time discloses a biomarker combination for evaluating red blood cell contamination, including the following biomarkers: PRPS1, ADD2, PA2G4, OLA1, ACLY, PIP4K2A, FLOT1, CFL1, ADD1, SNCA, STOM, RAP1B, PPIA, MPP1, EPB42, RAN, CD59, HSPA8, SLC4A1, HBD, EPB41, HBA1, BLVRB, GAPDH, PGK1, SPTB, HBB, SPTA1, CA1 and ANK1; based on this biomarker combination, red blood cell contaminated samples can be accurately identified and excluded, and the number of red blood cells can be quantified, which is used to exclude red blood cell contaminated samples obtained in enrichment methods. Moreover, this biomarker combination has universality and is not affected by the type of nanoparticles. It has high sensitivity and specificity and can be used for the identification of red blood cell contaminated samples in various nanoparticle-based or non-nanoparticle-based enrichment methods; it is also conducive to accurately screening biomarkers for evaluating diseases through proteomics and can be used in clinical scenarios such as early cancer screening and personalized medicine.
[0086] After accurately identifying and excluding red blood cell contaminated samples, the present application further constructs a model for differentiating benign lung nodules and lung cancer by machine learning method, and its AUC can reach 0.8, showing high diagnostic efficacy. BRIEF DESCRIPTION OF THE DRAWINGS
[0087] Figure 1 Shown for the development of the red blood cell contamination quality control biomarker group; (a) Red blood cell related biomarker screening scheme; (b) Number of peptide precursor ions identified in the discovery dataset; (c) Number of proteomes; (d) Abundance distribution of red blood cell biomarkers; (e) Spearman correlation analysis of 30 red blood cell biomarkers; (f) Biomarker validation experimental design; (g) Relationship between the Z-value normalized intensity of 30 biomarkers and the red blood cell incorporation ratio; (h) Correlation between red blood cell count and contamination index (RBC: red blood cell-rich plasma, PPP: red blood cell-poor plasma).
[0088] Figure 2 Is the red blood cell contamination index for the lung cancer cohort. (a) Distribution of the red blood cell contamination index for the lung cancer cohort; (b) Diagnostic efficacy in the test set; (c) Importance ranking of the top-30 features screened by machine learning.
[0089] Figure 3 Is the ROC curve for the training set and test set in Example 4. DETAILED DESCRIPTION OF THE INVENTION
[0090] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification.
[0091] Before further describing the specific embodiments of the present invention, it should be understood that the protection scope of the present invention is not limited to the following specific embodiments; it should also be understood that the terms used in the embodiments of the present invention are for describing specific embodiments, rather than for limiting the protection scope of the present invention. The test methods without specific conditions indicated in the following embodiments are usually carried out according to conventional conditions or according to the conditions recommended by each manufacturer.
[0092] When an embodiment gives a numerical range, it should be understood that unless otherwise specified in the present invention, any value between the two endpoints of each numerical range and either of the two endpoints can be selected. Unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those skilled in the art of this technology. In addition to the specific methods, devices, and materials used in the embodiments, according to the knowledge of those skilled in the art of this technology and the description of the present invention, any methods, devices, and materials of the prior art similar or equivalent to those described in the embodiments of the present invention can also be used to implement the present invention.
[0093] Example 1 Screening and construction of a pretreatment process for nanoparticle-based enrichment of plasma proteome
[0094] The pretreatment process for nanoparticle-based plasma proteome includes: the combination of plasma samples and nanoparticles to form a protein corona (soft protein corona and hard protein corona), the removal of the soft protein corona, protein denaturation, and enzymatic digestion, etc.
[0095] Silica nanoparticles are used to enrich low-abundance proteins in plasma. The specific steps are as follows:
[0096] First, take 15 - 30 μL of plasma sample and dilute it with 75 - 100 μL of 1× phosphate - buffered saline (PBS, called buffer 2, pH = 10 - 11) containing 0.05 w / w% 3 - [(3 - cholamidopropyl)dimethylammonio]-1 - propanesulfonate and 0.02 v / v% ammonia water. Subsequently, add 0.3 - 1.0 mg of solid spherical silica nanoparticles NP (particle size 300 - 1000 nm) solution and incubate it in a constant - temperature shaker at 30 °C for 30 - 120 minutes at a rotation speed of 100 - 500 revolutions per minute. After incubation, wash the soft corona layer with buffer 2 diluted with water to a concentration of 33 v / v%, and repeat the washing three times by centrifugation at 2500 - 10000 g for 10 minutes to remove the soft protein corona. In each centrifugation, discard the supernatant and repeat the washing of the precipitate. Finally, perform protein denaturation, enzymatic digestion, and desalting and purification processes on the hard corona layer. The specific process is as follows: First, use 50 μL of 8 M urea and 2 M thiourea to denature the protein. Subsequently, add Tris(2 - carboxyethyl)phosphine (TCEP) with a final concentration of 10 mM and iodoacetamide (IAA) with a concentration of 40 mM under light - protected conditions and react for 30 - 60 minutes to complete reduction and alkylation. Then, dilute the urea concentration to less than 1.2 M with 100 mM ABB diluent and add 0.5 μg - 2 μg of trypsin according to a mass ratio of 1:10 - 1:25 for overnight enzymatic digestion. Finally, add 30 - 50 μL of 10% trifluoroacetic acid (TFA) to terminate the reaction. The peptide segments produced by enzymatic digestion are purified using a peptide desalting column and subjected to vacuum drying treatment.
[0097] Finally, after the plasma sample undergoes the pretreatment process, in the 24 - minute Astral mass spectrometry nDIA analysis, 3000 - 6000 proteomes can be stably identified per sample, and the batch coefficient of variation (CV) is less than 12%. The parameters of the mass spectrometry analysis are as follows: First, load approximately 200 - 400 ng of peptide segments onto the trapping column, and then separate them using an analytical column (inner diameter 75 μm × length 15 cm, particle size 1.9 μm, customized) through a 24 - minute LC - MS method. The initial liquid - phase gradient condition is 8% buffer B (buffer B: 80% acetonitrile containing 0.1% formic acid (v / v); buffer A: 0.1% formic acid (v / v) dissolved in mass - spectrometry - grade ultrapure water), which rises to 10% B within 1.5 minutes, then rises to 30% B within 16 minutes, and finally rises to 40% B within 2 minutes. Each run includes a 4.3 - minute column cleaning and equilibration step before. The eluted peptide segments are analyzed by an Orbitrap Astral mass spectrometer, and the parameter settings are: FAIMS voltage - 42 V, full - scan resolution 240,000, mass - to - charge ratio scan range 380 - 980 Th; the MS / MS scan range is the same as the full - scan, using the DIA mode (data - independent acquisition), and the isolation window width is 2 Da.
[0098] Example 2: Discovery and Validation of Biomarkers for Red Blood Cell Contamination
[0099] Based on the nanoparticle-based plasma protein pretreatment method and the species-verified proteomic spectral library (download path: https: / / www.uniprot.org / ) established above, experiments were designed to establish markers and algorithms for evaluating red blood cell contamination of plasma samples.
[0100] First, pure red blood cell samples (RBC) and plasma samples without red blood cells (PPP) were collected and mixed in accordance with Figure 1 proportion a to obtain samples with different degrees of red blood cell contamination. After nanoparticle enrichment, more protein identifications were achieved in plasma samples with different red blood cell contaminations as the red blood cell content gradually increased.
[0101] The results showed that an average of about 3,500 proteomes were identified in RBC samples, and about 1,500 proteomes were identified in PPP samples ( Figure 1 b - c).
[0102] Then, proteins with a missing rate of less than 50% in all samples were screened. The top 100 proteins with the highest correlation with red blood cell concentration (Spearman r > 0.95) were screened by Mfuzz clustering analysis. The 30 proteins with the highest abundance were selected as markers for evaluating red blood cell contamination, including 30 markers such as HBA1 (hemoglobin α1), HBB (hemoglobin β), and SPTA1 (spectrin α). The abundance distribution of red blood cell markers is shown in Figure 1 d;. The median of the spearman correlation of these 30 proteins was 0.94 ( Figure 1 e). The specific information of the 30 markers is as follows: PRPS1, ADD2, PA2G4, OLA1, ACLY, PIP4K2A, FLOT1, CFL1, ADD1, SNCA, STOM, RAP1B, PPIA, MPP1, EPB42, RAN, CD59, HSPA8, SLC4A1, HBD, EPB41, HBA1, BLVRB, GAPDH, PGK1, SPTB, HBB, SPTA1, CA1, ANK1. The peptide sequences corresponding to these proteins are shown in the following table.
[0103] Table 1
[0104]
[0105]
[0106] The applicant established a contamination index based on the above 30 biomarkers, which is the ratio of the sum of the abundances of the 30 biomarkers in each sample to the abundance of all proteins in this sample. Specifically, the contamination index (erythrocyte contamination index = Σ biomarker intensity / total protein intensity) is used to evaluate the degree of erythrocyte contamination in this sample. The defined value is 0.04. If the contamination index is higher than 0.04, the sample can be determined to be a contaminated sample and needs to be excluded. The higher the contamination index, the more serious the erythrocyte contamination. Immediately afterwards, another dataset was used to verify the feasibility of the 30 proteins selected for evaluating erythrocyte contamination.
[0107] The applicant collected pure erythrocyte and pure PPP samples from 12 people, and then used an erythrocyte counter to count each sample to determine the blood cell distribution of each sample, so as to determine the purity of the erythrocyte and PPP samples. Then, the erythrocytes of every three patients were mixed and the erythrocytes were mixed. Finally, 4 groups of pure erythrocyte samples (RBC) and pure PPP samples ( Figure 1 f). We performed 10-fold dilutions of the erythrocyte and PPP samples on the 4 groups of samples respectively, and then used an erythrocyte counter to count each diluted sample to determine the absolute erythrocyte content of each sample.
[0108] A correlation analysis was performed between the erythrocyte concentration and the abundances of the 30 biomarkers, and the results were good. And the spline regression was used to fit the absolute erythrocyte content and contamination, and the result showed that R 2 = 0.94( Figure 1 h), showing a good correlation. It shows that the contamination index can be used to evaluate the degree of erythrocyte contamination in the sample.
[0109] The applicant further quantitatively analyzed the number of erythrocytes through the contamination index. Specifically, the constructed spline regression model was used to map the index to the actual cell count. The model construction method is as follows: Based on the 10-fold dilution data of the 4 groups of samples in the validation set (each dilution gradient contains 3 biological replicates, a total of 120 samples), a spline function (Univariate Spline) was used for fitting and modeling. With the contamination index as the independent variable and the actual erythrocyte count as the dependent variable, the number of nodes k = 3 and the smoothing coefficient s = 50 were set (through multiple cross-validations, this parameter setting can maximize the balance between the flexibility and smoothness of the model). Finally, the erythrocyte count prediction curve (R 2= 0.94), the model expression is: Red blood cell count = Spline(Contamination index), where coeffs[6.67520162, 6.34688444, 7.72838449, 7.53426373] / knots[0.12249392, 0.75213502].
[0110] To further verify these biomarkers, red blood cell and platelet-poor plasma (PPP) samples from six patients were collected and mixed into two groups (each group containing samples from three patients). Subsequently, the red blood cell samples and PPP samples were serially diluted in a ten-step dilution method ( Figure 1 f). Among them, it can be seen that with the gradual dilution, all 30 proteins used to evaluate red blood cell contamination showed a downward trend ( Figure 1 g).
[0111] Example 3 Universality of Red Blood Cell Contamination Biomarkers
[0112] To verify the applicability of the identified biomarkers in samples treated with other types of nanoparticles, two types of frequently reported nanoparticles in the literature were selected: NaY zeolite (particle size 300 - 700 nm) and silanol-functionalized iron oxide (particle size 400 - 700 nm). PPP and RBC samples were prepared by collecting donor blood, and 100 μL of each was mixed and serially diluted in a ten-step gradient. The diluted plasma samples were treated with NaY zeolite and silanol-functionalized iron oxide nanoparticles respectively according to the steps of Example 1, and nDIA analysis was performed.
[0113] The results showed that: with the increase in red blood cell concentration, the number of peptide precursors and proteome identifications in samples treated with both types of nanoparticles increased significantly, and the number of proteomes detected in a single injection exceeded 5500. The 30 red blood cell-related biomarkers showed a high correlation with the degree of red blood cell contamination in samples treated with NaY zeolite (median spearman correlation between the abundances of 30 proteins evaluating red blood cell contamination and the red blood cell contamination count: 0.92, range 0.92 - 0.96) and NP81-treated samples (median: 0.95, range 0.89 - 0.94).
[0114] It also shows that the method for evaluating the degree of red blood cell contamination of the present invention is compatible with various nanoparticles such as molecular sieves and Fe3O4, providing a unified framework for cross-laboratory data standardization.
[0115] Example 4 Application Example
[0116] To evaluate the application effect of the technology for enriching low-abundance proteins in plasma based on nanoparticles and assessing red blood cell contamination, 193 subjects were included in this study, including 42 patients with benign lung nodules and 151 patients with early-stage malignant tumors.
[0117] 1. Criteria for the diagnosis of benign lung nodules
[0118] 1.1 Benign lung nodules
[0119] a. The pathological report diagnosed it as a benign lung nodule;
[0120] b. In the case where a pathological report could not be obtained, imaging results showed small nodules, and imaging follow-up data over more than 1 year showed that the nodules were stable (volume change < 25%) or the volume doubling time was greater than 400 days.
[0121] 1.2 Lung cancer: The pathological report diagnosed it as lung cancer (both small cell carcinoma and non-small cell carcinoma were included)
[0122] a. Pathologically confirmed as primary lung cancer;
[0123] b. Stage: I - III;
[0124] c. The histological type was clearly recorded.
[0125] 2. Inclusion and exclusion criteria
[0126] 2.1 Sample inclusion criteria
[0127] a. Aged between 25 and 80 years old, regardless of gender;
[0128] b. The size of the lung nodules shown by CT screening was 5 - 30 mm;
[0129] c. Had not received any lung nodule / lung cancer-related treatments (including surgery, chemotherapy, radiotherapy, targeted therapy, immunotherapy, interventional therapy, etc.);
[0130] d. The clinical and imaging data were complete;
[0131] e. Could provide low-dose spiral CT reports within the recent 3 months;
[0132] f. Voluntarily signed the informed consent form.
[0133] 2.2 Sample exclusion criteria
[0134] a. Had a history of previous tumors;
[0135] b. Clinically uncontrolled active infections, such as acute pneumonia, pulmonary tuberculosis, etc.;
[0136] c. Had received any lung nodule-related treatments such as antibiotics and hormones within the past 4 weeks;
[0137] d. Combined with other tumors and serious diseases of the heart, liver, kidney, brain, blood and other systems;
[0138] e. Participated in other clinical trials within the past 3 months;
[0139] f. Combined with diseases affecting protein content such as hepatic and renal insufficiency, hypoproteinemia, etc.;
[0140] g. In the gestational or lactation period.
[0141] Such cases often present unclear diagnostic results in routine CT imaging examinations and need to be confirmed by surgical resection. All plasma samples were collected by vacuum blood collection tubes containing ethylenediaminetetraacetic acid. Some patients underwent secondary sampling to evaluate the stability of pretreatment; samples with repeated processing were not included in subsequent modeling. After centrifuging the blood (3000g, 4°C, 15 minutes), the plasma was collected and stored at 80°C. Then, peptide samples were obtained using the process of enriching low-abundance plasma proteins based on nanoparticles in Example 1, combined with Astral instrument analysis to achieve quantitative analysis of peptides and proteins. Proteomics analysis showed that an average of 4413 proteomes were identified in each plasma sample, and both biological and technical replicates showed stable coefficients of variation. The erythrocyte situation in the samples was evaluated by the aforementioned pollution degree assessment method, and it was found that 3 samples had detectable pollution indicators ( Figure 2 a), and these contaminated samples were excluded from subsequent analysis.
[0142] The subjects were divided into a training set and a test set at a ratio of 7:3; eight machine learning algorithms (logistic regression, KNeighbors Classifier, Support Vector Machine (SVC), Random Forest Classifier, Gradient Boosting Classifier, Multi-Layer Perceptron (MLPC lassifier), AdaBoost Classifier, and Extra TreesClassifier) were used to distinguish between patients with benign pulmonary nodules and lung cancer. After feature selection, they were ranked according to importance; the ROC curves of each model in the training set and the test set are referenced Figure 3, it can be seen that the modeling results of the previous 30 important proteins or clinical features show that the extreme random tree model performs optimally in terms of indicators such as F1 score, accuracy, precision, recall rate, and AUC value. This model screened out tumor size, CA-125, and 28 proteins as key features, among which the 28 proteins include IGHV4-4, IGLV9-49, TIMM44, GMDS, SELENOF, UGDH, FCGR3A, TNFSF13, PAPLN, IGLV1-51, SLC3A2, DBH, AZU1, ANOS1, COL15A1, HAMP, FGL1, SKIV2L, ATP6AP1, FSCN1, HID1, PLVAP, PDGFD, SPTBN5, IMPAD1, RRN3, VPS18, and IGLC7, confirming that the extreme random tree model constructed based on the above key features can effectively distinguish between patients with benign lung nodules and lung cancer, and the classification accuracy in the test set reaches 82% (AUC = 0.80)( Figure 2 b-c and Figure 3 ), with excellent diagnostic efficacy.
[0143] The above embodiments merely illustrate the principles and effects of the present invention and are not intended to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.
Claims
1. A marker combination for evaluating the degree of red blood cell contamination, the marker combination comprising PRPS1, ADD2, PA2G4, OLA1, ACLY, PIP4K2A, FLOT1, CFL1, ADD1, SNCA, STOM, RAP1B, PPIA, MPP1, EPB42, RAN, CD59, HSPA8, SLC4A1, HBD, EPB41, HBA1, BLVRB, GAPDH, PGK1, SPTB, HBB, SPTA1, CA1 and ANK1.
2. Use of the marker combination according to claim 1 or a reagent for detecting the marker combination according to claim 1 in at least one of the following: 1) Evaluating the degree of red blood cell contamination; 2) Preparing a product for evaluating the degree of red blood cell contamination; 3) Calculating the number of red blood cells; 4) Preparing a product for calculating the number of red blood cells; 5) Screening biomarkers for diseases; 6) Preparing a product for screening biomarkers for diseases; Preferably, the reagent is detected by at least one method among mass spectrometry, chromatography, surface enhanced Raman spectroscopy, western blotting, flow cytometry, protein array, immunoprecipitation, and immunosorption; Preferably, the reagent comprises one or more of a substance specific for the marker, a marker-specific probe, and a protein chip; Preferably, the substance specific for the marker comprises an antibody, a ligand protein, a polypeptide, a non-protein compound, and / or a nucleic acid aptamer.
3. A product, characterized in that, The product comprises a reagent for detecting the marker combination according to claim 1; Preferably, the reagent is detected by at least one method among mass spectrometry, chromatography, surface enhanced Raman spectroscopy, western blotting, flow cytometry, protein array, immunoprecipitation, and immunosorption; Preferably, the reagent comprises one or more of a substance specific for the marker, a marker-specific probe, and a protein chip; Preferably, the substance specific for the marker comprises an antibody, a ligand protein, a polypeptide, a non-protein compound, and / or a nucleic acid aptamer; Preferably, the product comprises a reagent, a kit, a test strip, a chip, a device, and a detection system.
4. A method for evaluating the degree of red blood cell contamination, comprising the following steps: 1) Collecting a sample to be tested and measuring the abundance of the markers in the marker combination according to claim 1; 2) Calculating the contamination index and / or the number of red blood cells of the sample to be tested, and the calculation formula of the contamination index is as follows: Pollution index = (abundance of PRPS1 + abundance of ADD2 + abundance of PA2G4 + abundance of OLA1 + abundance of ACLY + abundance of PIP4K2A + abundance of FLOT1 + abundance of CFL1 + abundance of ADD1 + abundance of SNCA + abundance of STOM + abundance of RAP1B + abundance of PPIA + abundance of MPP1 + abundance of EPB42 + abundance of RAN + abundance of CD59 + abundance of HSPA8 + abundance of SLC4A1 + abundance of HBD + abundance of EPB41 + abundance of HBA1 + abundance of BLVRB + abundance of GAPDH + abundance of PGK1 + abundance of SPTB + abundance of HBB + abundance of SPTA1 + abundance of CA1 + abundance of ANK1) / abundance of total protein; The calculation method of the number of red blood cells is as follows: mapping the pollution index to the actual cell count through a regression model; 3) Predict the pollution situation of the sample based on the calculated pollution index and / or the number of red blood cells of the sample to be tested: the lower the pollution index and the number of red blood cells of the sample to be tested, the lower the degree of red blood cell pollution of the sample to be tested; the higher the pollution index and / or the number of red blood cells of the sample to be tested, the higher the degree of red blood cell pollution of the sample to be tested; Preferably, the sample is at least one of plasma, serum, tissue and / or cell sample; Preferably, the sample includes an enriched sample obtained through pretreatment, and the pretreatment step includes or does not include the step of incubating with nanoparticles; Preferably, the pretreatment specifically includes the following steps: incubating the sample to be tested with nanoparticles, and performing washing, denaturation and enzymatic digestion after the incubation is completed.
5. A system for evaluating the degree of red blood cell pollution, comprising the following modules: a) Data collection module: collect the sample to be tested and measure the abundance of the markers in the marker combination described in claim 1; b) Model calculation module: calculate the pollution index and / or the number of red blood cells of the sample to be tested, and the calculation formula of the pollution index is as follows: Pollution index = (abundance of PRPS1 + abundance of ADD2 + abundance of PA2G4 + abundance of OLA1 + abundance of ACLY + abundance of PIP4K2A + abundance of FLOT1 + abundance of CFL1 + abundance of ADD1 + abundance of SNCA + abundance of STOM + abundance of RAP1B + abundance of PPIA + abundance of MPP1 + abundance of EPB42 + abundance of RAN + abundance of CD59 + abundance of HSPA8 + abundance of SLC4A1 + abundance of HBD + abundance of EPB41 + abundance of HBA1 + abundance of BLVRB + abundance of GAPDH + abundance of PGK1 + abundance of SPTB + abundance of HBB + abundance of SPTA1 + abundance of CA1 + abundance of ANK1) / abundance of total protein; The calculation method of the number of red blood cells is as follows: mapping the pollution index to the actual cell count through a regression model; c) Output prediction module: predict the pollution situation of the sample based on the calculated pollution index and / or the number of red blood cells of the sample to be tested: among them, the lower the pollution index and / or the number of red blood cells of the sample to be tested, the lower the pollution degree; the higher the pollution index and / or the number of red blood cells of the sample to be tested, the higher the pollution degree.
6. A computing device comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method described in claim 4 are implemented.
7. A method for constructing a model for differentiating benign pulmonary nodules and lung cancer, comprising: 1) Obtain imaging data and clinically relevant data of patients with benign pulmonary nodules and lung cancer, and construct a data set; 2) Screen out key features by machine learning methods and construct an algorithm model; Preferably, when obtaining proteomics clinical data in step 1), the method described in claim 4 can be used to evaluate the contamination level of red blood cells and thus exclude red blood cell-contaminated samples; Preferably, the machine learning method includes logistic regression, K-nearest neighbor, support vector machine, random forest, gradient boosting, multi-layer perceptron, adaptive boosting or extremely randomized trees; Preferably, the key features include tumor size, CA-125, IGHV4-4, IGLV9-49, TIMM44, GMDS, SELENOF, UGDH, FCGR3A, TNFSF13, PAPLN, IGLV1-51, SLC3A2, DBH, AZU1, ANOS1, COL15A1, HAMP, FGL1, SKIV2L, ATP6AP1, FSCN1, HID1, PLVAP, PDGFD, SPTBN5, IMPAD1, RRN3, VPS18 and IGLC7.
8. A biomarker combination for differentiating benign pulmonary nodules and lung cancer, comprising tumor size, CA-125, IGHV4-4, IGLV9-49, TIMM44, GMDS, SELENOF, UGDH, FCGR3A, TNFSF13, PAPLN, IGLV1-51, SLC3A2, DBH, AZU1, ANOS1, COL15A1, HAMP, FGL1, SKIV2L, ATP6AP1, FSCN1, HID1, PLVAP, PDGFD, SPTBN5, IMPAD1, RRN3, VPS18 and IGLC7.
9. Use of the biomarker combination according to claim 8 or a reagent for detecting the biomarker combination according to claim 8 in the preparation of a product for differentiating benign pulmonary nodules and lung cancer.
10. A system for differentiating benign pulmonary nodules and lung cancer, comprising: An acquisition module for obtaining the detection results of each biomarker in the biomarker combination of the sample to be tested; wherein, the biomarker combination is the biomarker combination according to claim 8; A prediction module for inputting the detection results of all biomarkers into the model according to claim 7 to obtain the differentiation result of the sample to be tested.
Citation Information
Patent Citations
Measuring method of complement sensitized erythrocyte
JP1995005171A
Laboratory diagnostic technique for paroxysmal nocturnal haemoglobinuria
RU2574968C1
Erythrocyte-derived extracellular vesicles and proteins associated with such vesicles as biomarkers for parkinson's disease
US20200271672A1
Classification, screening and diagnostic methods and apparatus
US20210123855A1