Marker combination for evaluating the degree of red blood cell contamination and application thereof

By assessing the degree of red blood cell contamination through a combination of biomarkers, this approach solves the problem of the inability to accurately quantify red blood cell contamination in existing technologies, thereby achieving accuracy in plasma proteomics analysis and specificity in disease biomarker screening.

CN120369957BActive Publication Date: 2025-12-30WESTLAKE LAB OF LIFE SCI & BIOMEDICINE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510492870.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-12-30
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

Current technologies cannot accurately quantify the degree of red blood cell contamination, leading to large errors in plasma proteomics analysis results and affecting the accuracy of disease biomarker screening.

Method used

A combination of biomarkers, including PRPS1, ADD2, PA2G4, OLA1, ACLY, PIP4K2A, FLOT1, CFL1, ADD1, SNCA, STOM, RAP1B, PPIA, MPP1, EPB42, RAN, CD59, HSPA8, SLC4A1, HBD, EPB41, HBA1, BLVRB, GAPDH, PGK1, SPTB, HBB, SPTA1, CA1, and ANK1, were used for quantitative detection by mass spectrometry, chromatography, and other methods. The contamination index was calculated to assess the red blood cell count and the degree of contamination.

Benefits of technology

This technology enables precise identification and quantification of the degree of red blood cell contamination, improves the accuracy of plasma proteomics analysis, reduces the false positive rate, and enhances the specificity of disease biomarker screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120369957B_ABST
    Figure CN120369957B_ABST
Patent Text Reader

Abstract

The application provides a marker combination for evaluating the degree of red blood cell contamination, the marker combination comprising PRPS1, ADD2, PA2G4, OLA1, ACLY, PIP4K2A, FLOT1, CFL1, ADD1, SNCA, STOM, RAP1B, PPIA, MPP1, EPB42, RAN, CD59, HSPA8, SLC4A1, HBD, EPB41, HBA1, BLVRB, GAPDH, PGK1, SPTB, HBB, SPTA1, CA1 and ANK1. Based on the marker combination, red blood cell contamination samples can be accurately identified and excluded, and biomarkers for evaluating diseases can be accurately screened through proteomics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biomedical technology, and in particular to a combination of biomarkers for assessing the degree of red blood cell contamination and their applications. Background Technology

[0002] Current methods for assessing red blood cell contamination rely on subjective experience or single biomarkers (such as HBA1) to evaluate contamination, failing to quantify the degree of contamination or correct for data bias, leading to a disconnect between analytical results and actual biological effects. Furthermore, circulating blood proteomics, as a core technology for non-invasive disease biomarker discovery, has significantly improved the depth of plasma protein detection in recent years, driven by mass spectrometry (MS) and nanoparticle (NP) enrichment methods. However, current technologies do not publicly address the contamination issues in NP-based plasma proteomics.

[0003] Furthermore, in clinical translation, current technologies often use red blood cell-specific proteins to screen biomarkers for disease evaluation. If proteins associated with red blood cell contamination are selected as biomarkers for disease, it can lead to increased false positive rates, decreased specificity, and failed clinical validation. For example, hemoglobin subunits released from residual red blood cells (such as HBA1) may interfere with the analysis of inflammation or tumor-related pathways, leading to incorrect associations with disease mechanisms. Therefore, there is an urgent need for a biomarker that can accurately assess the degree of red blood cell contamination in plasma proteomics samples. Summary of the Invention

[0004] In view of the shortcomings of the prior art described above, the purpose of this invention is to provide a combination of biomarkers for assessing the degree of red blood cell contamination, thereby solving the problems in the prior art.

[0005] To achieve the above and other related objectives, the present invention is obtained through the following technical solution.

[0006] In a first aspect, the present invention provides a combination of biomarkers for assessing the degree of red blood cell contamination, the biomarker combination comprising PRPS1, ADD2, PA2G4, OLA1, ACLY, PIP4K2A, FLOT1, CFL1, ADD1, SNCA, STOM, RAP1B, PPIA, MPP1, EPB42, RAN, CD59, HSPA8, SLC4A1, HBD, EPB41, HBA1, BLVRB, GAPDH, PGK1, SPTB, HBB, SPTA1, CA1, and ANK1.

[0007] A second aspect of the invention provides the use of a reagent for quantitatively detecting a combination of biomarkers in a sample in at least one of the following combinations, said combination including PRPS1, ADD2, PA2G4, OLA1, ACLY, PIP4K2A, FLOT1, CFL1, ADD1, SNCA, STOM, RAP1B, PPIA, MPP1, EPB42, RAN, CD59, HSPA8, SLC4A1, HBD, EPB41, HBA1, BLVRB, GAPDH, PGK1, SPTB, HBB, SPTA1, CA1, and ANK1;

[0008] (1) Assess the degree of red blood cell contamination;

[0009] (2) Prepare products for assessing the degree of red blood cell contamination;

[0010] (3) Calculate the number of red blood cells;

[0011] (4) Prepare a product for calculating the number of red blood cells;

[0012] (5) Screening for disease-related biomarkers;

[0013] (6) Prepare products for screening disease biomarkers.

[0014] In some embodiments of the present invention, the reagent is quantitatively detected by at least one of the following methods: mass spectrometry, chromatography, surface-enhanced Raman spectroscopy, Western blotting, flow cytometry, protein array, immunoprecipitation, and immunoadsorption.

[0015] In some embodiments of the present invention, the reagent is selected from one or more of substances specific to the marker, marker-specific probes, and protein chips.

[0016] In some embodiments of the present invention, the substances that are specific to the marker include antibodies, ligand proteins, polypeptides, non-protein compounds and / or nucleic acid aptamers.

[0017] In some embodiments of the present invention, the sample is at least one of plasma, serum, tissue, and cell samples.

[0018] In some embodiments of the present invention, the sample includes an enriched sample obtained after pretreatment, the pretreatment step of which may or may not include an incubation step with nanoparticles; that is, the biomarker combination of the present invention can be used to assess the degree of erythrocyte contamination in any sample, and can accurately identify the degree of contamination for samples with or without nanoparticle pretreatment.

[0019] In some embodiments of the present invention, the product includes reagents, reagent kits, test strips, chips, devices, and detection systems.

[0020] A third aspect of the present invention provides a product comprising a reagent for detecting the combination of the above-described markers.

[0021] In some embodiments of the present invention, the reagent is quantitatively detected by at least one of the following methods: mass spectrometry, chromatography, surface-enhanced Raman spectroscopy, Western blotting, flow cytometry, protein array, immunoprecipitation, and immunoadsorption.

[0022] In some embodiments of the present invention, the reagent is selected from one or more of substances specific to the marker, marker-specific probes, and protein chips.

[0023] In some embodiments of the present invention, the substances that are specific to the marker include antibodies, ligand proteins, polypeptides, non-protein compounds and / or nucleic acid aptamers.

[0024] In some embodiments of the present invention, the product includes reagents, reagent kits, test strips, chips, devices, and detection systems.

[0025] In some embodiments of the present invention, the product further includes a protein expression level detection reagent.

[0026] In some embodiments of the present invention, the test sample is at least one of plasma, serum, tissue, and cell samples.

[0027] In some embodiments of the present invention, the sample includes an enriched sample obtained through preprocessing.

[0028] A fourth aspect of the present invention provides a method for assessing the degree of red blood cell contamination, comprising the following steps:

[0029] 1) Collect the samples to be tested and determine the abundance of the markers in the above-mentioned combination of markers used to assess the degree of red blood cell contamination;

[0030] 2) Calculate the contamination index and / or red blood cell count of the sample to be tested. The formula for calculating the contamination index is as follows:

[0031] Pollution index = (PRPS1 abundance + ADD2 abundance + PA2G4 abundance + OLA1 abundance + ACLY abundance + PIP4K2A abundance + FLOT1 abundance + CFL1 abundance + ADD1 abundance + SNCA abundance + STOM abundance + RAP1B abundance)

[0032] +PPIA abundance +MPP1 abundance +EPB42 abundance +RAN abundance +CD59 abundance +HSPA8 abundance +SLC4A1 abundance +HBD abundance +EPB41 abundance +HBA1 abundance +BLVRB abundance +GAPDH abundance +PGK1 abundance

[0033] (+SPTB abundance+HBB abundance+SPTA1 abundance+CA1 abundance+ANK1 abundance) / total protein abundance;

[0034] The method for calculating the number of red blood cells is as follows: the pollution index is mapped to the actual cell count through a regression model.

[0035] 3) The contamination level of the sample is predicted based on the calculated contamination index and / or red blood cell count: the lower the contamination index and red blood cell count of the sample, the lower the degree of contamination; the higher the contamination index and / or red blood cell count of the sample, the higher the degree of contamination.

[0036] In some embodiments of the present invention, the regression model includes: polynomial regression, support vector regression, regression tree, and multivariate adaptive regression splines or spline regression; preferably spline regression.

[0037] In some embodiments of the present invention, the pollution index is compared with a threshold value. If the index is higher than the threshold value, the pollution level is higher; if the index is lower than the threshold value, the pollution level is lower.

[0038] In some embodiments of the present invention, the defined value is 0.04.

[0039] In some embodiments of the present invention, the sample is at least one of plasma, serum, tissue, and cell samples.

[0040] In some embodiments of the present invention, the sample includes an enriched sample obtained after pretreatment, the pretreatment step of which may or may not include an incubation step with nanoparticles; that is, the aforementioned combination of markers can be used to assess the degree of red blood cell contamination in any sample, and the degree of contamination can be accurately identified for samples with or without nanoparticle pretreatment.

[0041] In some embodiments of the present invention, the pretreatment including the incubation step with nanoparticles specifically includes the following steps: incubating the sample with nanoparticles, and after incubation, washing, denaturation, and enzymatic digestion; in the pretreatment, the sample is incubated with nanoparticles, and the sample and nanoparticles combine to form protein crowns (soft protein crowns and hard protein crowns), and the soft protein crowns are subsequently removed by washing, while the high-affinity hard protein crowns are retained, and the sample that can be used for proteomics analysis is obtained by denaturation and enzymatic digestion.

[0042] In some embodiments of the present invention, the incubation temperature is 20–40°C; it can also be 20–25°C, 25–30°C, 30–35°C, 35–40°C, or even 26°C, 27°C, 28°C, 29°C, 30°C, 31°C, 32°C, 33°C, or 34°C. In some embodiments of the present invention, the incubation time is 30–120 min; it can also be 30–120 min, 30–120 min, 30–120 min, 30–120 min, 30–120 min, or 30–120 min.

[0043] In some embodiments of the present invention, the incubation rotation speed is 100 to 500 rpm; it can also be 100 rpm, 200 rpm, 300 rpm, 400 rpm or 500 rpm.

[0044] In some embodiments of the present invention, the materials of the nanoparticles include, but are not limited to, silica, molecular sieves, Fe3O4, liposomes, polymer nanoparticles, etc.

[0045] In some embodiments of the present invention, the types of nanoparticles include, but are not limited to, solid spheres, mesoporous structures, hollow mesoporous structures, and hierarchical porous structures.

[0046] In some embodiments of the present invention, the particle size of the nanoparticles is 300-1000 nm, and may also be 400 nm, 500 nm, 600 nm, 700 nm, 800 nm or 900 nm.

[0047] In some embodiments of the present invention, the mass-to-volume ratio of the nanoparticles to the sample is 3–15 mg / mL; it may also be 3–6 mg / mL, 6–9 mg / mL, 9–12 mg / mL, or 12–15 mg / mL.

[0048] In some embodiments of the present invention, the sample is a diluted sample, specifically, the sample to be tested is diluted with a diluent.

[0049] In some embodiments of the present invention, the diluent includes an ionic surfactant and a buffer solution.

[0050] In some embodiments of the present invention, the ionic surfactant includes 3-[(3-cholamidopropyl)dimethylammonium]-1-propanesulfonate (CHAPS), 3-[(3-cholamidopropyl)dimethylammonium]-2-hydroxy-1-propanesulfonate (CHAPSO), hexadecyltrimethylammonium bromide (CTAB), sodium dodecyl sulfate (SDS), sodium sarkosyl dodecyl creatine (sarkosyl) or / and dodecyltrimethylammonium bromide (DTAB); preferably 3-[(3-cholamidopropyl)dimethylammonium]-1-propanesulfonate (CHAPS).

[0051] In some embodiments of the present invention, the concentration of the ionic surfactant is 0.01 to 0.1 w / w, and may also be 0.03 w / w%, 0.04 w / w%, 0.05 w / w%, 0.06 w / w%, or 0.07 w / w.

[0052] In some embodiments of the present invention, the buffer solution includes PBS buffer, Tris buffer, etc.

[0053] In some embodiments of the present invention, the diluent further includes a pH adjuster, such as ammonia, sodium hydroxide, sodium bicarbonate, sodium carbonate, etc.

[0054] In some embodiments of the present invention, the pH of the diluent is 7-11, and may also be 7-8, 8-9, 9-10 or 10-11, preferably 10-11.

[0055] In some embodiments of the present invention, the dilution factor is 2 to 8, wherein the dilution factor is the ratio of the volume after dilution to the volume before dilution; the dilution factor may also be 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5 or 8.

[0056] In some embodiments of the invention, the washing includes first washing the soft corona layer with a buffer solution, and then centrifuging to remove the soft protein crown.

[0057] In some embodiments of the present invention, the number of centrifugations is 2 to 5 times, and may also be 2, 3, 4 or 5 times.

[0058] In some embodiments of the present invention, the conditions for a single centrifugation are 2000-10000g and 5-15min; wherein the centrifugal force can also be 2000-4000g, 4000-6000g, 6000-8000g or 8000-10000g; and the centrifugation time can also be 5-8min, 8-12min or 12-15min.

[0059] In some embodiments of the present invention, the pretreatment further includes a purification step; the purification method includes, but is not limited to, using commercial purification kits; and can be column purification reagents, magnetic bead purification reagents, gel electrophoresis reagents, etc.

[0060] A fifth aspect of the present invention provides a system for assessing the level of red blood cell contamination, comprising the following modules:

[0061] a) Data collection module: Collects the samples to be tested and determines the abundance of the markers in the above-mentioned combination of markers used to assess the degree of red blood cell contamination;

[0062] b) Model calculation module: Calculates the contamination index and / or red blood cell count of the sample to be tested. The formula for calculating the contamination index is as follows:

[0063] Contamination index = (PRPS1 abundance + ADD2 abundance + PA2G4 abundance + OLA1 abundance + ACLY abundance + PIP4K2A abundance + FLOT1 abundance + CFL1 abundance + ADD1 abundance + SNCA abundance + STOM abundance + RAP1B abundance + PPIA abundance + MPP1 abundance + EPB42 abundance + RAN abundance + CD59 abundance + HSPA8 abundance + SLC4A1 abundance + HBD abundance + EPB41 abundance + HBA1 abundance + BLRRB abundance + GAPDH abundance + PGK1 abundance + SPTB abundance + HBB abundance + SPTA1 abundance + CA1 abundance + ANK1 abundance) / total protein abundance;

[0064] The method for calculating the number of red blood cells is as follows: the contamination index is mapped to the actual cell count through a regression model;

[0065] c) Output prediction module: Predicts the contamination status of the sample based on the calculated contamination index and / or red blood cell count of the sample to be tested: the lower the contamination index and / or red blood cell count of the sample to be tested, the lower the degree of contamination; the higher the contamination index and / or red blood cell count of the sample to be tested, the higher the degree of contamination.

[0066] In some embodiments of the present invention, the pollution index is compared with a threshold value. If the index is higher than the threshold value, the pollution level is higher; if the index is lower than the threshold value, the pollution level is lower.

[0067] A sixth aspect of the present invention provides a computing device comprising:

[0068] At least one processing unit; and

[0069] At least one memory coupled to the processing unit and storing a program for execution by the processing unit, which, when executed by the processor, causes the processor to perform the aforementioned assessment of red blood cell contamination levels.

[0070] A seventh aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps described above for assessing the degree of red blood cell contamination.

[0071] An eighth aspect of the present invention provides a method for constructing a model to differentiate between benign pulmonary nodules and lung cancer, comprising:

[0072] 1) Obtain imaging and clinically relevant data from patients with benign pulmonary nodules and lung cancer, and construct a dataset;

[0073] 2) Key features were selected using machine learning methods to construct an algorithm model that distinguishes between benign lung nodules and lung cancer.

[0074] In the model construction method, when acquiring relevant clinical data such as proteomics, the method described in the fourth aspect of this invention can be used to assess the contamination level of red blood cells and then exclude contaminated samples to prevent the screening of contamination-related proteins as biomarkers of disease.

[0075] In some embodiments of the present invention, the machine learning method includes logistic regression, K-nearest neighbors, support vector machine, random forest, gradient boosting, multilayer perceptron, adaptive boosting, or extreme random tree; preferably extreme random tree.

[0076] In some embodiments of the present invention, the key features include: tumor size, CA-125, IGHV4-4, IGLV9-49, TIMM44, GMDS, SELENOF, UGDH, FCGR3A, TNFSF13, PAPLN, IGLV1-51, SLC3A2, DBH, AZU1, ANOS1, COL15A1, HAMP, FGL1, SKIV2L, ATP6AP1, FSCN1, HID1, PLVAP, PDGFD, SPTBN5, IMPAD1, RRN3, VPS18, and IGLC7.

[0077] A ninth aspect of the invention provides a combination of biomarkers for differentiating between benign pulmonary nodules and lung cancer, including tumor size, CA-125, IGHV4-4, IGLV9-49, TIMM44, GMDS, SELENOF, UGDH, FCGR3A, TNFSF13, PAPLN, IGLV1-51, SLC3A2, DBH, AZU1, ANOS1, COL15A1, HAMP, FGL1, SKIV2L, ATP6AP1, FSCN1, HID1, PLVAP, PDGFD, SPTBN5, IMPAD1, RRN3, VPS18, and IGLC7.

[0078] In a tenth aspect, the invention provides the use of a substance for detecting the combination of markers for differentiating benign pulmonary nodules and lung cancer in the preparation of a product for differentiating benign pulmonary nodules and lung cancer.

[0079] In some embodiments of the present invention, the substances used to detect and differentiate the biomarkers for benign pulmonary nodules and lung cancer include substances that can be quantitatively detected by at least one of the following methods: mass spectrometry, chromatography, surface-enhanced Raman spectroscopy, Western blotting, flow cytometry, protein array, immunoprecipitation, and immunoadsorption, for CA-125, IGHV4-4, IGLV9-49, TIMM44, GMDS, SELENOF, UGDH, FCGR3A, TNFSF13, PAPLN, IGLV1-51, SLC3A2, DBH, AZU1, ANOS1, COL15A1, HAMP, FGL1, SKIV2L, ATP6AP1, FSCN1, HID1, PLVAP, PDGFD, SPTBN5, IMPAD1, RRN3, and VPS18; and substances that can be used to detect tumor size by imaging and pathology.

[0080] In some embodiments of the present invention, the product includes reagents, reagent kits, test strips, systems, devices, or chips.

[0081] An eleventh aspect of the present invention provides a system for differentiating between benign pulmonary nodules and lung cancer, comprising:

[0082] An acquisition module is used to acquire the detection result of each marker in the marker combination of the sample to be tested; wherein, the marker combination is the marker combination described in the ninth aspect of the present invention;

[0083] The prediction module is used to input the detection results of all markers into the model described in the eighth aspect of the present invention to obtain the identification result of the sample to be tested.

[0084] In a twelfth aspect, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, can perform the functions of the system described in the eleventh aspect of the present invention.

[0085] To address the common problem of sample contamination in proteomics, this invention discloses for the first time a combination of biomarkers for assessing erythrocyte contamination, comprising the following biomarkers: PRPS1, ADD2, PA2G4, OLA1, ACLY, PIP4K2A, FLOT1, CFL1, ADD1, SNCA, STOM, RAP1B, PPIA, MPP1, EPB42, RAN, CD59, HSPA8, SLC4A1, HBD, EPB41, HBA1, BLVRB, GAPDH, PGK1, SPTB, H BB, SPTA1, CA1, and ANK1; this biomarker combination can accurately identify and exclude red blood cell contamination samples and quantify the number of red blood cells, used to exclude red blood cell contamination samples obtained by enrichment methods. Moreover, this biomarker combination is universal, unaffected by the type of nanoparticle, and has high sensitivity and specificity, and can be used to identify red blood cell contamination samples in various enrichment methods, whether based on or not based on nanoparticles. It is also beneficial for accurately screening biomarkers for disease evaluation through proteomics, and can be used in clinical scenarios such as early cancer screening and personalized medicine.

[0086] Based on the accurate identification and exclusion of red blood cell contamination samples, this application further constructed a model for distinguishing between benign pulmonary nodules and lung cancer using machine learning, with an AUC of up to 0.8, demonstrating high diagnostic efficacy. Attached Figure Description

[0087] Figure 1 The following are presented: (a) Development of a quality control biomarker set for erythrocyte contamination; (b) Screening protocol for erythrocyte-related biomarkers; (c) Number of peptide precursor ions identified in the discovery dataset; (d) Number of proteomes; (e) Abundance distribution of erythrocyte biomarkers; (f) Spearman correlation analysis of 30 erythrocyte biomarkers; (g) Design of biomarker validation experiments; (h) Relationship between the Z-score normalization strength of 30 biomarkers and the proportion of erythrocyte incorporation; (h) Correlation between erythrocyte count and contamination index (RBC: erythrocyte-rich plasma, PPP: erythrocyte-deficient plasma).

[0088] Figure 2 Red blood cell contamination index for the lung cancer cohort. (a) Distribution of red blood cell contamination index in the lung cancer cohort; (b) Diagnostic efficacy in the test set; (c) Importance ranking of the top 30 features selected by machine learning.

[0089] Figure 3 The ROC curves are for the training and test sets in Example 4. Detailed Implementation

[0090] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification.

[0091] Before further describing specific embodiments of the present invention, it should be understood that the scope of protection of the present invention is not limited to the specific embodiments described below; it should also be understood that the terminology used in the embodiments of the present invention is for describing specific embodiments and not for limiting the scope of protection of the present invention. Test methods in the following embodiments that do not specify specific conditions are generally performed under conventional conditions or as recommended by the respective manufacturers.

[0092] When numerical ranges are given in the embodiments, it should be understood that, unless otherwise stated in the present invention, both endpoints of each numerical range and any value between the two endpoints may be selected. Unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art. In addition to the specific methods, apparatus, and materials used in the embodiments, based on the knowledge of the prior art possessed by one of ordinary skill in the art and the description of this invention, any prior art methods, apparatus, and materials similar to or equivalent to those described, apparatus, and materials in the embodiments of this invention may be used to implement the present invention.

[0093] Example 1: Screening and Construction of a Plasma Proteomics Pretreatment Process Based on Nanoparticle Enrichment

[0094] The pretreatment process for plasma proteomics based on nanoparticles includes: plasma sample and nanoparticle binding to form protein corona (soft and hard corona), soft corona removal, protein denaturation and enzymatic digestion, etc.

[0095] The specific steps for using silica nanoparticles to enrich low-abundance proteins in plasma are as follows:

[0096] First, take 15-30 μL of plasma sample and dilute it with 75-100 μL of 1× phosphate buffer (PBS, referred to as buffer 2, pH = 10-11) containing 0.05 w / w% 3-[(3-cholamidopropyl)dimethylamino]-1-propanesulfonic acid and 0.02 v / v% ammonia. Then add 0.3-1.0 mg of solid spherical silica nanoparticles (NP, particle size 300-1000 nm) and incubate in a 30°C constant temperature shaker at 100-500 rpm for 30-120 minutes. After incubation, wash the soft corona layer with buffer 2 diluted to 33 v / v% with water. Repeat the washing process three times by centrifuging at 2500-10000g for 10 minutes to remove the soft corona, discarding the supernatant after each centrifugation and repeatedly washing the precipitate. The hard corona layer was then subjected to a process of protein denaturation, enzymatic digestion, and desalting purification, as follows: First, protein denaturation was achieved using 50 μL of 8M urea and 2M thiourea. Then, under light-protected conditions, Tris(2-carboxyethyl)phosphine (TCEP) and 40 mM iodoacetamide (IAA) were added to a final concentration of 10 mM, and the reaction was carried out for 30–60 minutes to complete reduction and alkylation. Next, the urea concentration was reduced to below 1.2 M using 100 mM ABB dilution buffer, and 0.5 μg–2 μg of trypsin was added at a mass ratio of 1:10–1:25 for overnight enzymatic digestion. Finally, 30–50 μL of 10% trifluoroacetic acid (TFA) was added to terminate the reaction. The peptides generated from the enzymatic digestion were purified using a peptide desalting column and then vacuum dried.

[0097] Finally, after pretreatment, plasma samples were analyzed by 24-minute Astral mass spectrometry (nDIA), consistently identifying 3000–6000 proteomes per sample with a batch coefficient of variation (CV) of less than 12%. The mass spectrometry parameters were as follows: approximately 200–400 ng of peptides were first loaded onto a trap column, followed by separation using a custom-designed analytical column (75 μm inner diameter × 15 cm length, 1.9 μm particle size) via 24-minute LC-MS. The initial LC gradient conditions were: 8% buffer B (Buffer B: 80% acetonitrile containing 0.1% formic acid (v / v); Buffer A: 0.1% formic acid (v / v) dissolved in mass spectrometry-grade ultrapure water), increasing to 10% B within 1.5 minutes, then to 30% B within 16 minutes, and finally to 40% B within 2 minutes. Each run included a 4.3-minute column washing and equilibration step. Eluted peptides were analyzed using an Orbitrap Astral mass spectrometer with the following parameters: FAIMS voltage -42V, full scan resolution 240,000, mass-to-charge ratio scan range 380–980Th; MS / MS scan range was the same as the full scan, using DIA mode (data independent of acquisition), with an isolation window width of 2Da.

[0098] Example 2: Discovery and Validation of Biomarkers for Erythrocyte Contamination

[0099] Based on the established nanoparticle-based plasma protein pretreatment method and the human-approved proteomics spectral library (download path: https: / / www.uniprot.org / ), experiments were designed to establish biomarkers and algorithms for assessing red blood cell contamination of plasma samples.

[0100] First, pure red blood cell (RBC) samples and blood plasma samples without red blood cells (PPP) were collected, according to... Figure 1 The samples were mixed in proportion a to obtain samples containing different levels of red blood cell contamination. After being enriched by nanoparticles, the plasma samples with different red blood cell contaminations will achieve a greater amount of protein identification as the red blood cell count gradually increases.

[0101] The results showed that approximately 3500 proteomes were identified on average in RBC samples, while approximately 1500 proteomes were identified on average in PPP samples. Figure 1 bc).

[0102] Then, proteins with a missing rate of less than 50% in all samples were screened. Mfuzz clustering analysis was used to select the top 100 proteins with the highest correlation to erythrocyte concentration (Spearman r > 0.95). Based on protein abundance, the top 30 proteins were selected as biomarkers for evaluating erythrocyte contamination, including HBA1 (hemoglobin α1), HBB (hemoglobin β), and SPTA1 (spectrin α), among others. The abundance distribution of erythrocyte biomarkers is referenced. Figure 1 d. The median Spearman correlation of these 30 proteins was 0.94. Figure 1 e). The specific information for the 30 biomarkers is as follows: PRPS1, ADD2, PA2G4, OLA1, ACLY, PIP4K2A, FLOT1, CFL1, ADD1, SNCA, STOM, RAP1B, PPIA, MPP1, EPB42, RAN, CD59, HSPA8, SLC4A1, HBD, EPB41, HBA1, BLVRB, GAPDH, PGK1, SPTB, HBB, SPTA1, CA1, and ANK1. The peptide sequences corresponding to these proteins are detailed in the table below.

[0103] Table 1

[0104]

[0105]

[0106] Based on the aforementioned 30 biomarkers, the applicant established a contamination index, which is the ratio of the sum of the abundance of the 30 biomarkers in each sample to the abundance of all proteins in that sample. Specifically, this is the contamination index (erythrocyte contamination index = Σ biomarker intensity / total protein intensity). The contamination index is used to assess the degree of erythrocyte contamination in a sample; the cutoff value is 0.04. A contamination index higher than 0.04 indicates a contaminated sample that needs to be excluded. A higher contamination index indicates more severe erythrocyte contamination. Subsequently, the feasibility of using another dataset to evaluate the 30 selected proteins for erythrocyte contamination was verified.

[0107] The applicant collected pure red blood cell and pure PPP samples from 12 individuals. Each sample was then counted using a red blood cell counter to determine the blood cell distribution and purity of the red blood cells and PPP samples. Red blood cells from every three patients were then mixed, resulting in four groups of pure red blood cell (RBC) and pure PPP samples. Figure 1 f). We performed 10-stage dilutions of red blood cells and PPP samples in the four groups of samples, and then used a red blood cell counter to count the red blood cells in each diluted sample to determine the absolute red blood cell content of each sample.

[0108] Correlation analysis between erythrocyte concentration and the abundance of 30 biomarkers showed good results. Furthermore, spline regression was used to fit the absolute erythrocyte content and contamination, resulting in a positive correlation ratio (R0). 2 =0.94( Figure 1 h) showed a good correlation, indicating that the contamination index can be used to assess the degree of red blood cell contamination in a sample.

[0109] The applicant further quantified the red blood cell count using a contamination index, specifically by using a constructed spline regression model to map the index to the actual cell count. The model construction method is as follows: Based on 10 levels of dilution data from 4 sets of samples in the validation set (each dilution gradient contains 3 biological replicates, for a total of 120 samples), a univariate spline function was used for fitting and modeling. With the contamination index as the independent variable and the actual red blood cell count as the dependent variable, the number of nodes k=3 and the smoothing coefficient s=50 (through multiple cross-validations, this parameter setting maximizes the balance between flexibility and smoothness in the model), the final red blood cell count prediction curve (R²) was obtained. 2=0.94), the model expression is: Red blood cell count = Spline(contamination index), where coeffs[6.67520162,6.34688444,7.72838449,7.53426373] / knots[0.12249392,0.75213502].

[0110] To further validate these biomarkers, erythrocyte and thrombocytopenic plasma (PPP) samples were collected from six patients and mixed into two groups (each group containing samples from three patients). Subsequently, the erythrocyte and PPP samples were serially diluted using a ten-step serial dilution method. Figure 1 f). It can be seen that with gradual dilution, all 30 proteins used to evaluate erythrocyte contamination showed a decreasing trend (f). Figure 1 g).

[0111] Example 3: Universality of Red Blood Cell Contamination Biomarkers

[0112] To verify the applicability of the identified biomarkers to samples treated with other types of nanoparticles, two high-frequency nanoparticles reported in the literature were selected: NaY-type zeolite (particle size 300–700 nm) and silanol-functionalized iron oxide (particle size 400–700 nm). PPP and RBC samples were prepared by collecting donor blood, and 100 μL of each was mixed and serially diluted in 10 steps. The diluted plasma samples were then treated with NaY-type zeolite and silanol-functionalized iron oxide nanoparticles respectively according to the steps in Example 1, and nDIA analysis was performed.

[0113] The results showed that with increasing erythrocyte concentration, the number of peptide precursors and proteomes identified in both nanoparticle-treated samples significantly increased, with over 5500 proteomes detected in a single injection. Thirty erythrocyte-related biomarkers showed a high correlation with the degree of erythrocyte contamination in both NaY-type zeolite-treated samples (median spearman correlation between the abundance of 30 proteins evaluating erythrocyte contamination and the erythrocyte contamination count: 0.92, range 0.92–0.96) and NP81-treated samples (median: 0.95, range 0.89–0.94).

[0114] This also demonstrates that the method for assessing the degree of red blood cell contamination in this invention is compatible with various nanoparticles such as molecular sieves and Fe3O4, providing a unified framework for cross-laboratory data standardization.

[0115] Example 4 Application Case

[0116] To evaluate the effectiveness of nanoparticle-based techniques for enriching low-abundance plasma proteins and assessing erythrocyte contamination, this study included 193 participants, including 42 patients with benign pulmonary nodules and 151 patients with early-stage malignant tumors.

[0117] 1. Criteria for determining benign lung nodules

[0118] 1.1 Benign pulmonary nodules

[0119] a. The pathology report diagnosed it as a benign pulmonary nodule;

[0120] b. Pathology report is unavailable, imaging results show small nodules, and imaging follow-up data for more than 1 year show that the nodules are stable (volume change <25%) or the volume doubling time is greater than 400 days.

[0121] 1.2 Lung Cancer: Pathology report diagnosed lung cancer (both small cell carcinoma and non-small cell carcinoma were included).

[0122] a. Pathologically confirmed as primary lung cancer;

[0123] b. Stages: Stages I-III;

[0124] c. Histological type is clearly recorded.

[0125] 2. Inlet and outlet standards

[0126] 2.1 Sample Inclusion Criteria

[0127] a. Age between 25 and 80 years old, gender not limited;

[0128] b. CT screening results showed lung nodules ranging in size from 5 to 30 mm;

[0129] c. Has not undergone any treatment related to pulmonary nodules / lung cancer (including surgery, chemotherapy, radiotherapy, targeted therapy, immunotherapy, interventional therapy, etc.);

[0130] d. Complete clinical and imaging data;

[0131] e. Low-dose spiral CT reports from the past 3 months are available;

[0132] f. Voluntarily sign the informed consent form.

[0133] 2.2 Sample Exclusion Criteria

[0134] a. History of cancer;

[0135] b. Clinically uncontrolled active infections, such as acute pneumonia, pulmonary tuberculosis, etc.

[0136] c. Received any lung nodule-related treatments such as antibiotics and hormones within the past 4 weeks;

[0137] d. Comorbid with other tumors and serious diseases of the heart, liver, kidneys, brain, blood, etc.;

[0138] e. Participated in other clinical trials within the past 3 months;

[0139] f. Comorbidities affecting protein levels, such as liver and kidney dysfunction or hypoproteinemia;

[0140] g. During pregnancy or lactation.

[0141] Such cases often present unclear diagnoses on routine CT imaging and require surgical resection for definitive diagnosis. All plasma samples were collected using EDTA vacuum blood collection tubes; some patients underwent secondary sampling to evaluate the stability of the pretreatment process; samples from repeated processing were not included in subsequent modeling. Blood was centrifuged (3000g, 4℃, 15 minutes) to collect plasma, which was then stored at 80℃. Peptide samples were then obtained using the nanoparticle-based enrichment process for low-abundance plasma proteins described in Example 1, and analyzed using an Astral instrument for quantitative analysis of peptides and proteins. Proteomics analysis showed an average of 4413 proteomes identified per plasma sample, with stable coefficients of variation in both biological and technical replicates. The contamination level was assessed using the aforementioned methods, and detectable contamination indicators were found in three samples. Figure 2 a) These contaminated samples were excluded from subsequent analyses.

[0142] Subjects were divided into training and testing sets in a 7:3 ratio. Eight machine learning algorithms (logistic regression, K-nearest neighbors classifier, support vector machine (SVC), random forest classifier, gradient boosting classifier, multilayer perceptron (MLPC classifier), adaptive boosting classifier, and extra trees classifier) ​​were used to differentiate between benign lung nodules and lung cancer patients. After feature selection, the models were ranked according to their importance. The ROC curves of each model on the training and testing sets were referenced. Figure 3The modeling results based on the first 30 important proteins or clinical features show that the extreme random tree model performs best in terms of F1 score, accuracy, precision, recall, and AUC. This model selected tumor size, CA-125, and 28 proteins as key features. These 28 proteins include IGHV4-4, IGLV9-49, TIMM44, GMDS, SELENOF, UGDH, FCGR3A, TNFSF13, PAPLN, IGLV1-51, SLC3A2, DBH, AZU1, ANOS1, COL15A1, HAMP, FGL1, SKIV2L, ATP6AP1, FSCN1, HID1, PLVAP, PDGFD, SPTBN5, IMPAD1, RRN3, VPS18, and IGLC7. This confirms that the extreme random tree model constructed based on these key features can effectively distinguish between benign pulmonary nodules and lung cancer patients, achieving a classification accuracy of 82% (AUC = 0.80) in the test set. Figure 2 bc and Figure 3 It has excellent diagnostic efficacy.

[0143] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.

Claims

1. A marker combination for evaluating the degree of red blood cell contamination, the marker combination is SLC4A1, ADD2, PA2G4, SPTB, SPTA1, ADD1, HBA1, FLOT1, CA1, ANK1, PGK1, ACLY, EPB42, CD59, BLVRB, SNCA, STOM, PPIA, HSPA8, EPB41, MPP1, GAPDH, CFL1, RAP1B, PIP4K2A, HBB, HBD, OLA1, RAN and PRPS1; the peptide sequences of the markers in the marker combination are shown in SEQ ID NO. 1-30 respectively.

2. Use of the marker combination of claim 1 in at least one of the following, 1) evaluating the degree of red blood cell contamination; 2) preparing a product for evaluating the degree of red blood cell contamination; 3) calculating the number of red blood cells; 4) preparing a product for calculating the number of red blood cells.

3. Use according to claim 2, characterized in that, The marker combination is detected by at least one of mass spectrometry, chromatography, surface-enhanced Raman spectroscopy, Western blotting, flow cytometry, protein array, immunoprecipitation, and immunoadsorption.

4. A method for evaluating the degree of red blood cell contamination, comprising the following steps: 1) collecting the sample to be tested, and determining the abundance of the markers in the marker combination of claim 1; 2) calculating the contamination index and / or the number of red blood cells of the sample to be tested, the formula for calculating the contamination index is as follows: Contamination index = (PRPS1 abundance + ADD2 abundance + PA2G4 abundance + OLA1 abundance + ACLY abundance + PIP4K2A abundance + FLOT1 abundance + CFL1 abundance + ADD1 abundance + SNCA abundance + STOM abundance + RAP1B abundance + PPIA abundance + MPP1 abundance + EPB42 abundance + RAN abundance + CD59 abundance + HSPA8 abundance + SLC4A1 abundance + HBD abundance + EPB41 abundance + HBA1 abundance + BLVRB abundance + GAPDH abundance + PGK1 abundance + SPTB abundance + HBB abundance + SPTA1 abundance + CA1 abundance + ANK1 abundance) / total protein abundance; The method for calculating the number of red blood cells is to map the contamination index to the actual cell count by a regression model; 3) predicting the contamination of the sample according to the calculated contamination index and / or the number of red blood cells of the sample to be tested: the lower the contamination index and the number of red blood cells of the sample to be tested, the lower the degree of red blood cell contamination of the sample to be tested; The higher the contamination index and / or the number of red blood cells of the sample to be tested, the higher the degree of red blood cell contamination of the sample to be tested.

5. The method of claim 4, wherein, The sample is at least one of plasma, serum, tissue and / or cell sample.

6. The method of claim 4, wherein, The sample includes an enriched sample obtained after pretreatment, and the pretreatment step includes or does not include a step of incubation with nanoparticles.

7. The method of claim 6, wherein, The pretreatment specifically includes the following steps: incubating the sample to be tested with nanoparticles, and then washing, denaturing and enzymatic digestion after incubation.

8. A system for evaluating the degree of red blood cell contamination, comprising the following modules: a) a data collection module: collecting a sample to be tested, and determining the abundance of markers in the marker combination of claim 1; b) a model calculation module: calculating the contamination index and / or the red blood cell count of the sample to be tested, wherein the formula for calculating the contamination index is as follows: Contamination index = (PRPS1 abundance + ADD2 abundance + PA2G4 abundance + OLA1 abundance + ACLY abundance + PIP4K2A abundance + FLOT1 abundance + CFL1 abundance + ADD1 abundance + SNCA abundance + STOM abundance + RAP1B abundance + PPIA abundance + MPP1 abundance + EPB42 abundance + RAN abundance + CD59 abundance + HSPA8 abundance + SLC4A1 abundance + HBD abundance + EPB41 abundance + HBA1 abundance + BLVRB abundance + GAPDH abundance + PGK1 abundance + SPTB abundance + HBB abundance + SPTA1 abundance + CA1 abundance + ANK1 abundance) / total protein abundance; The method for calculating the red blood cell count is to map the contamination index to the actual cell count through a regression model; c) an output prediction module: predicting the contamination of the sample according to the calculated contamination index and / or red blood cell count of the sample to be tested: the lower the contamination index and / or red blood cell count of the sample to be tested, the lower the degree of contamination; the higher the contamination index and / or red blood cell count of the sample to be tested, the higher the degree of contamination. 9.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-8. The processor executes the computer program to realize the steps of the method of claim 4.

10. A model construction method for identifying benign lung nodules and lung cancer, comprising: 1) obtaining imaging data and clinical data of patients with benign lung nodules and lung cancer, and constructing a data set; 2) screening key features by a machine learning method, and constructing an algorithm model; In step 1), when obtaining proteomic clinical data, the method of claim 4 is used to evaluate the contamination level of red blood cells and exclude red blood cell contamination samples.

11. The method of claim 10, wherein, The machine learning method includes logistic regression, K-nearest neighbors, support vector machines, random forests, gradient boosting, multi-layer perceptron, adaptive boosting, or extreme random trees.

12. The method of claim 10, wherein, The key features include tumor size, CA-125, IGHV4-4, IGLV9-49, TIMM44, GMDS, SELENOF, UGDH, FCGR3A, TNFSF13, PAPLN, IGLV1-51, SLC3A2, DBH, AZU1, ANOS1, COL15A1, HAMP, FGL1, SKIV2L, ATP6AP1, FSCN1, HID1, PLVAP, PDGFD, SPTBN5, IMPAD1, RRN3, VPS18, and IGLC7.

Citation Information

Patent Citations

  • Measuring method of complement sensitized erythrocyte

    JP1995005171A

  • Laboratory diagnostic technique for paroxysmal nocturnal haemoglobinuria

    RU2574968C1