Application of biomarker in evaluation of platelet contamination degree
The degree of platelet contamination is evaluated through marker combination and regression model, and the problem of plasma proteomic analysis error caused by platelet contamination in the prior art is solved, and the accurate identification and quantification of platelet contamination is achieved, which improves the accuracy and clinical application of plasma proteomic analysis.
Patent Information
- Application Number
- CN202510492911.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-18
AI Technical Summary
The prior art cannot accurately evaluate the degree of platelet contamination, resulting in large errors in plasma proteomic analysis results, high false positive rate, and decreased specificity, affecting the accuracy of disease marker screening.
The marker combination includes 11 proteins such as IDH2, LIMS1, MLEC, and DIAPH1. The pollution index and platelet number were calculated by quantitative detection through mass spectrometry, chromatography and other methods, and the degree of pollution was evaluated in combination with the regression model.
Accurate identification and quantification of platelet contamination is achieved, the accuracy of plasma proteomics analysis is improved, the false positive rate is reduced, and it is helpful for the large-scale application of clinical scenarios.
Smart Images

Figure BDA0005366085690000081 
Figure BDA0005366085690000091 
Figure HDA0005366085700000011
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biomedical technologies, and particularly to the application of biomarkers in the assessment of platelet contamination degree. Background Art
[0002] As a core technology for non-invasive disease biomarker discovery, circulating blood proteomics has significantly improved the detection depth of plasma proteins in recent years driven by mass spectrometry (MS) and nanoparticle (NP) enrichment methods. In previous literature reports, approximately 54% of blood proteomics studies had platelet residue contamination, and the contaminants falsely increased the number of protein identifications through non-specific adsorption (an average increase of 38%).
[0003] Existing methods for evaluating platelet contamination of plasma samples rely on subjective experience or a single biomarker (such as PF4) to assess contamination, and cannot quantify the contamination degree or correct data deviation, resulting in the disconnection between the analysis results and the true biological effects. Moreover, it has been reported in the literature that there are more proteins in platelets than in plasma. It can be expected that samples with platelet contamination processed by the nanoparticle enrichment method may be more likely to cause platelet contamination of plasma proteins.
[0004] In terms of clinical translation, biomarkers screened by existing NP technologies for evaluating diseases often have platelet-specific proteins. If proteins related to contamination are screened out as disease biomarkers, it will lead to an increase in incorrect results, specifically manifested as an increase in false positive rate, a decrease in specificity, and failure of clinical verification. For example, the platelet contamination biomarker PF4 (Platelet Factor 4) may be misjudged as a potential biomarker for cardiovascular diseases, while in fact its abundance change only reflects the degree of platelet activation during sample processing, resulting in incorrect association of disease mechanisms. Therefore, providing a biomarker that can accurately assess platelet contamination in plasma proteomics samples is an urgent problem to be solved in this field. Summary of the Invention
[0005] In view of the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a biomarker combination for evaluating the degree of platelet contamination to solve the problems in the prior art.
[0006] To achieve the above purpose and other related purposes, the present invention is obtained through the following technical solutions.
[0007] In the first aspect of the present invention, there is provided a marker combination for evaluating the degree of platelet contamination, and the marker combination includes IDH2, LIMS1, MLEC, DIAPH1, GNAI2, ATP5F1B, RAB14, TBXAS1, JAM3, VDAC1, ATP2A3, ARHGAP6, VDAC3, ESAM, SLC25A5, ITGA6, PDLIM7, PF4V1, HSP90B1, HSPD1, RDH11, ATP2A2, VDAC2, RAB35, TREML1, GP5, PRDX3, ITGB1, MDH2, and GNB1.
[0008] In the second aspect of the present invention, there is provided the use of a reagent for quantitatively detecting the marker combination in a sample in at least one of the following, and the combination includes IDH2, LIMS1, MLEC, DIAPH1, GNAI2, ATP5F1B, RAB14, TBXAS1, JAM3, VDAC1, ATP2A3, ARHGAP6, VDAC3, ESAM, SLC25A5, ITGA6, PDLIM7, PF4V1, HSP90B1, HSPD1, RDH11, ATP2A2, VDAC2, RAB35, TREML1, GP5, PRDX3, ITGB1, MDH2, and GNB1;
[0009] (1) Evaluating the degree of platelet contamination;
[0010] (2) Preparing a product for evaluating the degree of platelet contamination;
[0011] (3) Calculating the number of platelets;
[0012] (4) Preparing a product for calculating the number of platelets.
[0013] In some embodiments of the present invention, the reagent performs quantitative detection by at least one method of mass spectrometry, chromatography, surface-enhanced Raman spectroscopy, Western blotting, flow cytometry, protein array, immunoprecipitation, immunoadsorption.
[0014] In some embodiments of the present invention, the reagent is selected from one or more of a substance specific to the marker, a probe specific to the marker, and a protein chip.
[0015] In some embodiments of the present invention, the substance specific to the marker includes an antibody, a ligand protein, a polypeptide, a non-protein compound, and / or a nucleic acid aptamer.
[0016] In some embodiments of the present invention, the sample includes at least one of plasma, serum, tissue, and cells.
[0017] In some embodiments of the present invention, the sample includes an enriched sample obtained through pretreatment, and the pretreatment includes a pretreatment method involving or not involving nanoparticles.
[0018] In some embodiments of the present invention, the product includes a reagent, a kit, a test strip, a chip, a device or a detection system.
[0019] In a third aspect of the present invention, there is provided a product, which includes a reagent for detecting the above-mentioned marker combination for evaluating the degree of platelet contamination.
[0020] In some embodiments of the present invention, the reagent is quantitatively detected by at least one of mass spectrometry, chromatography, surface-enhanced Raman spectroscopy, Western blotting, flow cytometry, protein array, immunoprecipitation, immunosorption.
[0021] In some embodiments of the present invention, the reagent is selected from one or more of a substance specific to the marker, a probe specific to the marker, and a protein chip.
[0022] In some embodiments of the present invention, the substance specific to the marker includes an antibody, a ligand protein, a polypeptide, a non-protein compound, and / or a nucleic acid aptamer.
[0023] In some embodiments of the present invention, the product includes a reagent, a kit, a test strip, a chip, a device or a detection system.
[0024] In some embodiments of the present invention, the product further includes a reagent for assisting in detecting protein abundance.
[0025] In some embodiments of the present invention, the test sample includes at least one of plasma, serum, tissue, and cells.
[0026] In some embodiments of the present invention, the sample includes an enriched sample obtained through pretreatment, and the pretreatment includes a pretreatment method involving or not involving nanoparticles.
[0027] In a fourth aspect of the present invention, there is provided a method for evaluating the degree of platelet contamination, including the following steps:
[0028] 1) Collect a test sample and measure the abundance of the markers in the above-mentioned marker combination for evaluating the degree of platelet contamination;
[0029] 2) Calculate the contamination index and / or the number of platelets of the test sample, and the calculation formula of the contamination index is as follows:
[0030] Contamination index = (abundance of IDH2 + abundance of LIMS1 + abundance of MLEC + abundance of DIAPH1 + abundance of GNAI2
[0031] + Abundance of ATP5F1B + Abundance of RAB14 + Abundance of TBXAS1 + Abundance of JAM3 + Abundance of VDAC1 + Abundance of ATP2A3 + Abundance of ARHGAP6 + Abundance of VDAC3 + Abundance of ESAM + Abundance of SLC25A5 + Abundance of ITGA6 + Abundance of PDLIM7 + Abundance of PF4V1 + Abundance of HSP90B1 + Abundance of HSPD1 + Abundance of RDH11 + Abundance of ATP2A2 + Abundance of VDAC2 + Abundance of RAB35 + Abundance of TREML1 + Abundance of GP5 + Abundance of PRDX3 + Abundance of ITGB1 + Abundance of MDH2
[0032] + Abundance of GNB1) / Total protein abundance;
[0033] The calculation method of the platelet count is: mapping the contamination index to the actual cell count through a regression model;
[0034] 3) Predict the contamination situation of the sample based on the calculated contamination index and / or platelet count of the sample to be tested: the lower the contamination index and / or platelet count of the sample to be tested, the lower the contamination degree; the higher the contamination index and / or platelet count of the sample to be tested, the higher the contamination degree.
[0035] In some embodiments of the present invention, the contamination index is compared with a defined value. If it is higher than the defined value, the contamination degree is higher; if it is lower than the defined value, the contamination degree is lower.
[0036] In some embodiments of the present invention, the defined value is 0.006. If it is higher than 0.006, it is a platelet - contaminated sample and needs to be excluded when performing subsequent proteomic analysis.
[0037] In some embodiments of the present invention, the regression model includes: Polynomial Regression, Support Vector Regression, Regression Tree, and Multivariate Adaptive Regression Splines or Spline Regression; preferably Spline Regression.
[0038] In some embodiments of the present invention, the sample includes at least one of plasma, serum, tissue, and cells.
[0039] In some embodiments of the present invention, the sample includes an enriched sample obtained through pretreatment, and the pretreatment includes a pretreatment method with or without nanoparticles.
[0040] In some embodiments of the present invention, the pre-treatment specifically includes the following steps:
[0041] Incubate the sample to be tested with nanoparticles. After incubation, wash, denature, and enzymatically digest. In the pre-treatment, incubate the plasma sample with nanoparticles. The plasma sample and nanoparticles combine to form a protein corona (including a soft protein corona and a hard protein corona). Remove the soft protein corona by washing and retain the hard protein corona with high affinity. Subsequently, obtain a sample that can be used for proteomic analysis through denaturation and enzymatic digestion.
[0042] In some embodiments of the present invention, the temperature of the incubation is 20 - 40 °C; it can also be 20 - 25 °C, 25 - 30 °C, 30 - 35 °C, or 35 - 40 °C, and it can also be 26 °C, 27 °C, 28 °C, 29 °C, 30 °C, 31 °C, 32 °C, 33 °C, or 34 °C.
[0043] In some embodiments of the present invention, the incubation time is 30 - 120 min; it can also be 30 - 120 min, 30 - 120 min, 30 - 120 min, 30 - 120 min, 30 - 120 min, 30 - 120 min, or 30 - 120 min.
[0044] In some embodiments of the present invention, the rotation speed of the incubation is 100 - 500 revolutions per minute; it can also be 100 revolutions per minute, 200 revolutions per minute, 300 revolutions per minute, 400 revolutions per minute, or 500 revolutions per minute.
[0045] In some embodiments of the present invention, the material of the nanoparticles includes but is not limited to silica, molecular sieve, Fe3O4, liposome, polymer nanoparticles, etc.
[0046] In some embodiments of the present invention, the types of the nanoparticles include but are not limited to solid spherical, mesoporous structure, hollow mesoporous, and hierarchical porous structure.
[0047] In some embodiments of the present invention, the particle size of the nanoparticles is 300 nm - 1000 nm; it can also be 400 nm, 500 nm, 600 nm, 700 nm, 800 nm, 900 nm, or 1000 nm.
[0048] In some embodiments of the present invention, the mass-volume ratio of the nanoparticles to the sample to be tested is 3 - 15 mg / mL; it can also be 3 - 6 mg / mL, 6 - 9 mg / mL, 9 - 12 mg / mL, or 12 - 15 mg / mL.
[0049] In some embodiments of the present invention, the sample to be tested is a diluted sample, specifically diluted by a diluent.
[0050] In some embodiments of the present invention, the diluent comprises an ionic surfactant and a buffer solution.
[0051] In some embodiments of the present invention, the ionic surfactant comprises 3-[(3-cholamidopropyl)dimethylammonio]-1-propanesulfonate (CHAPS), 3-[(3-cholamidopropyl)dimethylammonio]-2-hydroxy-1-propanesulfonate (CHAPSO), cetyltrimethylammonium bromide (CTAB), sodium dodecyl sulfate (SDS), sodium lauroyl sarcosinate (sarkosyl) or / and dodecyltrimethylammonium bromide (DTAB); preferably 3-[(3-cholamidopropyl)dimethylammonio]-1-propanesulfonate (CHAPS).
[0052] In some embodiments of the present invention, the concentration of the ionic surfactant is 0.01 - 0.1 w / w%, and can also be 0.03 w / w%, 0.04 w / w%, 0.05 w / w%, 0.06 w / w% or 0.07 w / w%.
[0053] In some embodiments of the present invention, the buffer solution comprises PBS buffer solution, Tris buffer solution, etc.
[0054] In some embodiments of the present invention, the diluent further comprises a pH regulator, such as ammonia water, sodium hydroxide, sodium bicarbonate, sodium carbonate, etc., for adjusting the pH value of the diluent.
[0055] In some embodiments of the present invention, the pH value of the diluent is 7 - 11, and can also be 7 - 8, 8 - 9, 9 - 10 or 10 - 11, preferably 10 - 11.
[0056] In some embodiments of the present invention, the dilution factor is 2 - 8, where the dilution factor is the ratio of the volume after dilution to the volume before dilution; the dilution factor can also be 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5 or 8.
[0057] In some embodiments of the present invention, the washing includes first cleaning the soft corona layer with a buffer solution, and then centrifuging and washing to remove the soft protein corona.
[0058] In some embodiments of the present invention, the number of centrifugations is 2 - 5.
[0059] In some embodiments of the present invention, the conditions for single centrifugation are 2000 - 10000g, 5 - 15 min; wherein the centrifugal force can also be 2000 - 4000g, 4000 - 6000g, 6000 - 8000g, or 8000 - 10000g; and the centrifugation time can also be 5 - 8 min, 8 - 12 min, or 12 - 15 min.
[0060] In some embodiments of the present invention, the pretreatment further includes a purification step; the purification method includes but is not limited to methods such as using commercial purification kits; it can be column purification reagents, magnetic bead purification reagents, gel electrophoresis reagents, etc.
[0061] In a fifth aspect of the present invention, a system for evaluating the degree of platelet contamination is provided, including the following modules:
[0062] a) Data collection module: Collect a sample to be tested and measure the abundances of the markers in the above - mentioned marker combination for evaluating the degree of platelet contamination.
[0063] b) Model calculation module: Calculate the contamination index and / or the number of platelets of the sample to be tested. The calculation formula for the contamination index is as follows:
[0064] Contamination index = (abundance of IDH2 + abundance of LIMS1 + abundance of MLEC + abundance of DIAPH1 + abundance of GNAI2 + abundance of ATP5F1B + abundance of RAB14 + abundance of TBXAS1 + abundance of JAM3 + abundance of VDAC1 + abundance of ATP2A3 + abundance of ARHGAP6 + abundance of VDAC3 + abundance of ESAM + abundance of SLC25A5 + abundance of ITGA6 + abundance of PDLIM7 + abundance of PF4V1 + abundance of HSP90B1 + abundance of HSPD1 + abundance of RDH11 + abundance of ATP2A2 + abundance of VDAC2 + abundance of RAB35 + abundance of TREML1 + abundance of GP5 + abundance of PRDX3 + abundance of ITGB1 + abundance of MDH2 + abundance of GNB1) / total protein abundance;
[0065] The calculation method for the number of platelets is: mapping the contamination index to the actual cell count through a regression model.
[0066] c) Output prediction module: Predict the contamination situation of the sample based on the calculated contamination index and / or the number of platelets of the sample to be tested: the lower the contamination index and / or the number of platelets of the sample to be tested, the lower the degree of contamination; the higher the contamination index and / or the number of platelets of the sample to be tested, the higher the degree of contamination.
[0067] In some embodiments of the present invention, the contamination index is compared with a defined value. If it is higher than the defined value, a higher degree of contamination is output; if it is lower than the defined value, a lower degree of contamination is output.
[0068] In some embodiments of the present invention, the defined value is 0.006. If it is higher than 0.006, it is a platelet - contaminated sample, which needs to be excluded when performing subsequent proteomic analysis.
[0069] In some embodiments of the present invention, the regression model includes: Polynomial Regression, Support Vector Regression, Regression Tree, and Multivariate Adaptive Regression Splines or Spline Regression; preferably Spline Regression.
[0070] In a sixth aspect of the present invention, there is provided a computing device, comprising:
[0071] At least one processing unit; and
[0072] At least one memory, the memory being coupled to the processing unit and storing a program for execution by the processing unit. When the program is executed by the processor, the processor realizes the evaluation of the degree of platelet contamination.
[0073] In a seventh aspect of the present invention, there is provided a computer - readable storage medium storing a computer program, which when executed by a processor implements the steps of the above - mentioned method
[0074] It includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it realizes the steps of the above - mentioned method.
[0075] Advantageous effects:
[0076] In view of the problem of platelet contamination in samples that often occurs in proteomics, the present invention first discloses a marker combination for evaluating platelet contamination, comprising the following markers: IDH2, LIMS1, MLEC, DIAPH1, GNAI2, ATP5F1B, RAB14, TBXAS1, JAM3, VDAC1, ATP2A3, ARHGAP6, VDAC3, ESAM, SLC25A5, ITGA6, PDLIM7, PF4V1, HSP90B1, HSPD1, RDH11, ATP2A2, VDAC2, RAB35, TREML1, GP5, PRDX3, ITGB1, MDH2, and GNB1; based on this marker combination, the platelet contamination index in the sample can be accurately identified and the platelet count can be quantified, which is used to exclude platelet-contaminated samples obtained by pretreatment in methods such as proteomics. Moreover, this marker combination has universality and is not affected by the type of nanoparticles in the pretreatment method based on nanoparticles. It has high sensitivity and specificity and can be used for the identification of platelet-contaminated samples in samples obtained by various enrichment treatment methods; it is also conducive to the large-scale implementation of plasma proteomics in clinical scenarios such as early cancer screening and personalized medicine. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] Figure 1 It is the development process of the quality control marker group for platelet contamination. (a) Platelet-related marker screening scheme; (b) The number of peptide precursors identified in the discovery dataset; (c) The number of proteomes; (d) Spearman correlation analysis of platelet markers; (e) Marker validation experimental design; (f) Correlation between platelet count and contamination index; (g) Marker application scenario; (h) Calculation of contamination index for PRP and PPP samples (PRP: platelet-rich plasma, PPP: platelet-poor plasma).
[0078] Figure 2 It is a correlation statistical chart of 30 proteins used to evaluate the degree of blood cell contamination during pretreatment with NaY zeolite (a) and iron oxide nanoparticles (b).
[0079] Figure 3 It is the platelet contamination index of the lung cancer cohort. DETAILED DESCRIPTION OF THE INVENTION
[0080] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification.
[0081] Before further describing the specific embodiments of the present invention, it should be understood that the protection scope of the present invention is not limited to the specific embodiments described below; it should also be understood that the terms used in the embodiments of the present invention are for describing specific embodiments, rather than limiting the protection scope of the present invention. The test methods without specific conditions noted in the following embodiments are generally carried out under conventional conditions or according to the conditions recommended by each manufacturer.
[0082] When the embodiments give a numerical range, it should be understood that unless otherwise specified in the present invention, both endpoints of each numerical range and any value between the two endpoints can be selected. Unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those skilled in the art of this technology. In addition to the specific methods, equipment, and materials used in the embodiments, according to the knowledge of those skilled in the art of this technology and the description of the present invention, any methods, equipment, and materials of the prior art similar or equivalent to those described in the embodiments of the present invention can also be used to implement the present invention.
[0083] Example 1 Screening and construction of a pretreatment process for nanoparticle-based enrichment of plasma proteome
[0084] The pretreatment process for nanoparticle-based plasma proteome includes: the combination of plasma samples and nanoparticles to form a protein corona (soft protein corona and hard protein corona), the removal of the soft protein corona, protein denaturation, and enzymatic digestion, etc.
[0085] In this example, silica nanoparticles are used to enrich low-abundance proteins in plasma. The specific steps are as follows:
[0086] First, take 15 - 30 μL of plasma samples and dilute them with 75 - 100 μL of 1× phosphate buffer (PBS, referred to as buffer 2, pH 10 - 11) containing 0.05 w / w% 3-[(3-cholamidopropyl)dimethylammonio]-1-propanesulfonate and 0.02 v / v% ammonia water. Subsequently, add 0.3 - 1.0 mg of solid spherical silica nanoparticles (NP, particle size 300 - 1000 nm) solution and incubate in a constant temperature shaker at 30 °C at a rotation speed of 100 - 500 revolutions per minute for 30 - 120 minutes. After incubation, wash the soft corona layer with buffer 2 diluted to 33 v / v% concentration, then centrifuge at 2500 - 10000 g for 10 minutes. After discarding the supernatant, add the diluted buffer 2 to the precipitate and centrifuge again. Repeat the centrifugation and washing three times to remove the soft protein corona.
[0087] Then, perform protein denaturation, enzymatic digestion, and desalting and purification processes on the enriched samples. The specific processes are as follows: First, use 50 μL of 8 M urea and 2 M thiourea to achieve protein denaturation. Subsequently, add Tris(2-carboxyethyl)phosphine (TCEP) at a final concentration of 10 mM and iodoacetamide (IAA) at 40 mM under light protection conditions, and react for 30 - 60 minutes to complete reduction and alkylation. Then, dilute the urea concentration to below 1.2 M with 100 mM ABB diluent, and add 0.5 μg - 2 μg of trypsin according to a mass ratio of 1:10 - 1:25 for overnight enzymatic digestion. Finally, add 30 - 50 μL of 10% trifluoroacetic acid (TFA) to terminate the reaction. The peptide segments generated by enzymatic digestion are purified using a peptide desalting column and subjected to vacuum drying treatment.
[0088] Finally, the plasma samples processed through the pretreatment process are subjected to 24-minute Astral mass spectrometry nDIA analysis. The parameters of the mass spectrometry analysis are as follows: First, load approximately 200 - 400 ng of peptide segments onto the capture column, and then use an analytical column (inner diameter 75 μm × length 15 cm, particle size 1.9 μm, customized) for separation by a 24-minute LC-MS method. The initial liquid phase gradient condition is 8% buffer B (buffer B: 80% acetonitrile containing 0.1% formic acid (v / v); buffer A: 0.1% formic acid (v / v) dissolved in mass spectrometry-grade ultrapure water), rising to 10% B within 1.5 minutes, then rising to 30% B within 16 minutes, and finally rising to 40% B within 2 minutes. Each run includes a 4.3-minute column cleaning and equilibration step before. The eluted peptide segments are analyzed by an Orbitrap Astral mass spectrometer, and the parameter settings are: FAIMS voltage -42 V, full scan resolution 240,000, mass-to-charge ratio scan range 380 - 980 Th; the MS / MS scan range is the same as the full scan, using the DIA mode (data-independent acquisition), and the isolation window width is 2 Da.
[0089] The results show that 3000 - 6000 proteomes can be stably identified per sample, and the batch coefficient of variation (CV) is less than 12%.
[0090] Example 2 Discovery and Validation of Platelet Contamination Biomarkers
[0091] Based on the nanoparticle-based plasma protein pretreatment method established above, combined with the plasma proteomics spectral library, design experiments to establish biomarkers and algorithms for evaluating platelet contamination of plasma samples.
[0092] First, collect platelet-rich plasma (PRP) and platelet-poor plasma (PPP) samples, and mix them in a certain proportion to obtain samples with different degrees of platelet contamination ( Figure 1a), First, after nanoparticle enrichment, different plasma samples contaminated with platelets were annotated for peptides and proteins using common proteomics software in combination with a spectral library (https: / / www.uniprot.org / ). The results showed that plasma samples with different degrees of platelet contamination achieved a greater number of protein identifications as the platelet level gradually increased.
[0093] The results showed that the PRP samples identified an average of 4,580 proteomes, and the PPP samples identified an average of 2,492 proteomes ( Figure 1 b - c).
[0094] Then, proteins with a missing rate of less than 50% in all samples were screened. The top 100 proteins with the highest correlation related to platelet concentration (Spearman r > 0.95) were screened by Mfuzz clustering analysis. Then, based on the protein abundance, the 30 proteins with the highest abundance were selected as markers for evaluating platelet contamination, including IDH2, LIMS1, MLEC, DIAPH1, GNAI2, ATP5F1B, RAB14, TBXAS1, JAM3, VDAC1, ATP2A3, ARHGAP6, VDAC3, ESAM, SLC25A5, ITGA6, PDLIM7, PF4V1, HSP90B1, HSPD1, RDH11, ATP2A2, VDAC2, RAB35, TREML1, GP5, PRDX3, ITGB1, MDH2, GNB1. The median of the spearman correlation of these 30 proteins was 0.94 ( Figure 1 d). The peptide sequences corresponding to these proteins are shown in the following table.
[0095] Table 1
[0096]
[0097]
[0098] The present invention establishes a contamination index indicator, which is the ratio of the sum of the abundances of 30 biomarkers in each sample to the abundances of all proteins in this sample. Specifically, the contamination index (platelet contamination index = Σ biomarker abundances / total protein abundances). The platelet contamination degree of this sample is evaluated through the contamination index. The defined value is 0.006. Samples with a contamination index higher than 0.006 can be determined as contaminated samples and need to be excluded. The higher the contamination index, the more serious the platelet contamination situation.
[0099] Immediately afterwards, the feasibility of the 30 proteins selected for evaluating platelet contamination was verified using another dataset.
[0100] Twelve groups of pure platelet and pure PPP samples were collected and then each sample was counted using a platelet counter to determine the blood cell distribution of each sample, so as to determine the purity of the platelet and PPP samples. Then, every three groups of platelets and PPPs were mixed, and finally 4 groups of pure platelet samples (platelet) and pure PPP samples were obtained. Ten-fold dilutions of the platelet and PPP samples were made for the 4 groups of samples ( Figure 1 e), and then each sample was counted using a platelet counter to determine the absolute platelet content of each sample.
[0101] Correlation analysis was performed between the platelet concentration and the abundances of 30 biomarkers, and the results were good. And a spline regression was used to fit the absolute platelet content and the contamination index, and the results showed that R 2 = 0.95, showing a good correlation ( Figure 1 f).
[0102] Using the spline regression model, the index was mapped to the actual cell count; the model construction method was as follows: Based on the ten-fold dilution data of the 4 groups of samples in the validation set (each dilution gradient included 3 biological replicates, and a total of 117 samples after removing 3 abnormal samples), a spline function (Univariate Spline) was used for fitting and modeling. With the contamination index as the independent variable and the actual platelet count as the dependent variable, the number of knots k = 3 and the smoothing coefficient s = 50 were set (through multiple cross-validations, this parameter setting could maximize the model's balance between flexibility and smoothness). Finally, a platelet count prediction curve (R 2 = 0.95) was obtained, and the model expression was: platelet count = Spline(contamination index), where the coefficients (coeffs) and knots of the spline function were as follows: coeffs[6.20722104,6.21479977,6.71490564,7.23311078] / knots[0.00935505,0.0804042].
[0103] Finally, the spline regression was applied to 11 pairs of PPP and PRP samples ( Figure 1 g), and the results showed that the model showed good discrimination for the two types of samples ( Figure 1 h). It shows that the contamination index can be used to evaluate the number of platelets and the degree of platelet contamination in the sample.
[0104] Subsequently, a new set of plasma samples was selected, and the enriched plasma was obtained using the procedure for enriching low-abundance plasma proteins based on nanoparticles in Example 1. The contaminated samples were detected by the aforementioned method for evaluating platelet contamination, and an ROC curve was plotted. The results showed that it had high sensitivity and specificity.
[0105] Example 3 Universality of Platelet Contamination Biomarkers
[0106] To verify the applicability of the screened biomarkers in samples treated with other types of nanoparticles, two frequently reported nanoparticles in the literature were selected: NaY zeolite (particle size 300 - 700 nm) and silanol-functionalized iron oxide (particle size 400 - 700 nm). PPP (platelet-poor plasma) and PRP (platelet-rich plasma) samples were prepared by collecting donor blood, and 100 μL of each was mixed and subjected to 10-step serial gradient dilution. The diluted plasma samples were treated with NP74 and NP81 nanoparticles respectively, and nDIA analysis was performed.
[0107] The results showed that: as the platelet concentration increased, the number of peptide precursors and proteome identifications in samples treated with both types of nanoparticles increased significantly, and the number of proteomes detected in a single injection exceeded 6000. Thirty platelet-related biomarkers showed a high degree of correlation with the degree of platelet contamination in samples treated with NaY zeolite (median spearman correlation between the protein abundance of 30 proteins evaluating platelet contamination and the platelet contamination count: 0.95, range 0.89 - 0.96) and iron oxide (median: 0.93, range 0.87 - 0.94) Figure 2 ).
[0108] This also indicates that the method for evaluating platelet contamination of the present invention is compatible with various nanoparticles such as molecular sieves and Fe3O4, and the biomarker correlation is maintained at 0.87 - 0.96, providing a unified framework for cross-laboratory data standardization.
[0109] Example 4 Application Example
[0110] To evaluate the application effect of the technology for enriching low-abundance plasma proteins based on nanoparticles and evaluating platelet contamination, 193 subjects were included in this study, including 42 patients with benign lung nodules and 151 patients with early-stage malignant tumors.
[0111] All plasma samples were collected by drawing blood using ethylenediaminetetraacetic acid (EDTA) vacuum tubes. For some patients, secondary sampling was performed to evaluate the stability of the pretreatment. Repeatedly processed samples were not included in subsequent modeling. After centrifuging the blood (3000 g, 4 °C, 15 minutes), the plasma was collected and stored at 80 °C. Then, peptide samples were obtained using the process for enriching low-abundance plasma proteins based on nanoparticles in Example 1, combined with Astral instrument analysis to achieve quantitative analysis of peptides and proteins. Proteomics analysis showed that an average of 4413 proteomes were identified in each plasma sample, and both biological replicates and technical replicates exhibited stable coefficients of variation. The platelet situation in the samples was evaluated by the aforementioned method for assessing the degree of contamination, and it was found that 1 sample had detectable contamination indicators ( Figure 3 ), and these contaminated samples were excluded from subsequent analyses, which was beneficial for the subsequent screening of benign pulmonary nodules / lung cancer markers.
[0112] The above embodiments merely illustrate the principles and effects of the present invention, rather than limiting the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.
Claims
1. A marker combination for evaluating the degree of platelet contamination, the marker combination comprising IDH2, LIMS1, MLEC, DIAPH1, GNAI2, ATP5F1B, RAB14, TBXAS1, JAM3, VDAC1, ATP2A3, ARHGAP6, VDAC3, ESAM, SLC25A5, ITGA6, PDLIM7, PF4V1, HSP90B1, HSPD1, RDH11, ATP2A2, VDAC2, RAB35, TREML1, GP5, PRDX3, ITGB1, MDH2 and GNB1.
2. Use of the marker combination according to claim 1 or a reagent for detecting the marker combination according to claim 1 in at least one of the following: 1) Evaluating the degree of platelet contamination; 2) Preparing a product for evaluating the degree of platelet contamination; 3) Calculating the platelet count; 4) Preparing a product for calculating the platelet count.
3. The application according to claim 2, wherein The reagent is detected by at least one method of mass spectrometry, chromatography, surface-enhanced Raman spectroscopy, western blotting, flow cytometry, protein array, immunoprecipitation, immunosorption; preferably, the reagent comprises one or more of a substance specific for the marker, a probe specific for the marker, and a protein chip; Preferably, the substance specific for the marker comprises an antibody, a ligand protein, a polypeptide, a non-protein compound, and / or a nucleic acid aptamer.
4. A product, characterized in that, The product comprises a reagent for detecting the marker combination according to claim 1.
5. The product according to claim 4, wherein The reagent is detected by at least one method of mass spectrometry, chromatography, surface-enhanced Raman spectroscopy, western blotting, flow cytometry, protein array, immunoprecipitation, immunosorption; preferably, the reagent comprises one or more of a substance specific for the marker, a probe specific for the marker, and a protein chip; Preferably, the substance specific for the marker comprises an antibody, a ligand protein, a polypeptide, a non-protein compound, and / or a nucleic acid aptamer; Preferably, the product comprises a reagent, a kit, a test strip, a chip, a device, a detection system.
6. A method for evaluating the degree of platelet contamination, comprising the following steps: a) Collecting a sample to be tested and measuring the abundance of the markers in the marker combination according to claim 1; b) Calculating the contamination index and / or the platelet count of the sample to be tested, and the calculation formula of the contamination index is as follows: Pollution index = (IDH2 abundance + LIMS1 abundance + MLEC abundance + DIAPH1 abundance + GNAI2 abundance + ATP5F1B abundance + RAB14 abundance + TBXAS1 abundance + JAM3 abundance + VDAC1 abundance + ATP2A3 abundance + ARHGAP6 abundance + VDAC3 abundance + ESAM abundance + SLC25A5 abundance + ITGA6 abundance + PDLIM7 abundance + PF4V1 abundance + HSP90B1 abundance + HSPD1 abundance + RDH11 abundance + ATP2A2 abundance + VDAC2 abundance + RAB35 abundance + TREML1 abundance + GP5 abundance + PRDX3 abundance + ITGB1 abundance + MDH2 abundance + GNB1 abundance) / total protein abundance; The calculation method of the platelet count is as follows: mapping the pollution index to the actual cell count through a regression model; 3) Predict the pollution situation of the sample according to the calculated pollution index and / or platelet count of the sample to be tested: the lower the pollution index and / or platelet count of the sample to be tested, the lower the platelet pollution degree of the sample to be tested; the higher the pollution index and / or platelet count of the sample to be tested, the higher the platelet pollution degree of the sample to be tested.
7. The method according to claim 6, characterized in that, The sample is a plasma, tissue, cell and / or serum sample; preferably, the sample includes the sample obtained through pretreatment, and the pretreatment includes or does not include the step of incubating with nanoparticles; Preferably, the pretreatment specifically includes the following steps: incubating the plasma sample with nanoparticles, and performing washing, denaturation, and enzymatic digestion after the incubation is completed.
8. A system for evaluating the degree of platelet pollution, comprising the following modules: a) Data collection module: collecting the sample to be tested and measuring the abundance of the markers in the above marker combination; b) Model calculation module: calculating the pollution index and / or platelet count of the sample to be tested, and the calculation formula of the pollution index is as follows: Pollution index = (IDH2 abundance + LIMS1 abundance + MLEC abundance + DIAPH1 abundance + GNAI2 abundance + ATP5F1B abundance + RAB14 abundance + TBXAS1 abundance + JAM3 abundance + VDAC1 abundance + ATP2A3 abundance + ARHGAP6 abundance + VDAC3 abundance + ESAM abundance + SLC25A5 abundance + ITGA6 abundance + PDLIM7 abundance + PF4V1 abundance + HSP90B1 abundance + HSPD1 abundance + RDH11 abundance + ATP2A2 abundance + VDAC2 abundance + RAB35 abundance + TREML1 abundance + GP5 abundance + PRDX3 abundance + ITGB1 abundance + MDH2 abundance + GNB1 abundance) / total protein abundance; The calculation method of the platelet count is as follows: mapping the pollution index to the actual cell count through a regression model; c) Output prediction module: predicting the contamination situation of the sample according to the calculated contamination index and / or platelet count of the sample to be tested: wherein, the lower the contamination index and / or platelet count of the sample to be tested, the lower the contamination degree; the higher the contamination index and / or platelet count of the sample to be tested, the higher the contamination degree.
9. A computing device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 6 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 6 to 7 are implemented.
Citation Information
Patent Citations
Method for evaluating blood pollution degree in sample
CN118518802A
KR20250022647A
Cited By
Protein molecule combination for predicting diabetic complications, product and application
CN122063277A