Novel biomarkers for beta-thalassemia screening and diagnosis and uses thereof

By using PLXDC2[90]_C, CDH1[152]_C and IGF2[126]_C polypeptide cluster biomarkers, combined with LC-MS/MS technology, the problems of high misdiagnosis rate and high missed detection rate in existing methods were solved, and high sensitivity and high specificity of β-thalassemia screening and diagnosis were achieved.

CN121090843BActive Publication Date: 2026-03-03INSTITUTE OF BIOPHYSICS CHINESE ACADEMY OF SCIENCES
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511308661.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2026-03-03
Estimated Expiration
2045-09-15

AI Technical Summary

Technical Problem

Existing screening and diagnostic methods for β-thalassemia have high rates of misdiagnosis and missed detection, and they are difficult to detect uncommon gene mutation types. Traditional gene testing methods are limited to pre-specified mutations, resulting in the missed detection of unassociated mutation types.

Method used

The peptide cluster biomarkers, including PLXDC2[90]_C, CDH1[152]_C and IGF2[126]_C peptide clusters, were detected by liquid chromatography-tandem mass spectrometry (LC-MS/MS) combined with high performance liquid chromatography (HPLC) and capillary electrophoresis (CE) to screen and diagnose β-thalassemia.

Benefits of technology

It improves the accuracy and specificity of β-thalassemia screening and diagnosis, reduces the misdiagnosis rate, provides a highly sensitive detection method, simplifies the diagnostic process, and can identify multiple gene mutation types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121090843B_ABST
    Figure CN121090843B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of biology, and particularly relates to a novel biomarker for screening and diagnosis of beta-thalassemia and application thereof. The application provides a novel polypeptide cluster biomarker for screening and diagnosis of beta-thalassemia based on polypeptides in plasma / serum obtained in vitro from patients, which has the advantages of high sensitivity and high specificity. The biomarker provided by the application can be detected by mass spectrometry, and the detection method is stable and efficient, and does not need to adopt cumbersome steps for inspection, thereby providing laboratory support for screening and diagnosis and treatment of beta-thalassemia.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biotechnology, and in particular relates to a novel biomarker for screening and diagnosis of β-thalassemia and its application. Background Technology

[0002] Beta-thalassemia is a hereditary blood disorder caused by mutations in the hemoglobin subunit beta gene (HBB). Current research and literature reports over 350 mutations associated with beta-thalassemia, exhibiting diverse mutation types including point mutations, small deletions, small insertions, and gene rearrangements. Heterozygous individuals with beta-thalassemia mutations are typically asymptomatic, a condition known as trisomy 2 (TT), while homozygous or compound heterozygous beta-thalassemia mutations lead to intermediate beta-thalassemia (TI) and severe beta-thalassemia (TM). Early detection of beta-thalassemia gene carriers is a key strategy for the prevention and control of this disease.

[0003] Currently, combining information such as red blood cell markers and family history to qualitatively and quantitatively analyze different types of hemoglobin is the first-line screening strategy for β-thalassemia diagnosis. Commonly used analytical techniques include high-performance liquid chromatography (HPLC) and capillary electrophoresis; however, these methods have insufficient accuracy, sensitivity, and specificity. Before making a diagnosis, it is necessary to rule out some suspected diseases, such as iron deficiency anemia, structural hemoglobin variations, and anemia caused by chronic diseases. Therefore, this screening strategy has a high rate of misdiagnosis and missed detection.

[0004] In addition, genetic testing is another widely used method for screening β-thalassemia. For example, reverse dot blot hybridization (RDB) is currently the most commonly used method in first-line clinical practice in my country. It can simultaneously detect 17 β-thalassemia mutations, primarily targeting common globin gene mutation types in my country. Cross-breakpoint PCR (GAP-PCR) is a routine method for detecting β-thalassemia deletion mutations, mainly used to detect the more common Chinese-type Gγ+ (Aγδβ). 0 Deletions and Southeast Asian HPFH deletions, among others. The commonly used gene testing methods mentioned above are mainly used to detect pre-specified mutations, leading to the missed diagnosis of uncommon mutation types or mutation types not yet associated with the disease.

[0005] Over the years, numerous studies have revealed significant changes in erythrocytes and their fluid environment composition due to HBB gene mutations. Enhanced proteolysis in erythroid cells of thalassemia is a long-standing observation, and protease activity in β-thalassemia may be correlated with disease severity. Furthermore, modulating protein quality control pathways, particularly the proteasome and autophagy, may be a potential therapeutic strategy for β-thalassemia. Our plasma proteomic analysis revealed that among the top five most population-discriminating proteins, in addition to two classic ferritin subunits, the remaining candidate proteins—platelet-activating factor acetylhydrolase and cathepsin S—were clearly classified as proteases. Simultaneously, upregulation of many proteases, such as those in the lysosomal pathway, was observed in patients with β-thalassemia. Furthermore, in plasma extracellular vesicle proteomics studies, we also identified complement C1s subcomponents belonging to proteases in the final selected protein composition with diagnostic potential. Importantly, inhibition of membrane protease-2 (TMPRSS6), a type II transmembrane serine protease, led to upregulation of hepcidin (HAMP) and improved many disease symptoms associated with β-thalassemia. The close association between proteases and β-thalassemia suggests that protease dysregulation and function in β-thalassemia is an important but under-explored area. However, given the diversity, wide range of sources, and complex catalytic mechanisms of enzymes, enzymology has long been considered a mysterious "black box." Tracing changes in enzyme systems from their substrates may be a more effective approach.

[0006] With advancements in mass spectrometry technology, peptidomics research has made significant progress, enabling large-scale identification and quantitative analysis of endogenous peptides. It is increasingly being used in research on clinical biomarkers for various diseases. Peptides are short-chain molecules composed of amino acids. Endogenous peptides are mainly produced by the cleavage and degradation of precursor proteins. Changes in their amino acid sequence composition and abundance can reveal changes in the activity of upstream enzyme systems. However, peptides generated by the cleavage of proteins by dysregulated specific enzymes in disease states can produce peptide clusters with identical terminal sequences and a ladder-like sequence under uncontrolled non-specific enzyme cleavage. This redundant information is one of the main bottlenecks in peptidomics research. The discovery process of peptide cluster biomarkers includes the discovery of amino acid sites with diagnostic potential and the mining of representative peptide clusters. Theoretically, using only a single peptide to diagnose a disease reduces or minimizes the differences in enzyme activity. Peptide clusters more accurately reflect these differences and are suitable as potential biomarkers. Therefore, using LC-MS / MS technology to discover novel polypeptide cluster biomarkers that can be used for screening and diagnosis of β-thalassemia, and developing efficient screening methods that are alternatives to or complementary to existing methods, will simplify the diagnostic process, improve the accuracy and comprehensiveness of diagnosis, and reduce the rate of missed diagnoses. Summary of the Invention

[0007] To address the aforementioned problems, the present invention provides a polypeptide cluster for detecting β-thalassemia, the polypeptide cluster comprising...

[0008] PLXDC2

[90] _C, which is a polypeptide cluster whose C-terminus is the 90th amino acid site of the PLXDC2 protein; CDH1

[152] _C, which is a polypeptide cluster whose C-terminus is the 152nd amino acid site of the CDH1 protein; IGF2

[126] _C, which is a polypeptide cluster whose C-terminus is the 126th amino acid site of the IGF2 protein.

[0009] In a preferred embodiment, the polypeptide cluster comprises the above-described combination of polypeptide clusters.

[0010] The aforementioned polypeptide cluster contains one or more polypeptides. Preferably, the PLXDC2

[90] _C polypeptide cluster contains four polypeptides;

[0011] The CDH1

[152] _C polypeptide cluster contains two polypeptides;

[0012] The IGF2

[126] _C polypeptide cluster contains 5 polypeptides.

[0013] In one specific implementation, the amino acid sites to which these polypeptide clusters belong, the specific amino acid sequences, the start and end sites of the sequences, the gene names of the proteins to which they belong, and the Accession ID are shown in Table 1.

[0014] Table 1

[0015]

[0016] Another aspect of the present invention provides a kit comprising the above-described detection of β-thalassemia, the kit comprising reagents for detecting the above-described polypeptide clusters.

[0017] The reagents are selected from primers, probes, antibodies, or other reagents that can quantitatively detect peptides in the above-mentioned peptide clusters;

[0018] In a more preferred technical solution, one or more of the following techniques can be selected: liquid chromatography-tandem mass spectrometry (LC-MS / MS), high performance liquid chromatography (HPLC) coupled with UV, and capillary electrophoresis (CE) coupled with UV.

[0019] In one specific implementation, the kit further comprises synthetic peptides corresponding to the peptide clusters and / or mixtures of synthetic peptides and / or references;

[0020] In one specific implementation, the kit also includes reagents commonly used for enriching plasma peptides;

[0021] Preferably, the kit may further include standards for mass spectrometry retention time correction.

[0022] A third aspect of the present invention provides the application of the above-described kit, which can determine whether a subject carries or has β-thalassemia or is at risk of having β-thalassemia by detecting the content of the above-described polypeptide clusters.

[0023] In one specific implementation plan, the prognosis of a patient can be determined by the content of the aforementioned polypeptide clusters in serum or plasma;

[0024] In one specific implementation plan, the effectiveness of a patient's surgery, drug treatment, etc., or when to discontinue treatment can be determined by the content of the aforementioned polypeptide clusters in serum or plasma. Beneficial effects

[0025] (1) The present invention provides a novel polypeptide cluster biomarker based on polypeptides obtained in vitro from the plasma / serum of patients as screening and diagnosis of β-thalassemia, which has the advantages of high sensitivity and high specificity.

[0026] (2) The biomarkers provided by this invention can be detected by mass spectrometry. The detection method is stable and efficient, and there is no need to use cumbersome steps for examination, which provides laboratory support for the screening, diagnosis and treatment of β-thalassemia.

[0027] (3) Integrating data and information obtained through omics technology with clinical information will help to better understand the pathophysiological process of β-thalassemia, find new intervention targets, and thus improve the clinical management of patients. Attached Figure Description

[0028] Figure 1 These are polypeptides with diagnostic potential in the polypeptide group during the implementation of this invention.

[0029] Figure 2 These are important disease-related amino acid sites discovered in the polypeptide group during the implementation of this invention.

[0030] Figure 3 This describes the overlap between amino acid sites corresponding to peptides with diagnostic potential and important amino acid sites related to diseases, discovered during the implementation of this invention. There are a total of 7 stable amino acid sites belonging to peptides with diagnostic potential, 6 of which are located in amino acid sites shared by both.

[0031] Figure 4Figure A shows the differences in individual aa-scores of amino acid sites in different groups during the implementation of this invention (PRM targeting analysis) and the diagnostic performance of representative polypeptide clusters at amino acid sites (Figure B). The individual aa-score is calculated based on the absolute concentration of the polypeptide.

[0032] Figure 5 Figure A shows the differences in individual aa-scores of amino acid sites in different groups during the implementation of this invention (PRM targeted analysis) and the diagnostic performance of representative polypeptide clusters at amino acid sites (Figure B). The individual aa-score is calculated based on a mixture of relabeled polypeptides, i.e., internal standard.

[0033] Figure 6 Figure A shows the differences in individual aa-scores of amino acid sites in different groups during the implementation of this invention (PRM targeting analysis) and the diagnostic performance of representative polypeptide clusters at amino acid sites (Figure B). The individual aa-scores are calculated based on reference samples.

[0034] Figure 7 This shows the distribution of individual aa-scores of amino acid sites in different groups during the implementation of this invention to evaluate the reliability of the detection method. Detailed Implementation

[0035] The following description, with reference to the accompanying drawings, illustrates several preferred embodiments of the present invention to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms, and the scope of protection of the present invention is not limited to the embodiments mentioned herein.

[0036] Example 1: Discovery of amino acid sites and their polypeptide clusters with diagnostic potential

[0037] Patients and healthy participants involved in this invention were recruited by the First Affiliated Hospital of Guangxi Medical University (Guangxi Zhuang Autonomous Region, China). All participants provided written informed consent; for patients under 18 years of age, consent from their parents or legal guardians was required. Two weeks after their last transfusion, patients provided blood samples and completed a questionnaire including information on their first transfusion, treatment, and other medical conditions. This study recruited 286 patients or carriers of β-thalassemia and 51 healthy controls. Patients included in the plasma polypeptide study received both transfusions and iron chelation therapy; none underwent splenectomy. Whole blood was collected in EDTA vacuum aspirators and centrifuged at 3000 × g for 10 minutes. Plasma collected from the supernatant was stored at -80˚C until use. This study was designed and conducted in accordance with the Declaration of Helsinki. The Ethics Committee of the First Affiliated Hospital of Guangxi Medical University approved this study.

[0038] The above samples were diagnosed by BGI Genomics Clinical Laboratory (Shenzhen, China) using Gap-PCR and SNP testing to determine the mutation type of the patients. Mild β-thalassemia carriers were identified as β0βN / β+βN, intermediate β-thalassemia patients as β+β+ / β+β0, and severe β-thalassemia patients as β0β0. Gene mutations are categorized as follows: β+ mutations include -28A>G (HBB: c. -78A>G), -29A>G (HBB: c. -79A>G), codon 26G>A (HbE, HBB: c. 79G>A), IVS-II-654C654>T (HBB: c. 316-197C>T) and IVS-II-5G>C (HBB: c. 315+5G>C). β0 mutations include codons 41 / 42-TTCT (HBB: c. 126_129delCTTT), codon 17A>T (HBB: c. 52A>T), codons 71 / 72+A (HBB: c. 216_217insA), IVS-I-1G>T (HBB: c. 92+1G>T), codon 43G>T (HBB: c.130G>T), IVS-I-130 G>C (HBB: c.93-1G>C), codon37 G>A (HBB: c.114G>A), codons 27 / 28 + C (HBB: c.84_85insC), and codon 30A>G (HBB: c.91A>G).

[0039] Clinical parameters including HbF, SF, HbA2, HGB, MCV, and MCH were measured using standard techniques. These included the CELL-DYN fully automated hematology analyzer (Abbott Diagnostics) for measuring MCV, MCH, and HGB; high-performance liquid chromatography (VARIANT II, ​​Bio-Rad) for measuring HbF and HbA2; and electrochemiluminescence immunoassay (Cobas e601, Roche) for measuring serum ferritin (SF).

[0040] 1. Sequential precipitation defatting (SPD) method for separating and enriching peptides in plasma samples

[0041] 1.1 Add 250 µL of methyl tert-butyl ether (MTBE), 50 µL of deionized water and 150 µL of methanol to 50 µL of plasma in sequence, mix well and let stand at 4°C for 30 min.

[0042] 1.2 Centrifuge the above sample at 21,000 g, 4℃ for 30 min, and take out the supernatant after centrifugation.

[0043] 1.3 Add 500 µL MTBE and 100 µL deionized water to the supernatant above, mix well, and centrifuge at 1,000 g and 4 °C for 10 min.

[0044] 1.4 The above samples were divided into two phases. The upper layer, rich in hydrophobic interfering substances such as lipids, was removed, and the lower clear liquid was retained and dried at 4°C.

[0045] The obtained samples were processed by a desalting column and then ready for testing.

[0046] 2. Whole-peptide genome study population cohort

[0047] The first phase enrolled 54 participants, including 13 healthy controls (Ctr), 8 mild β-thalassemia carriers (TT), 17 intermediate β-thalassemia patients (TI), and 16 severe β-thalassemia patients (TM).

[0048] 3. Detection of the whole peptide genome using liquid chromatography-mass spectrometry (LC-MS / MS)

[0049] 3.1 Dissolve the peptide in 15 µL of 0.1% FA aqueous solution, centrifuge at 21,000 g, 4℃ for 30 min, and then take 10 µL of the supernatant into a sample vial.

[0050] 3.2 Data acquisition was performed using an Orbitrap Exploris 480 mass spectrometer equipped with FAIMS Pro.

[0051] The LC-MS / MS instrument used was an OrbitrapExploris 480 mass spectrometer (Thermo Fisher Scientific) equipped with an EASY-nLC 1200 HPLC system and FAIMS Pro. The analytical liquid chromatography and mass spectrometry conditions were as follows: a Dr. Maisch GmbH ReproSil-Pur C18 AQ column (75 μm id × 20 cm, 3 μm) was used; the mobile phase was: A: water containing 0.1% formic acid, B: acetonitrile containing 0.1% formic acid; the flow rate was 300 nL / min; gradient elution was used: 4–11% B, 4 min; 11–21% B, 28 min; 21–30% B, 29 min; 30–42% B, 27 min; 42–95% B, 5 min; 95% B, 10 min. Two compensation voltages (-45 V and -65 V) were used for FAIMS separation. In data-dependent acquisition (DDA) mode, MS1 scan range: 350-1600 m / z, normalized AGC: 300%; maximum injection time: 80 ms; MS1 resolution: 60,000 (m / z 200); isolation window width: 1.6 m / z; HCD collision energy: 28%. MS / MS resolution: 60,000 (m / z 200); normalized AGC: 150%; maximum ion injection time: 118 ms; dynamic exclusion time: 30 s.

[0052] 3.3 Protein Discoverer 2.4 software was used to retrieve mass spectrometry data.

[0053] Mass spectrometry data were retrieved to obtain quantitative information on the whole peptide genome. Protein Discoverer 2.4 software was used, and the key parameters for data retrieval were set as follows: the database selected was the human database downloaded from Uniprot in September 2019; protease was set to No enzyme; the maximum errors for precursor and daughter ions were 10 ppm and 0.02 Da, respectively; variable modifications included oxidation of methionine and proline, cysteine ​​modification, and conversion of glutamine to pyroglutamic acid; the free dose ratio (FDR) for peptide level was set to 1%; label-free quantification was performed based on chromatographic area.

[0054] 4. Analysis of whole-peptide genome detection results

[0055] 4.1 Differential analysis and ROC curve analysis of peptides.

[0056] The p-value for comparison between the two groups was calculated using the t-test statistical method. The screening criteria for differentially regulated peptides were: fold change > 2 and p < 0.05 for upregulated peptides, and fold change < 0.5 and p < 0.05 for downregulated peptides. ROC curve analysis was performed on the differentially regulated proteins.

[0057] 4.2 Discovering stable amino acid sites for the classification of differentially expressed peptides with diagnostic potential.

[0058] In the three comparison groups, peptides with AUC > 0.9 were considered to have diagnostic potential; a total of 43 such differentially expressed peptides were identified. Figure 1 It contains 76 protein amino acid sites, of which 7 are stable amino acid sites.

[0059] 4.3 Identification of important amino acid sites related to the disease based on the grouped aa-score algorithm

[0060] Data analysis was performed using the R package aascore (1.0.0). First, amino acid site assignment analysis was performed on the differentially expressed peptides, identifying 1049 changing points. The grouped aa-score algorithm was used to calculate the aa-score value of each amino acid site associated with the differentially expressed peptide. Simultaneously, the average aa-score value for each site across three pairs of differential comparisons was calculated. The average aa-score was normalized using the full length of the protein to obtain the normalized mean aa-score. Transition points showing significant changes in trend (i.e., positive, negative, or zero change in the curve slope) were identified among the changing points on the waterfall map curve. A total of 186 transition points were present in all subgroups compared to healthy controls; these sites were considered important amino acid sites associated with the disease. Figure 2 ).

[0061] 4.4 Screening for amino acid sites and polypeptide clusters with diagnostic potential

[0062] Comparative analysis of stable amino acid sites and disease-related important amino acid sites contained in differentially expressed peptides with diagnostic potential revealed that 6 sites were common amino acid sites. Figure 3 These shared amino acid sites, along with the first few disease-related important amino acid sites that have both high mean of aa-score and normalized mean of aa-score, are considered to have the greatest diagnostic potential. The polypeptide clusters of the amino acid sites with the greatest diagnostic potential include all differentially expressed polypeptides at that site.

[0063] Example 2: Screening of representative polypeptide clusters with diagnostic potential amino acid sites, identification and validation of potential biomarkers for polypeptide clusters.

[0064] 5. Preparation of peptide samples from targeted analysis cohort plasma and mixed plasma

[0065] This cohort included 49 participants, comprising 16 healthy controls, 17 carriers of mild β-thalassemia, and 16 patients with severe β-thalassemia. Peptides were extracted from the plasma of the validation cohort using the SPD method, following the same procedure as in step 1. Equal volumes of plasma samples from the cohort were mixed, and peptides were extracted using the SPD method. These peptides were then used in the preliminary and formal PRM experiments.

[0066] 6. PRM preliminary screening of amino acid sites and representative polypeptide clusters for targeted analysis

[0067] 6.1 Establishment of PRM Pre-experiment Methods

[0068] The method was established using SpectroDive v12.1 (Biognosys) software. When developing the PRM detection method, the Pulsar engine was used to retrieve raw data from the entire peptide genome, generating a spectral library. Key search parameters were set as follows: non-enzymatic digestion; peptide length 7-40; modification: variable modification: Gln->pyro-Glu, Oxidation (P, M). Amino acid sites with diagnostic potential and their contained peptides were included in the PRM preliminary analysis list as much as possible, with the following parameter settings: precursor ion mass-to-charge ratio range, 350-1500; precursor ion charge number, 2-6; peptide length, 8-40; daughter ion mass-to-charge ratio range, 300-1800; maximum daughter ion charge number, 3; ion type, b, y ions; allowed neutral loss types, H2O, NH3, and no loss; top 6 daughter ions were selected. Unscheduled PRM was selected to establish the preliminary experimental method.

[0069] 6.2 PRM Data Acquisition

[0070] The LC-MS / MS instrument used in the parallel reaction monitoring (PRM) targeted quantitative analysis stage was an Orbitrap Eclipse Tribrid mass spectrometer (Thermo Fisher Scientific) equipped with an EASY-nLC1200 HPLC system. The chromatographic column and flow were the same as in step 2.2. The chromatographic elution gradient was set as follows: 5-10%, 3 min; 10-20% B, 22 min; 20-30% B, 22 min; 30-40% B, 13 min; 40-99% B, 4 min; 95% B, 9 min. The mass spectrometry parameters for PRM data acquisition were set as follows: full MS mass-to-charge ratio scan range 350-1200 m / z, resolution 60,000 (m / z 200), AGC 6 × 10⁻⁶. 5 The maximum ion implantation time was 50 ms. The MS / MS resolution was 30,000 (m / z 200), the isolation window was 1.0 Da, the collision energy was 30%, and the AGC was 2 × 10⁻⁶. 5 The maximum ion implantation time is 80 ms.

[0071] 6.3 PRM mass spectrometry data analysis and relabeled peptide synthesis

[0072] Data were analyzed using SpectroDive v12.1 (Biognosys) software. Based on the preliminary PRM experimental results, the target peptides were further screened. The screening criteria were as follows: a. Targetable by PRM; b. Good peak shape with no interference; c. High abundance. Finally, nine amino acid sites and their representative peptide clusters were selected for further analysis: PLXDC2

[90] _C, C3

[1320] _N, CDH1

[152] _C, AHSG

[339] _C, SRGN

[130] _N, SRGN

[72] _N, IGF2

[126] _C, APOC3

[21] _N, and ITIH4

[668] _C. Multiple peptides with diagnostic potential at the corresponding amino acid sites were preferentially selected to synthesize relabeled peptides; when no peptides with diagnostic potential were found, relabeled peptides were synthesized based on abundance.

[0073] 7. Establish a targeted analysis method to calculate the individual aa-score.

[0074] 7.1 Establish a targeted analysis method for formal experiments, and conduct targeted analysis on the cohort samples.

[0075] Based on the abundance ratio of the target endogenous peptide in the preliminary experimental samples, a mixed labeled synthetic peptide was used as an internal standard. A targeted analysis method was established using this mixed sample. The internal standard and a 10×iRT standard peptide (Biognosys) were added to the mixed sample for unscheduled PRM analysis. This established a scheduled PRM targeted quantitative analysis method that includes the m / z of the labeled peptide precursor ion, the m / z of the endogenous target peptide precursor ion, and the retention time of the precursor ion. The retention time window was set to ±2.5 minutes.

[0076] Based on the above method, the scheduled PRM data collection for the queue samples was formally carried out.

[0077] 7.2 Establish a standard curve for the recalibrated peptides for absolute quantitative analysis.

[0078] Based on the expected abundance of the target peptide, serially diluted recalibrated peptide standards are prepared. The serially diluted recalibrated peptide standards and iRT peptides are added to the mixed sample for scheduled PRM analysis to establish a standard curve. This mixed sample is referred to as the reference sample.

[0079] 7.3 PRM data analysis: Calculation of individual aa-score values ​​for amino acid sites using different methods

[0080] Data was processed in SpectroDive software to exclude daughter ions with significant interfering signals. For each parent ion, at least three daughter ion pairs were selected for quantitative analysis. The data was then exported and further processed in R language.

[0081] Individual aa-score is calculated in three ways: a. Based on the established standard curve, the absolute concentration of peptide clusters at amino acid sites is calculated; b. Based on the relabeled peptide mixture standard added to a single sample, i.e., the internal reference standard, the relative content of peptide clusters at amino acid sites is calculated; c. Based on the average abundance of endogenous peptides in the reference sample, the relative content of peptide clusters at amino acid sites in a single sample is calculated. Methods a and b are based on relabeled peptides. The representative peptide clusters involved in the individual aa-score calculation of the two amino acid sites PLXDC2

[90] _C and CDH1

[152] _C are shown in Table 2; c is a calculation method based on reference samples. The representative peptide clusters involved in the individual aa-score calculation of the three amino acid sites PLXDC2

[90] _C, CDH1

[152] _C and IGF2

[126] _C are shown in Table 3.

[0082] Preferably, the representative polypeptide clusters of the two amino acid sites PLXDC2

[90] _C and CDH1

[152] _C are shown in Table 2.

[0083] Table 2. Representative polypeptide clusters involved in the amino acid sites of PLXDC2

[90] _C, CDH1

[152] _C, and IGF2

[126] _C when analyzed based on relabeled polypeptides.

[0084] Sequence Gene name gene position aa position DTNRASVGQDSPEPR PLXDC2 PLXDC2[76-90] PLXDC2

[90] _C FLKAVDTNRASVGQDSPEPR PLXDC2 PLXDC2[71-90] PLXDC2

[90] _C SGIQAELLTFPNSSPGLRRQ CDH1 CDH1[133-152] CDH1

[152] _C SVSGIQAELLTFPNSSPGLRRQ CDH1 CDH1[131-152] CDH1

[152] _C

[0085] Table 3. Representative polypeptide clusters involved in the amino acid sites of PLXDC2

[90] _C, CDH1

[152] _C, and IGF2

[126] _C when analyzed based on the Reference sample.

[0086] Sequence Gene name gene position aa position DTNRASVGQDSPEPR PLXDC2 PLXDC2[76-90] PLXDC2

[90] _C VDTNRASVGQDSPEPR PLXDC2 PLXDC2[75-90] PLXDC2

[90] _C AVDTNRASVGQDSPEPR PLXDC2 PLXDC2[74-90] PLXDC2

[90] _C FLKAVDTNRASVGQDSPEPR PLXDC2 PLXDC2[71-90] PLXDC2

[90] _C SGIQAELLTFPNSSPGLRRQ CDH1 CDH1[133-152] CDH1

[152] _C SVSGIQAELLTFPNSSPGLRRQ CDH1 CDH1[131-152] CDH1

[152] _C DTWKQSTQRL IGF2 IGF2 [117-126] IGF2

[126] _C Q[Gln->pyro-Glu]YDTWKQSTQRL IGF2 IGF2 [115-126] IGF2

[126] _C YDTWKQSTQRL IGF2 IGF2 [116-126] IGF2

[126] _C FFQYDTWKQSTQRL IGF2 IGF2 [113-126] IGF2

[126] _C QYDTWKQSTQRL IGF2 IGF2 [115-126] IGF2

[126] _C

[0087] Based on the calculation method of method a, the differences in individual aa-score of amino acid sites in different groups were calculated using the two independent samples Wilcoxon test. Figure 4 a). The two amino acid sites PLXDC2

[90] _C and CDH1

[152] _C showed statistically significant differences between the β-thalassemia group and the healthy control group, with p-values ​​of 5.97e-13 and 6.21e-12, respectively. The diagnostic potential of the individual aa-score of the amino acid site was calculated using the ROC curve analysis method in the pROC package of R language. Figure 4 b) The AUC of PLXDC2

[90] _C was 100% (95% CI: 1-1, p<0.05), with both specificity and sensitivity at 100%, and the optimal cutoff value for absolute content was 0.5346 fmol / µL; the AUC of CDH1

[152] _C was 99.4% (95% CI: 0.977-1, p<0.05), with specificity at 100.0%, sensitivity at 96.9%, and the optimal cutoff value for absolute content was 6.2011 fmol / µL. (The optimal cutoff value was obtained by maximizing the Youden index).

[0088] Based on the calculation method of method b, the differences in individual aa-scores of amino acid sites in different groups were calculated using the two independent samples Wilcoxon test. Figure 5a). The two amino acid sites PLXDC2

[90] _C and CDH1

[152] _C showed statistically significant differences between the β-thalassemia group and the healthy control group, with p-values ​​of 5.97e-13 and 8.87e-13, respectively. The diagnostic potential of the individual aa-score of the amino acid site was calculated using the ROC curve analysis method in the pROC package of R language. Figure 5 b) The AUC of PLXDC2

[90] _C was 100% (95% CI: 1-1, p<0.05), with both specificity and sensitivity at 100%, and the optimal cutoff value for relative content was 0.0191; the AUC of CDH1

[152] _C was 100% (95% CI: 1-1, p<0.05), with both specificity and sensitivity at 100%, and the optimal cutoff value for relative content was 0.0745.

[0089] Based on the calculation method of method c, the differences in individual aa-score of amino acid sites in different groups were calculated using the two independent samples Wilcoxon test. Figure 6 a). The three amino acid sites PLXDC2

[90] _C, CDH1

[152] _C, and IGF2

[126] _C showed statistically significant differences between the β-thalassemia group and the healthy control group, with p-values ​​of 5.97e-13, 1.77e-12, and 5.97e-13, respectively. The diagnostic potential of the individual aa-score of the amino acid site was calculated using the ROC curve analysis method in the pROC package of R language. Figure 6 b) The AUC of PLXDC2

[90] _C was 100% (95% CI: 1-1, p<0.05), with both specificity and sensitivity at 100%, and the optimal cutoff value for relative content was 0.4466; the AUC of CDH1

[152] _C was 99.8% (95% CI: 0.988-1, p<0.05), with both specificity and sensitivity at 100.0%, and the optimal cutoff value for relative content was 0.0010; the AUC of IGF2

[126] _C was 100% (95% CI: 1-1, p<0.05), with both specificity and sensitivity at 100.0%, and the optimal cutoff value for relative content was 18.3369.

[0090] Three different individual aa-score calculation methods and ROC analysis showed that the peptide clusters at these three sites have a strong ability to distinguish between the β-thalassemia group and the healthy control group. This indicates broad application prospects in the screening and diagnosis of β-thalassemia.

[0091] Example 3: Evaluation of the reliability of detection methods for potential peptide cluster biomarkers for β-thalassemia screening and diagnosis.

[0092] 8. Preparation of plasma polypeptide samples

[0093] This example included 91 participants, comprising 40 healthy controls, 39 carriers of mild β-thalassemia, and 12 patients with severe β-thalassemia. 10 µL of plasma from each sample was collected and thoroughly mixed to serve as a correction sample for retention time correction. Plasma peptide extraction was performed using the same method as in step 1.

[0094] 9. PRM test method retention time correction

[0095] First, the scheduled PRM detection method established in Example 2 was modified by changing the retention time (RT) to an unscheduled PRM detection method; then, data was collected from the calibration samples; finally, a scheduled PRM detection method incorporating the latest RT was established based on the retention time of the precursor ion in the calibration samples, and then each sample was detected.

[0096] 10. Results Analysis

[0097] The data analysis method is the same as in step 7.3. Specifically, the individual aa-score of the two amino acid sites PLXDC2

[90] _C and CDH1

[152] _C is calculated according to method a; while for IGF2

[126] _C, since there is no relabeled polypeptide, the individual aa-score is calculated using method c.

[0098] In this cohort, the AUC of PLXDC2

[90] _C remained at 100% (95% CI: 1-1, p<0.05), with both specificity and sensitivity at 100%, and the optimal cutoff for absolute content was 0.4182 fmol / µL; the AUC of CDH1

[152] _C was 100% (95% CI: 1-1, p<0.05), with both specificity and sensitivity at 100%, and the optimal cutoff for absolute content was 4.4632 fmol / µL; the AUC of IGF2

[126] _C was 100% (95% CI: 1-1, p<0.05), with both specificity and sensitivity at 100.0%, and the optimal cutoff for relative content was 8.1008 ( Figure 7 ).

[0099] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1.A kit for diagnosing β-thalassemia, comprising reagents for detecting the expression level of any one of PLXDC2[90]_C, CDH1[152]_C, and IGF2[126]_C polypeptide clusters, wherein the specific information of the polypeptide clusters is shown in the following table The PLXDC2[90]_C is a polypeptide cluster with the C-terminal end at the 90th amino acid site of the PLXDC2 protein; the CDH1[152]_C is a polypeptide cluster with the C-terminal end at the 152th amino acid site of the CDH1 protein; and the IGF2[126]_C is a polypeptide cluster with the C-terminal end at the 126th amino acid site of the IGF2 protein. 2.The kit of claim 1, further comprising reagents commonly used in proteomics, as well as standard and / or control samples. 3.The kit of claim 1 or 2, wherein the expression level is determined by LC-MS / MS. 4.Use of the kit of any one of claims 1-3 in the preparation of a product for diagnosing β-thalassemia. 5.The use of claim 4, wherein the sample is derived from plasma. 6.The use of claim 4, comprising the following steps: (1) extracting plasma from the subject to be tested; (2) determining the expression level of any one of the PLXDC2[90]_C, CDH1[152]_C, and IGF2[126]_C polypeptide clusters; and (3) determining whether the subject has β-thalassemia by the content of the polypeptide clusters. ​ ​

Citation Information

Patent Citations

  • Chimeric antigen receptors for the treatment of cancer

    CN110225927A

  • Biomarker for beta-thalassemia disease and subtype typing diagnosis and application thereof

    CN118311262A