Early diagnosis of chronic obstructive pulmonary disease markers and their applications

By separating and enriching IgG glycopeptides in plasma/serum, and combining mass spectrometry and random forest models, the accuracy problem of early COPD diagnosis was solved, enabling efficient identification and diagnosis of high-risk individuals for early COPD.

CN117783536BActive Publication Date: 2026-02-03INSTITUTE OF BASIC MEDICAL SCIENCES CHINESE ACADEMY OF MEDICAL SCIENCES
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311770936.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-21
Publication Date
2026-02-03
Estimated Expiration
2043-12-21

AI Technical Summary

Technical Problem

Current technologies lack effective biomarkers for the early identification and diagnosis of COPD, especially when pulmonary function tests and imaging examinations are not sensitive to changes in small airway pulmonary function, making it difficult to accurately identify patients with early COPD.

Method used

Using IgG standards as molecular markers, high-abundance proteins in plasma/serum were removed by non-denaturing polyacrylamide gel electrophoresis, IgG heavy chains were separated by denaturing polyacrylamide gel electrophoresis, and IgG hydrolysates were obtained by trypsin digestion. Glycopeptides in the IgG hydrolysates were enriched using modified polydopamine magnetic nanomaterials. The relative intensities of IgG glycopeptides in serum or plasma were detected and identified by mass spectrometry. A random forest model was constructed to distinguish between high-risk individuals and healthy individuals in the early stages of COPD.

Benefits of technology

It provides a highly accurate early COPD diagnosis method. By combining immunoglobulin glycosylation modification indicators in plasma/serum with pulmonary function test indicators, it achieves accurate identification of high-risk groups for early COPD and has the advantages of low cost, high throughput and ease of operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117783536B_ABST
    Figure CN117783536B_ABST
Patent Text Reader

Abstract

The application discloses a diagnosis marker for early chronic obstructive pulmonary disease (COPD) and application thereof, and relates to the technical field of biomedicine. The application is based on the fact that the immunoglobulin glycosylation modification index in plasma / serum can be used as a marker when being combined with the lung function detection index FEV1 / FVC, so as to distinguish the high-risk population of early COPD from the healthy population. The method has the advantages of high prediction accuracy, low cost, high throughput and easy operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biomedical technology, and more specifically to early COPD diagnostic markers and their applications. Background Technology

[0002] Chronic obstructive pulmonary disease (COPD) is a common chronic respiratory disease characterized by persistent respiratory symptoms and airflow limitation. In recent years, the prevalence of COPD has been increasing annually, reaching 13.7% (age ≥40 years), severely impacting patients' quality of life. In the early stages of COPD, even individuals with normal lung ventilation function may experience recurrent acute exacerbations, structural changes in the lungs, and abnormalities in other lung function indicators. This has led to the concept of early COPD, representing the initial stage of the natural course of COPD.

[0003] Early identification, prevention, and management of COPD are key aspects of COPD control, whole-disease management, and personalized treatment, playing a positive role in early diagnosis and treatment of COPD and improving patient prognosis. Currently, international diagnostic criteria for early COPD mainly refer to the definition proposed by Martinez et al. in 2018. The diagnosis and classification of COPD are primarily based on pulmonary function test indicators such as forced expiratory volume in one second (FEV1), forced vital capacity (FVC), and FEV1 / FVC. However, routine pulmonary function tests and imaging examinations are not sensitive to changes in small airway function. Therefore, although early COPD patients are close to the onset of the disease, there is a lack of effective biomarkers for identification or quantification in clinical practice. Furthermore, COPD is a complex and heterogeneous disease, associated with multiple factors such as genetics, sex, age, smoking, environmental pollution, and comorbidities. Therefore, the Global Initiative for Chronic Obstructive Lung Disease (2023 revised edition) proposes a strategy based on "treatable characteristics," aiming to gain a deeper understanding of key causal pathways through phenotypic identification and / or through validated biomarkers. Regarding phenotypic identification, existing data have shown that elevated circulating eosinophil counts in COPD patients are associated with acute exacerbations. However, few biomarkers have been reported to predict the early development of COPD.

[0004] Therefore, in order to improve the success rate of early COPD diagnosis, developing biomarkers with high predictive accuracy and their detection and application methods is an urgent problem that needs to be solved by those skilled in the art. Summary of the Invention

[0005] In view of this, the present invention provides a biomarker for the early diagnosis of COPD. Using IgG standards as molecular markers, the inventors employed non-denaturing polyacrylamide gel electrophoresis to remove high-abundance proteins such as albumin and transferrin from plasma / serum; denaturing polyacrylamide gel electrophoresis to separate IgG heavy chains (molecular weight approximately 55 kDa); trypsin digestion to obtain IgG hydrolysates; enrichment of glycopeptides in the IgG hydrolysates using modified polydopamine magnetic nanomaterials; and mass spectrometry to detect and identify the relative intensity of the enriched IgG glycopeptides in serum or plasma. Based on this, the inventors discovered that immunoglobulin glycosylation modification indicators in plasma / serum, alone or in combination with the pulmonary function test indicator FEV1 / FVC, can serve as biomarkers to differentiate between high-risk individuals and healthy individuals in the early stages of COPD.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] The application of reagents for detecting IgG Fc N-glycopeptides in the preparation of early COPD diagnostic products, wherein the peptide sequence of the IgG Fc N-glycopeptide is the IgG1 isotype polypeptide sequence shown in SEQ ID NO: 1 or the IgG2 isotype polypeptide sequence shown in SEQ ID NO: 2, and the glycoform of the IgG Fc N-glycopeptide is selected from any one of the following tables:

[0008]

[0009]

[0010] Preferably, the glycoform of the IgG Fc N-glycopeptide is a glycosylation modification at the Asn site in the sequence shown in SEQ ID NO: 1 or SEQ ID NO: 2. In other words, the glycoform of the IgG Fc N-glycopeptide is a glycosylation modification at the Asn site in the sequence shown in SEQ ID NO: 3. 180 Glycosylation modification at the (IgG1 subtype) site, or at the Asn site shown in SEQ ID NO: 4. 176 Glycosylation modification at the (IgG2 subtype) site.

[0011] Preferably, the test sample for the early COPD diagnostic product is serum or plasma.

[0012] Preferably, the detection method for the IgG Fc N-glycopeptide includes the following steps:

[0013] (1) Using IgG standard as the marker molecule, high-abundance proteins such as albumin and transferrin in serum / plasma are removed by gradient (4% to 10%) or isocratic (7.5%) non-denaturing polyacrylamide gel electrophoresis to obtain the total IgG target band.

[0014] (2) The target band obtained in step (1) was separated by denaturing polyacrylamide gel electrophoresis, and the IgG heavy chain was obtained at a molecular weight of 55 kDa.

[0015] (3) Perform intragel enzymatic hydrolysis on the IgG heavy chain obtained in step (2) to obtain IgG protein hydrolysis product;

[0016] (4) Enrich the glycopeptides in the IgG protease hydrolysate obtained in step (3) using modified polydopamine magnetic nanomaterials;

[0017] (5) Identify the IgG Fc N-glycopeptides enriched in step (4) by mass spectrometry, including IgG1 G0F, IgG1 G0FN, IgG1G1FN, IgG2 G0FN, IgG2 G1FN, IgG2 G2FN, IgG2G2F, IgG2 G0, IgG2 G1, IgG2 G2, IgG2 G0F, and IgG2 G2F, and obtain their relative intensities. Calculate the relative intensity ratio of glycopeptides in the same IgG subtype whose Fc peptide segments have the same glycosylation modification sites and differ by one or two identical monosaccharide residues.

[0018] (6) Construct a random forest model, and output the result of whether the population is at high risk of early COPD by inputting the relative intensity ratio of glycopeptides obtained in step (5).

[0019] More preferably, the algorithm steps of the random forest model are as follows:

[0020] (1) Random sampling: Multiple random subsets are obtained from the original training dataset using a sampling method with replacement;

[0021] (2) Random feature selection: For each subset, a portion of features are randomly selected for the construction of the decision tree; preferably, if the total number of features is W, for regression problems, W / 3 features are usually randomly selected, but not less than 5 features; for classification problems, features are usually randomly selected. One feature;

[0022] (3) Decision tree construction: For each subset, a decision tree is constructed based on the selected features, and feature learning and training are performed;

[0023] (4) Integrated decision tree: The constructed decision trees are combined into a forest, and the final result is obtained by taking the average or voting. Preferably, taking the average means taking the average of the regression results of multiple decision trees as the final result, while voting is based on the principle of majority rule as the final result.

[0024] More preferably, the relative strength indices of glycopeptides input in step (6) are IgG1 G1FN to G0FN, IgG1 G0F to G0FN, IgG2 G0F to G0, IgG2 G2F to G2, IgG2 G1 to G0, IgG2 G2 to G0, IgG2 G1FN to G2FN, IgG2 G0FN to G2FN, and IgG2 G2F to G2FN;

[0025] Furthermore, an output value of 0 indicates that the individual is not at high risk for early-stage COPD, while an output value of 1 indicates that the individual is at high risk for early-stage COPD.

[0026] More preferably, the random forest model also includes the input lung function test index FEV1 / FVC.

[0027] Another objective of this invention is to provide the application of the ratio of the relative intensities of IgG Fc N-glycopeptides in the serum or plasma of a subject, or the ratio of the relative intensities of IgG Fc N-glycopeptides combined with the pulmonary function test index FEV1 / FVC, as a biomarker in the preparation of early COPD diagnostic products. The relative intensities of the glycopeptides are IgG1 G1FN to G0FN, IgG1 G0F to G0FN, IgG2 G0F to G0, IgG2 G2F to G2, IgG2 G1 to G0, IgG2 G2 to G0, IgG2 G1FN to G2FN, IgG2 G0FN to G2FN, and IgG2 G2F to G2FN.

[0028] Another object of the present invention is a system for predicting early high-risk COPD, the system comprising:

[0029] Analysis device: used to detect the expression level of IgG Fc N-glycopeptide as described in any one of claims 1-7 in the subject's biological sample, and input it into the evaluation model for predictive analysis;

[0030] Output device: Used to output the above prediction results.

[0031] Preferably, the evaluation model is a random forest model; a value of 0 in the output result indicates that the individual is not at high risk of early COPD, and a value of 1 indicates that the individual is at high risk of early COPD.

[0032] Beneficial Effects: Immune inflammatory response is one of the main characteristics of COPD. The occurrence of COPD is related to the enhanced chronic inflammation of the airways and lungs caused by the release of inflammatory factors and the expression of autoantibodies. Immunoglobulins (Ig) are key glycoproteins of the immune system, which can play an immunomodulatory and anti-inflammatory role. Their glycosylation modification affects the structure, effector function, and anti-inflammatory activity of Ig, and is closely related to the pathological state of the body. This invention, based on plasma / serum immunoglobulin glycosylation modification indicators or combined with the pulmonary function test indicators FEV1 / FVC, is used to differentiate between high-risk and healthy individuals in the early stages of COPD. It has the advantages of high predictive accuracy, low cost, high throughput, and ease of operation. Attached Figure Description

[0033] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0034] Figure 1 The representative mass spectra of serum IgG glycopeptides are shown below (dashed lines represent peptide sequences SEQ ID NO: 1 (Glu GluGln Tyr Asn)). 180 The solid line represents the IgG1 glycopeptide shown in Ser Thr Tyr Arg; the solid line represents the peptide sequence SEQ ID NO: 2 (Glu Glu Gln Phe Asn). 176 Ser Thr Phe Arg) IgG2 glycopeptide; ■(N), N-acetylglucosamine; ○(M), mannose; ○(G), galactose; △(F), fucose).

[0035] Figure 2 To determine the ROC curve of early COPD using a random forest model, we consider: (a) FEV1 / FVC, (b) IgG carbohydrate index, and (c) IgG carbohydrate index combined with FEV1 / FVC. Detailed Implementation

[0036] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0037] Reagent preparation in this embodiment of the invention:

[0038] (1) Preparation of Tris-boric acid-magnesium chloride (TBM) solution (stored at room temperature): Weigh 10.8g Tris, 5.6g boric acid and 0.3g magnesium chloride hexahydrate and dissolve in 200mL ultrapure water.

[0039] (2) Prepare 10×native-PAGE electrophoresis buffer (store at room temperature): Weigh 144g glycine and 30.2g Tris and dissolve them in 1L ultrapure water for later use. Dilute with ultrapure water to prepare 1×native-PAGE electrophoresis buffer for use during electrophoresis.

[0040] (3) Prepare 0.2M Tris-HCl (pH 7.5) (store at 4℃): Weigh 4.85g Tris and 2.02g magnesium chloride hexahydrate and dissolve them in ultrapure water. Add about 2.5mL HCl to adjust the pH to 7.5, and then make up to 200mL with ultrapure water.

[0041] (4) 5×native-PAGE loading buffer (long-term storage at -20℃, short-term storage at 4℃): Weigh 2.5mL of 0.2MTris-HCl (pH 7.5) and 5mL of glycerol and mix with ultrapure water. Add a trace amount of xylene cyanol FF and bring the volume up to 10mL with ultrapure water. Store at 4℃.

[0042] (5) 7.5% native-PAGE separating gel: prepared from 53.9% (v / v) ultrapure water, 25% (v / v) 30% polyacrylamide (ACR) solution, 20% (v / v) TBM solution, 1% 10% APS solution and 0.05% TEMED.

[0043] (6) 4% native-PAGE stacking gel: prepared from 75.1% (v / v) ultrapure water, 13.4% (v / v) 30% ACR solution, 10.1% (v / v) 0.2M Tris-HCl (pH=7.5), 1.3% (v / v) 10% APS solution and 0.1% (v / v) TEMED.

[0044] (7) 8% SDS-PAGE separating gel: prepared from 46.3% (v / v) ultrapure water, 26.7% (v / v) 30% ACR solution, 25.0% (v / v) 1.5M Tris-HCl (pH=8.8), 1.0% (v / v) 10% SDS solution, 1.0% (v / v) 10% APS solution and 0.04% (v / v) TEMED.

[0045] (8) 5% SDS-PAGE stacking gel: prepared from 69.2% (v / v) ultrapure water, 16.3% (v / v) 30% ACR solution, 12.4% (v / v) 1M Tris-HCl (pH=6.8), 1.0% (v / v) 10% SDS solution, 1.0% (v / v) 10% APS solution and 0.1% (v / v) TEMED.

[0046] (9) 0.2M dithiothreitol (DTT) solution: Weigh 0.30g DTT, dissolve it in equilibration solution and make up to 10mL. Prepare fresh before use.

[0047] (10) 0.3M iodoacetamide (IAA) solution: Weigh 0.6g IAA, dissolve it in equilibration solution and make up to 10mL. Prepare fresh before use.

[0048] (11) 25mM ammonium bicarbonate solution: Weigh 0.5g of ammonium bicarbonate, dissolve it in deionized water and make up to 250mL.

[0049] (12) 50% ACN decolorizing solution: ACN and 25mM ammonium bicarbonate solution are mixed in equal volume ratio.

[0050] (13) 10 μg / μL human IgG standard: Weigh 10 mg of human IgG standard powder, dissolve it in ultrapure water and make up to 1 mL.

[0051] (14) 12.5 ng / μL pancreatic enzyme solution: Dilute the sequencing grade pancreatic enzyme stock solution with 25 mM ammonium bicarbonate solution to a final concentration of 12.5 ng / μL working solution.

[0052] (15) 10 mg / mL 2,5-dihydroxybenzoic acid (DHB) matrix solution: Weigh 10 mg DHB, dissolve it in working solution (ACN / H2O / TFA=49.95 / 49.95 / 0.1, v / v / v) and make up to 1 mL.

[0053] (16) Preparation method of modified polydopamine magnetic nanomaterials: Weigh 1.35g of ferric chloride hexahydrate into a beaker, add 75mL of ethylene glycol, and stir continuously at room temperature until the ferric chloride hexahydrate is completely dissolved. Then slowly add 3.6g of sodium acetate and stir continuously at room temperature until completely dissolved. Transfer the resulting solution to a reaction vessel and heat in an oven at 200℃ for 16h, then cool naturally to obtain Fe3O4 material. After thoroughly washing the Fe3O4 material with anhydrous ethanol and ultrapure water alternately, dry it in an oven at 60℃ until the material is completely dry to obtain iron oxide nanomagnetic beads. 20 mg of dopamine was dissolved in 10 mL of 10 mM Tris-HCl (pH 8.8), and Fe3O4 magnetic nanomaterials were added. The mixture was stirred in the dark for 8-24 h. The product was washed and dried in an oven at 60 °C to obtain polydopamine magnetic beads. The washing method was to wash three times with N,N-dimethylformamide and water alternately, and twice with N,N-dimethylformamide. 16 μL of diethylenetriamine was dissolved in 2 mL of N,N-dimethylformamide, and 65 mg of N,N-carbonyldiimidazole was added. The mixture was shaken at room temperature for 10-60 minutes. 10 mg of polydopamine nanomaterial magnetic beads were dispersed in 1 mL of N,N-dimethylformamide, and 8 μL of diethylenetriamine and 1 mL of diethylenetriamine carbamoylimidazole were added. The mixture was shaken at room temperature for 4-20 hours. After magnetic separation, the supernatant was discarded, and the mixture was washed three times with N,N-dimethylformamide to obtain polyurea-modified polydopamine magnetic beads. The beads were then dispersed in water for at least 3 hours.

[0054] Example 1: Establishment of an early diagnostic method for COPD

[0055] 1. Isolation of human serum / plasma IgG

[0056] 2 μL of serum / plasma and native-PAGE loading buffer were thoroughly mixed and loaded onto a 7.5% isocratic native-PAGE gel. Electrophoresis was performed at a constant voltage of 100V for 15 min, followed by constant current electrophoresis at 200V for 45 min. After native-PAGE, the gel was stained with rapid protein electrophoresis gel staining solution for 40 min, and then repeatedly destained with ultrapure water until the gel background was colorless. Human IgG standard was used as a molecular marker to observe the migration pattern of IgG in native-PAGE, with a loading amount of 10 μg.

[0057] Alternatively, mix 2 μL of serum / plasma thoroughly with native-PAGE loading buffer, load the mixture onto a 4-10% gradient native-PAGE gel, and perform electrophoresis at 10 mA constant current for 2.5 h.

[0058] 2. Obtain IgG heavy chain

[0059] Cut gel strips containing IgG and place them in 12-well plates. Add 3 mL of 0.2 M DTT solution and incubate at 37 °C for 1 h for reduction. Discard the DTT solution and wash the strips three times with ultrapure water. Then add 3 mL of 0.5 M IAA solution and incubate at room temperature in the dark for 1 h for alkylation. Place the processed native-PAGE strips on a glass plate containing 8% SDS-PAGE separating gel, with the lower edge of the strip approximately 1 cm from the upper edge of the separating gel. Then pour in a 5% SDS-PAGE stacking gel. After the stacking gel solidifies, perform electrophoresis at 100 V for 40 min, then adjust to 200 V for 40 min. Subsequently, stain with protein electrophoresis PAGE gel rapid staining solution for 40 min, and then repeatedly destain with ultrapure water until the gel background is colorless.

[0060] 3. Obtain IgG protease hydrolysate

[0061] The SDS-PAGE band corresponding to the IgG heavy chain was excised [at molecular weight 55 kDa, i.e., the IgG heavy chain, with amino acid sequences as shown in SEQ ID NO: 3 (IgG1 subtype) or SEQ ID NO: 4 (IgG2 subtype)]. The band was then shredded into approximately 0.1 mm pieces using a grinder. 3 The gel particles were placed in centrifuge tubes, and 50% acetonitrile decolorizing solution (prepared with 25 mM ammonium bicarbonate solution) was added and shaken on a mixer for decolorization. The decolorizing solution was discarded, and the gel particles were completely dehydrated with ACN and then freeze-dried under vacuum. 20 μL of 12.5 ng / μL trypsin solution was added dropwise to the dried gel particles, and the mixture was incubated at 4 °C for 1 h. Then, 40 μL of 25 mM ammonium bicarbonate solution was added, and the enzymatic hydrolysis reaction was carried out at 37 °C for 16 h. 140 μL of water was added to each tube and mixed thoroughly. The obtained enzymatic hydrolysis product was collected and freeze-dried under vacuum for later use.

[0062] 4. Enrichment of IgG Fc N-glycopeptides in IgG protease hydrolysates

[0063] Modified polydopamine magnetic nanomaterials were washed three times with an enrichment solution of 85% acetonitrile (acetonitrile / water / formic acid = 85 / 14.9 / 0.1, v / v / v). The enriched material was then redispersed in the enrichment solution to prepare a final concentration of 2 mg / mL. 100 μL of the enrichment solution was added to dried IgG enzymatic hydrolysate, vortexed to dissolve, and then agitated on a mixer to adsorb glycopeptides (25℃, 1000 rpm, 1 h). The magnetic nanomaterials adsorbed with glycopeptides were magnetically separated using a magnet, the enrichment solution was discarded, and the nanomaterials were washed three times with 100 μL of 85% acetonitrile, vortexed for 20 s, magnetically separated again, and the washing solution was discarded. This process was repeated three times. The washed magnetic nanomaterials were dispersed in 100 μL of water and agitated on a mixer to elute the glycopeptides (25℃, 1000 rpm, 45 min). The magnetic nanomaterials were magnetically separated and discarded, and the eluent was collected and freeze-dried under vacuum for later use.

[0064] 5. Identify the glycoform of IgG Fc N-glycopeptides and calculate the ratio of relative intensities of glycopeptides differing by one or two identical monosaccharide residues at the same glycosylation modification site in the Fc peptide segment of the same IgG subtype.

[0065] The glycopeptide sample obtained in the previous step was redissolved in 5 μL of deionized purified water. 0.25 μL of the reconstituted solution was spotted onto the MTPancherChip™ target. After evaporation, 0.25 μL of 10 mg / mL 2,5-dihydroxybenzoic acid solution was applied over the sample residue. After natural evaporation, mass spectrometry analysis was performed. SolariX FTICRMS was used for detection. The mass spectrometry detection mode was positive ion mode, the laser power was set to 50%, and the acquisition range was set to 2000–5000 Da.

[0066] Mass spectrometry data were processed using DataAnalysis 4.4 software. Data acquisition: Mass spectrum information (mass-to-charge ratio and relative intensity) was obtained using the software, with selection criteria including signal-to-noise ratio ≥3, relative intensity >0.1%, absolute intensity >100,000, and reliable isotope peaks. Results were imported into Microsoft Excel for analysis. The mass and signal intensity of single isotope peaks of glycopeptides with a signal-to-noise ratio greater than 3.0 were extracted. The ratio of relative intensities of glycopeptides differing by one or two identical monosaccharide residues at the same glycosylation modification site in the Fc peptide segment of the same IgG subtype was calculated.

[0067] 6. Early COPD Prediction Based on Random Forest

[0068] Random forest is a machine learning method based on the Bagging ensemble algorithm. Conventional decision tree methods complete tasks such as classification and regression through feature selection, subset partitioning, and recursive tree structure construction. However, decision tree methods suffer from drawbacks such as overfitting and predictive randomness. Random forest introduces randomness and ensemble thinking into decision trees, solving these common problems. The Bagging algorithm is an ensemble learning algorithm. Its specific principle is as follows: From a dataset of size N, a subset of size M is selected with replacement, K times in total. These K subsets are used to train K models. Finally, the K models are used for prediction, and the final result is obtained by averaging or voting.

[0069] The algorithm steps for Random Forest are as follows:

[0070] (1) Random sampling: Multiple random subsets are obtained from the original training dataset using a sampling method with replacement;

[0071] (2) Random Feature Selection: For each subset, a portion of features are randomly selected for constructing the decision tree. Specifically, if the total number of features is W, for regression problems, W / 3 features are often randomly selected, but no fewer than 5 features. For classification problems, features are often randomly selected... One feature;

[0072] (3) Decision tree construction: For each subset, a decision tree is constructed based on the selected features, and feature learning and training are performed;

[0073] (4) Ensemble Decision Tree: The constructed decision trees are arranged into a forest, and the final result is obtained by averaging or voting. Specifically, averaging means taking the average of the regression results of multiple decision trees as the final result, while voting is based on the principle of majority rule.

[0074] Specifically, the input and output of each decision tree are determined. The input of each decision tree consists of various indicators of serum / plasma IgG glycosylation modification and FEV1 / FVC. The output of the random forest is whether the subject is at high risk for early COPD. An output of 0 indicates that the subject is not at high risk for early COPD, and an output of 1 indicates that the subject is at high risk for early COPD.

[0075] The specific process is as follows: a random forest model is built using the RandomForestClassifier class in the Python library sklearn, with the parameter n_estimators=80. The model is fitted to the input training dataset and training labels. Finally, the trained model is used on the test dataset to obtain the predicted labels. The model performance is evaluated by comparing the predicted labels with the test labels.

[0076] Further preferably, in a specific embodiment of the present invention, IgG1 G0F, IgG1 G0FN, IgG1 G1FN, IgG2 G0FN, IgG2 G1FN, IgG2 G2FN, IgG2 G2F, IgG2 G0, IgG2 G1, IgG2 G2, IgG2 G0F, and IgG2 G2F glycopeptides are detected in the patient sample to be tested, and their relative intensities are obtained;

[0077] Further optimized, the subject's serum / plasma IgG glycosylation modification index input for each decision tree:

[0078] IgG1 G1FN vs G0FN, IgG1 G0F vs G0FN, IgG2 G0F vs G0, IgG2G2F vs G2, IgG2 G1 vs G0, IgG2 G2 vs G0, IgG2 G1FN vs G2FN, IgG2 G0FN vs G2FN, IgG2 G2F vs G2FN.

[0079] Furthermore, each decision tree also inputs the subject's lung function test index FEV1 / FVC.

[0080] Example 2: Obtaining Plasma IgG Glycotype Indicators

[0081] 1. Isolation, enzymatic digestion, and enrichment of plasma IgG glycopeptides

[0082] Take 2 μL of the plasma to be tested, and separate the IgG heavy chain (~55 kDa) sequentially using 7.5% non-denaturing polyacrylamide gel electrophoresis and 8% sodium dodecyl sulfonate denaturing polyacrylamide gel electrophoresis. Stain with Coomassie brilliant blue overnight, and then remove the background color with water.

[0083] Cut off the gel dots of the IgG heavy chain, crushing them to approximately 0.1 mm. 3 Add 50% acetonitrile / 25mM ammonium bicarbonate for decolorization. Decolorization is complete after about 1 hour. Discard the supernatant. Add 100% acetonitrile to dehydrate and harden the gel block. Discard the supernatant and let the gel block stand until dry. Add 20 μL of trypsin solution (12.5 ng / μL, prepared with 25mM ammonium bicarbonate) and place at 4℃ for about 1 hour to fully hydrate the gel block. Add 25mM ammonium bicarbonate until the gel surface is submerged, and react at 37℃ for about 16 hours. The enzymatic hydrolysis product is then freeze-dried under vacuum for later use.

[0084] Modified polydopamine magnetic nanomaterials were washed three times with an enrichment solution of 85% acetonitrile (acetonitrile / water / formic acid = 85 / 14.9 / 0.1, v / v / v). The enriched material was then redispersed in the enrichment solution to prepare a final concentration of 2 mg / mL. 100 μL of the enrichment solution was added to dried IgG enzymatic hydrolysate, vortexed to dissolve, and then agitated on a mixer to adsorb glycopeptides (25℃, 1000 rpm, 1 h). The magnetic nanomaterials adsorbed with glycopeptides were magnetically separated using a magnet, the enrichment solution was discarded, and the nanomaterials were washed three times with 100 μL of 85% acetonitrile, vortexed for 20 s, magnetically separated again, and the washing solution was discarded. This process was repeated three times. The washed magnetic nanomaterials were dispersed in 100 μL of water and agitated on a mixer to elute the glycopeptides (25℃, 1000 rpm, 45 min). The magnetic nanomaterials were magnetically separated and discarded, and the eluent was collected and freeze-dried under vacuum for later use.

[0085] 2. Detection and identification of IgG glycopeptides

[0086] (1) The above glycopeptide sample was redissolved in 5 μL of deionized purified water. 0.25 μL of the redissolved solution was spotted onto the MTPancherChip™ target. After evaporation, 0.25 μL of 10 mg / mL 2,5-dihydroxybenzoic acid solution was applied over the sample residue. After natural evaporation, mass spectrometry analysis was performed. Detection was performed using a SolariX FTICR MS. The mass spectrometry detection mode was positive ion mode, the laser power was set to 50%, and the acquisition range was set to 2000–5000 Da.

[0087] (2) Acquisition of mass spectrometry data: The data information (mass-to-charge ratio and relative intensity) in the mass spectra were acquired using software. The selection criteria were a signal-to-noise ratio ≥3, a relative intensity >0.1%, an absolute intensity >100,000, and the presence of reliable isotope peaks. The results were imported into Microsoft Excel for analysis.

[0088] (3) Identification of glycopeptides: High-resolution mass spectrometry obtains accurate molecular mass information (observed value), and the information provided by the database is used to infer the corresponding possible glycopeptides and calculate their theoretical molecular weight. If the error between the observed value and the theoretical value obtained by mass spectrometry detection is within 20 ppm, they are considered to be the same substance.

[0089] (4) Obtaining IgG glycoform indices: Calculate the relative intensity ratio of glycopeptides with the same glycosylation modification site in the Fc peptide of the same IgG subtype that differ by one or two identical monosaccharide residues.

[0090] This example provides the plasma IgG glycopeptide mass spectrometry results of a patient with early-stage COPD (see Appendix). Figure 1 ). In the appendix Figure 1 The table below identifies 11 glycopeptide types corresponding to the 8 glycoforms G0, G0F, G1, G2, G2F, G0FN, G1FN, and G2FN involved in this invention (see Table 1 below). These glycopeptides are all N-glycopeptides, with the glycan chain linked to the Asn of the polypeptide chain. Specifically, the N-glycosylation modification site is located on the Asn of the IgG1 subtype polypeptide sequence shown in SEQ ID NO: 1 (as shown in SEQ ID NO: 3, at position 180 of the IgG1 heavy chain constant region amino acid sequence); or, on the Asn of the IgG2 subtype polypeptide sequence shown in SEQ ID NO: 2 (as shown in SEQ ID NO: 4, at position 176 of the IgG2 heavy chain constant region amino acid sequence).

[0091] Table 1

[0092]

[0093]

[0094] ■(N), N-acetylglucosamine; ○(M), mannose; ○(G), galactose; △(F), fucose

[0095] The IgG molecule has a conserved N-glycosylation modification site. The glycans attached to this site are mainly complex N-glycans. The complex N-glycans attached to Fc generally have a common hepta-glycan basic structure GlcNAc2-Man3-GlcNAc2 (see Epp A, Sullivan KC, Herr AB, et al., Immunoglobulin Glycosylation Effects in Allergy and Immunity[J]. Current Allergy & Asthma Reports, 2016, 16(11):79). Based on this common hepta-glycan basic structure, Fuc, Gal, and the subtype GlcNAc can be further attached. For example, in this invention, IgG2 G0F refers to Asn in the constant region of the IgG2 heavy chain. 176 The site connects to a basic heptaglycose structure (GlcNAc2-Man3-GlcNAc2), followed by a fucose residue, as shown in the diagram. As shown; IgG2G1F refers to the Asn in the IgG2 peptide. 176 The site connects to a basic heptaose structure (GlcNAc2-Man3-GlcNAc2), followed by a galactose residue and a fucose residue, as shown in the diagram. As shown; the remaining sugar types are represented in a similar manner.

[0096] That is, the glycoforms of the present invention have the glycosylation modification patterns shown in Table 2:

[0097] Table 2

[0098]

[0099] Note: GlcNAc2-Man3-GlcNAc2 represents the basic structure of IgG Fc N-glycans, GlcNAc represents N-acetylglucosamine residues, Man represents mannose residues, Fuc(F) represents fucose residues, Gal(G) represents galactose residues, and the numbers represent the number of monosaccharide residues. (See Epp A, Sullivan KC, Herr AB, et al. Immunoglobulin Glycosylation Effects in Allergy and Immunity[J]. Current Allergy & Asthma Reports, 2016, 16(11):79).

[0100] Example 3: Construction and Validation of a Random Forest Model for Predicting Early-Stage COPD

[0101] 1. Construction of a random forest model for predicting early-stage COPD

[0102] Plasma was collected from 50 healthy volunteers and 50 volunteers with early-stage COPD. Basic information of the healthy volunteers and early-stage COPD patients is shown in Table 3. Plasma IgG glycopeptide levels were obtained according to the method in Example 1. The Mann-Whitney U test was used to analyze the expression differences of various glycopeptide levels and the pulmonary function test index FEV1 / FVC between healthy individuals and early-stage COPD patients. It was found that IgG1 G1FN / G0FN, IgG1 G0F / G0FN, IgG2 G0F / G0, IgG2G1 / G0, IgG2 G2 / G0, IgG2 G1FN / G2FN, IgG2 G0FN / G2FN, and IgG2 G2F / G2FN were significantly higher in early-stage COPD patients than in healthy individuals; the pulmonary function test index FEV1 / FVC was significantly lower in early-stage COPD patients than in healthy individuals. Therefore, a machine learning model using plasma IgG glycopeptide levels and / or combined with FEV1 / FVC is proposed to predict high-risk groups for early-stage COPD.

[0103] Table 3

[0104]

[0105]

[0106] Using the population cohorts shown in Table 3 as the training set, a random forest model was constructed according to the method described in Example 1. The input training dataset consisted of the ratio of relative intensities of plasma IgG glycopeptides and / or the pulmonary function test index FEV1 / FVC. The input training label was whether the patient had early COPD as determined in the patient's basic information registration form. A random forest prediction model for early COPD was established with a 100% accuracy rate.

[0107] 2. Validation of the random forest model for predicting early COPD

[0108] To verify the predictive accuracy of the established random forest model for early COPD, plasma samples from 50 healthy volunteers and 50 early COPD patients were collected. Plasma IgG glycopeptide indexes were obtained according to the method in Example 1. The basic information and glycopeptide index data of the healthy volunteers and early COPD patients are shown in Table 4.

[0109] Table 4

[0110]

[0111] The population cohort shown in Table 4 is set as the test set. Using the method in Example 1 and the random forest model constructed from the test set data, the input is the ratio of the relative intensities of the predetermined serum / plasma IgG glycosylation modification index (glycoform index) and / or the pulmonary function test index FEV1 / FVC. The output indicates whether the subject is at high risk for early COPD. An output of 0 indicates that the subject is not at high risk for early COPD, and an output of 1 indicates that the subject is at high risk for early COPD. When FEV1 / FVC was input, the random forest model output 72 normal individuals and 28 individuals at high risk of early COPD, of which 49 normal individuals and 27 individuals at high risk of early COPD were correctly identified. When the IgG carbohydrate indices IgG1 G1FN / G0FN, IgG1 G0F / G0FN, IgG2 G0F / G0, IgG2 G2F / G2, IgG2 G1 / G0, IgG2 G2 / G0, IgG2 G1FN / G2FN, IgG2G0FN / G2FN, and IgG2 G2F / G2FN were input, the random forest model output 54 normal individuals and 46 individuals at high risk of early COPD, of which 48 normal individuals and 46 individuals at high risk of early COPD were correctly identified. When the IgG carbohydrate indices IgG1 G1FN / G0FN, IgG1G0F / G0FN, IgG2 G0F / G0, IgG2G2F / G2, and IgG2 G1 / G0 were input, the random forest model output 54 normal individuals and 46 individuals at high risk of early COPD, of which 48 normal individuals and 46 individuals at high risk of early COPD were correctly identified. When the ratios were G2 / G0, IgG2 G1FN / G2FN, IgG2 G0FN / G2FN, IgG2 G2F / G2FN, and FEV1 / FVC, the random forest model output 51 normal individuals and 49 individuals at high risk of early COPD. Among them, the 49 normal individuals and 48 individuals at high risk of early COPD were correctly identified.

[0112] Receiver Operating Curve (ROC) analysis was performed on the prediction results of the above random forest model (see Appendix). Figure 2 ). As attached Figure 2 As shown in Figure a, when FEV1 / FVC is input, the area under the ROC curve (AUC) of the random forest model for early COPD is 0.762, specificity is 0.98, sensitivity is 0.54, and accuracy is 0.76; see attached figure. Figure 2 As shown in b, when the IgG carbohydrate indices IgG1 G1FN / G0FN, IgG1 G0F / G0FN, IgG2 G0F / G0, IgG2G2F / G2, IgG2 G1 / G0, IgG2 G2 / G0, IgG2 G1FN / G2FN, IgG2G0FN / G2FN, and IgG2 G2F / G2FN are input, the area under the ROC curve for early COPD using the random forest model is 0.943, specificity is 0.96, sensitivity is 0.88, and accuracy is 0.92; see attached. Figure 2As shown in Figure c, when the input IgG carbohydrate indices IgG1 G1FN / G0FN, IgG1G0F / G0FN, IgG2 G0F / G0, IgG2 G2F / G2, IgG2 G1 / G0, IgG2G2 / G0, IgG2 G1FN / G2FN, IgG2 G0FN / G2FN, IgG2 G2F / G2FN and FEV1 / FVC are used, the area under the ROC curve for early COPD identification by the random forest model is 0.997, specificity is 0.98, sensitivity is 0.96, and accuracy is 0.97. This indicates that the random forest model characterized by IgG carbohydrate indices and FEV1 / FVC has a better ability to distinguish early COPD.

[0113] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0114] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. The application of the ratio of the relative intensities of IgG Fc N-glycopeptides in the plasma of the subject, or the ratio of the relative intensities of IgG Fc N-glycopeptides combined with the pulmonary function test index FEV1 / FVC, as a biomarker in the preparation of early COPD diagnostic products, characterized in that... The relative strength indices of the IgG Fc N-glycopeptides are IgG1 G1FN to G0FN, IgG1 G0F to G0FN, IgG2 G0F to G0, IgG2 G2F to G2, IgG2 G1 to G0, IgG2 G2 to G0, IgG2 G1FN to G2FN, IgG2 G0FN to G2FN, and IgG2 G2F to G2FN; The peptide sequence of the IgG Fc N-glycopeptide is the IgG1 isotype polypeptide sequence shown in SEQ ID NO: 1 or the IgG2 isotype polypeptide sequence shown in SEQ ID NO: 2, and the glycoform of the IgG Fc N-glycopeptide is a glycosylation modification at the Asn site in the sequence shown in SEQ ID NO: 1 or SEQ ID NO:

2. The glycoform of the IgG Fc N-glycopeptide is selected from any one of the following: 。 2. The application according to claim 1, characterized in that, The detection method for IgG Fc N-glycopeptide includes the following steps: (1) Using IgG standard as the marker molecule, albumin and transferrin, which are high-abundance proteins in plasma, were removed by non-denaturing polyacrylamide gel electrophoresis to obtain the total IgG target band; (2) The target band obtained in step (1) was separated by denaturing polyacrylamide gel electrophoresis, and the IgG heavy chain was obtained at a molecular weight of 55 kDa. (3) Perform intragel enzymatic hydrolysis on the IgG heavy chain obtained in step (2) to obtain IgG protease hydrolysis products; (4) Enrich the glycopeptides in the IgG protease hydrolysate obtained in step (3) using modified polydopamine magnetic nanomaterials; (5) Identify the IgG Fc N-glycopeptides enriched in step (4) by mass spectrometry, including IgG1 G0F, IgG1 G0FN, IgG1 G1FN, IgG2 G0FN, IgG2 G1FN, IgG2 G2FN, IgG2 G2F, IgG2 G0, IgG2 G1, IgG2 G2, IgG2 G0F, and IgG2 G2F, and obtain their relative intensities. Calculate the relative intensity ratio of glycopeptides in the same IgG subtype whose Fc peptide segments differ by one or two identical monosaccharide residues at the same glycosylation modification site. (6) Construct a random forest model, and output the result of whether the population is at high risk of early COPD by inputting the relative intensity ratio of glycopeptides obtained in step (5).

3. The application according to claim 2, characterized in that, The algorithm steps of the random forest model are as follows: (1) Random sampling: Multiple random subsets are obtained from the original training dataset using a sampling method with replacement; (2) Random feature selection: For each subset, a portion of features are randomly selected for the construction of the decision tree; If the total number of features is W, for regression problems, W / 3 features are usually randomly selected, but no fewer than 5 features. For classification problems, features are usually randomly selected. One feature; (3) Decision tree construction: For each subset, a decision tree is constructed based on the selected features, and feature learning and training are performed; (4) Integrated decision tree: The constructed decision trees are combined into a forest, and the final result is obtained by taking the average or voting. Taking the average means taking the average of the regression results of multiple decision trees as the final result, while voting is based on the principle of majority rule as the final result.

4. The application according to claim 2, characterized in that, The relative strength indices of IgG Fc N-glycopeptides input in step (6) are IgG1 G1FN to G0FN, IgG1 G0F to G0FN, IgG2 G0F to G0, IgG2 G2F to G2, IgG2 G1 to G0, IgG2 G2 to G0, IgG2 G1FN to G2FN, IgG2 G0FN to G2FN, and IgG2 G2F to G2FN; Furthermore, an output value of 0 indicates that the individual is not at high risk for early-stage COPD, while an output value of 1 indicates that the individual is at high risk for early-stage COPD.

5. The application according to claim 2, characterized in that, The random forest model also includes input lung function test indicators FEV1 / FVC.

6. A system for predicting early high-risk COPD, characterized in that, The system includes: Analysis device: used to detect the expression level of IgG Fc N-glycopeptide as described in any one of claims 1-5 in the subject's biological sample, and input it into the evaluation model for predictive analysis; Output device: Used to output the results of the above predictive analysis.

7. The system for predicting early high-risk COPD according to claim 6, characterized in that, The evaluation model is a random forest model; in the results, a value of 0 indicates that the individual is not at high risk of early COPD, and a value of 1 indicates that the individual is at high risk of early COPD.

Citation Information

Patent Citations

  • Data analysis system for predicting chronic obstructive pulmonary disease high-risk group

    CN117198491A