Polypeptide markers based on autoantibodies and their use in the early diagnosis of lung cancer

By screening 14 peptide probes and establishing a classification model based on signal value accumulation, the problem of insufficient sensitivity and specificity in the early diagnosis of lung cancer in existing technologies has been solved, and efficient early diagnosis of lung adenocarcinoma has been achieved.

CN115785214BActive Publication Date: 2026-02-10PEKING UNION MEDICAL COLLEGE HOSPITAL +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211203457.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-29
Publication Date
2026-02-10
Estimated Expiration
2042-09-29

AI Technical Summary

Technical Problem

In the early diagnosis of lung cancer, existing technologies based on ELISA for detecting autoantibody combinations have limited sensitivity and specificity, and the detection techniques are complex and cumbersome, making it difficult to discover new autoantibodies with diagnostic efficacy superior to known antigens.

Method used

By studying plasma samples from early-stage lung adenocarcinoma and healthy individuals, 14 peptide probes were identified and screened. The sum of the signal values ​​of these peptide probes was used as an indicator, and a classification model was established using machine learning methods. The classification threshold was obtained using Youden's J statistic method, enabling effective diagnosis of early-stage lung adenocarcinoma.

Benefits of technology

It improves the sensitivity and specificity of early lung adenocarcinoma diagnosis, and achieves low-cost and effective disease risk prediction. The sensitivity is 67%, the specificity is 84%, and the AUC reaches 0.79, which is superior to existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115785214B_ABST
    Figure CN115785214B_ABST
Patent Text Reader

Abstract

The application discloses polypeptide markers based on autoantibodies and application thereof in early lung cancer diagnosis, and relates to the technical field of biological detection. The application finds that there are 143 polypeptide probes with significant signal differences between a lung adenocarcinoma group in early stage and a healthy population group through research and analysis of plasma samples of the two groups, and further screens 14 polypeptide probes that can produce more stable signal differences between the two groups. Through attempts of various machine learning methods, it is found that a classification model can be established by taking the signal value accumulation of the 14 polypeptide probes as an index for model construction, thereby providing a path for effective diagnosis of early lung adenocarcinoma.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biodetection technology, and more specifically, to peptide biomarkers based on autoantibodies and their application in the early diagnosis of lung cancer. Background Technology

[0002] Lung cancer ranks second in global incidence among malignant tumors and remains the leading cause of cancer death. Lung adenocarcinoma is the most common pathological type, accounting for approximately 40% of all pathological classifications. Early detection can reduce lung cancer-related mortality and improve prognosis, further promoting the development of new, easily detectable blood biomarkers. Tumor cells, due to gene mutations, abnormal protein structure, abnormal modifications, or abnormal expression levels, may release abnormal proteins, including neoantigens, thereby triggering an immune response and producing autoantibodies through the clonal proliferation of immune B cells. Compared to autoantigens, autoantibodies are more stable and, amplified by the immune response, are easier to detect.

[0003] Currently, numerous studies have reported the detection of autoantibodies in the early stages of various cancers, often preceding clinical signs or imaging findings. Therefore, tumor-associated autoantibodies are among the promising tumor biomarkers. Because blood samples are readily available, and autoantibody signals are richer and less variable compared to tumor proteins, these easily detectable biomarkers are attractive for early diagnostic applications.

[0004] Based on the one-to-one recognition model of antibodies against antigens, both domestically and internationally, there are already ELISA methods for detecting autoantibody combinations targeting 6 or 7 protein antigens for lung cancer diagnosis. The earliest study in Europe and the United States in 2010 reported that using ELISA to detect autoantibodies against 6 tumor-associated antigens, including p53, CAGE, NY-ESO-1, Annexin 1, GBU4-5, and SOX2, the sensitivity in lung cancer diagnosis was 36%-39%, the specificity was about 90%, and the AUC value was 0.63-0.71 (Reference 1: Boyle, P, et al. (2011). Clinical validation of an autoantibody test for lung cancer. Ann Oncol. 22(2):383-389.). The team subsequently adjusted the autoantibody combination and finally developed a commercial ELISA kit with 7 autoantibody combinations (p53, NY-ESO-1, CAGE, GBU4-5, SOX2, HuD, and MAGE A4). Multiple studies have shown that, It can assist CT scans in improving the assessment of the risk of pulmonary nodules.

[0005] Similarly, the composition of domestically produced commercial ELISA kits for seven autoantibodies against lung cancer (p53, CAGE, GBU4-5, GAGE ​​7, SOX2, PGP9.5, and MAGE A1) is similar to... Similarly, its performance in the Chinese population was 61% sensitivity and 90% specificity (as indicated by the Keboro autoantibody kit). Many other autoantibody studies have also explored the performance of these two classes of seven autoantibody combinations in different populations using ELISA.

[0006] All of the above studies on autoantibodies have focused on detecting autoantibodies against known common cancer-associated proteins or embryonic testis antigens using ELISA methods. There are relatively few studies exploring unknown autoantibodies with diagnostic efficacy superior to those already reported. This is related to the fact that the experimental procedures for discovering new autoantibody detection technologies, such as serological proteome analysis (SERPA), serological analysis of recombinantly expressed cDNA clones (SEREX), and phage display technology, are cumbersome and complex.

[0007] In view of this, the present invention is proposed. Summary of the Invention

[0008] The purpose of this invention is to provide peptide biomarkers based on autoantibodies and their application in the early diagnosis of lung cancer.

[0009] This invention is implemented as follows:

[0010] In a first aspect, embodiments of the present invention provide a polypeptide combination comprising: polypeptides with sequences as shown in SEQ ID No. 1 to 14.

[0011] Secondly, embodiments of the present invention provide the application of polypeptide combinations in the preparation of products for the diagnosis or auxiliary diagnosis of early lung adenocarcinoma, wherein the polypeptide combination is the polypeptide combination described in the foregoing embodiments.

[0012] Thirdly, embodiments of the present invention provide a product comprising the polypeptide combination described in the foregoing embodiments.

[0013] Fourthly, embodiments of the present invention provide a device for predicting early-stage lung adenocarcinoma, comprising an acquisition module and a prediction module. The acquisition module is used to acquire the sum of detection results of a combination of peptide probes in a sample to be tested; the combination of peptide probes is the combination of peptides described in the foregoing embodiments, and the detection results include the detection signal value of the peptide probes or the level of proteins bound to the peptide probes. The prediction module is used to compare the sum of the acquired detection results of the combination of peptide probes with a set threshold to obtain a prediction result; the set threshold is obtained by the following method: acquiring the sum of detection results of the combination of peptide probes in a training sample and the corresponding labeling results, wherein the labeling results are labels indicating the risk or prognosis of early-stage lung adenocarcinoma in the sample; using the Youden's J statistic method, obtaining a classification threshold for the risk or prognosis of early-stage lung adenocarcinoma in the training sample, and taking the median value of ≥100 results as the final classification threshold.

[0014] Fifthly, embodiments of the present invention provide an electronic device, the electronic device including a processor and a memory; the memory is used to store a program, which, when executed by the processor, causes the processor to implement the training method or the prediction method for early-stage lung adenocarcinoma described in the foregoing embodiments. The prediction method includes: acquiring the sum of detection results of a combination of peptide probes in a test sample; the combination of peptide probes is the combination of peptides described in the foregoing embodiments, and the detection results include the detection signal value of the peptide probes or the level of proteins bound to the peptide probes; comparing the sum of the acquired detection results with a set threshold to obtain a prediction result; wherein the set threshold is obtained by the following method: acquiring the sum of detection results of a combination of peptide probes in a training sample and the corresponding labeling results, the labeling results being a label of the risk of early-stage lung adenocarcinoma in the sample; using the Youden's J statistic method to obtain a classification threshold for the risk of early-stage lung adenocarcinoma in the training sample, and taking the median value of ≥100 results as the final classification threshold.

[0015] In a sixth aspect, embodiments of the present invention provide a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the method for predicting early lung adenocarcinoma as described in the foregoing embodiments.

[0016] The present invention has the following beneficial effects:

[0017] This invention, through the study and analysis of plasma samples from early-stage lung adenocarcinoma and healthy individuals, identified 143 polypeptide probes with significant signal differences between the two groups. Based on this, 14 polypeptide probes that produced more stable signal differences between the two groups were further screened. Various machine learning methods were explored to establish a classification model. It was found that using the sum of the signal values ​​of the 14 polypeptide probes as the indicator for model construction could establish a good classification method, providing a way to diagnose the risk of early-stage lung adenocarcinoma. Attached Figure Description

[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a comparison of Early-LUAD and NHC peptide probe combinations. The horizontal axis represents the group label, and the vertical axis represents the signal value of the peptide probe. Each point represents a sample. The contour curve describes the kernel density distribution of the peptide probe signal values ​​of all samples in the corresponding group. From the vertical perspective, the denser the sample area, the higher the peak of the contour curve, which means there are more samples with that signal value. Comparing the peaks of the distribution curves of the Early-LUAD group and the NHC group, the peaks of the Early-LUAD group are all higher. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Where specific conditions are not specified in the embodiments, conventional conditions or conditions recommended by the manufacturer shall apply. Reagents or instruments whose manufacturers are not specified are all conventional products that can be purchased commercially.

[0021] This invention focuses more on the very early stages of lung adenocarcinoma (LUAD) (TNM stages 0 and I). From the perspective of autoantibodies that bind to various peptide-mimicking antigenic epitopes, it studies the autoantibody profile of early-stage lung adenocarcinoma, which is of great significance for identifying these early-stage patients with the best long-term survival as early as possible.

[0022] This application utilizes a peptide chip containing 130,000 peptide probes to study and analyze plasma samples from an early-stage lung adenocarcinoma group and a healthy population. It identified 143 peptide probes with significant signal differences between the two groups. Further screening revealed 14 peptide probes that produced more stable signal differences between the two groups. Various machine learning methods were explored to establish a classification model. The study found that using the sum of the signal values ​​of the 14 peptide probes as the model's metric, and employing Youden's J statistic method to obtain a classification threshold for the risk of early-stage lung adenocarcinoma in the training samples, established a good classification method, providing a pathway for the effective diagnosis of early-stage lung adenocarcinoma.

[0023] Specifically, the present invention provides a polypeptide combination comprising: polypeptides with sequences as shown in SEQ ID No. 1 to 14.

[0024] Table 1. Peptide sequences

[0025] SEQ ID No. sequence 1 RYLWANKRS 2 WRKESDNVEG 3 LYKSGNKYNVD 4 VKFDKYNLE 5 WLPLKRPAKG 6 RSGGFSWLG 7 RFQVSAWAG 8 PFNSQLGG 9 YRDKLRAKFS 10 RFNGDGWG 11 QHSAWNRYSYNKR 12 ADLSKFAQSD 13 KARSFFSG 14 VHKDYWQHG

[0026] The main inventive point of this application lies in identifying peptide probes that can specifically distinguish between early-stage lung adenocarcinoma and healthy individuals. Detection methods based on these peptide probes can be implemented using existing methods for peptide probe detection, and will not be elaborated further. Compared to existing technologies or other detection proteins / peptides, given a fixed detection method, using the peptide combination proposed in this application as detection probes can effectively capture autoantibodies that specifically distinguish between early-stage lung adenocarcinoma and healthy individuals. The detection signal values ​​from these probes can effectively and cost-effectively predict the risk of developing early-stage lung adenocarcinoma.

[0027] In some embodiments, the polypeptide is modified with a detectable marker.

[0028] Optionally, the marker includes at least one of fluorescent dyes and nanoparticle markers.

[0029] Optionally, the fluorescent dye is selected from at least one of fluorescein dyes, rhodamine dyes, Cy series dyes, Alexa series dyes, and protein dyes.

[0030] Optionally, the nanoparticle-based markers include either nanoparticles or colloids.

[0031] Optionally, the nanoparticles include at least one of organic nanoparticles, magnetic nanoparticles, quantum dot nanoparticles, and rare earth complex nanoparticles.

[0032] Optionally, the colloid is selected from at least one of latex, colloidal selenium, colloidal metal, disperse dye, and dye-labeled microspheres.

[0033] Optionally, the detectable marker is labeled at the N-terminus and / or C-terminus of the polypeptide sequence, preferably at the N-terminus.

[0034] On the other hand, embodiments of the present invention provide the application of polypeptide combinations in the preparation of products for the diagnosis or auxiliary diagnosis of early lung adenocarcinoma, wherein the polypeptide combination is the polypeptide combination described in any of the foregoing embodiments.

[0035] On the other hand, embodiments of the present invention provide a product comprising the polypeptide combination described in any of the foregoing embodiments.

[0036] In some embodiments, the product includes at least one of reagents, kits, and chips, and is used for the diagnosis or auxiliary diagnosis of early-stage lung adenocarcinoma. The polypeptide combination described in any of the foregoing embodiments, as a detection probe, can specifically bind to autoantibodies in a sample. By obtaining the signal value of the marker on the detection probe, the content / level of the binding protein in the sample can be obtained.

[0037] In some embodiments, the chip can be a peptide microarray chip.

[0038] In some embodiments, the polypeptide assembly is coupled to a solid support. The solid support includes at least one of magnetic beads, a chip, glass, a nylon membrane, nitrocellulose, and a PVDF membrane.

[0039] On the other hand, embodiments of the present invention provide a device for predicting early-stage lung adenocarcinoma, comprising:

[0040] The acquisition module is used to acquire the cumulative sum of the detection results of the peptide probe combination in the sample to be tested; the peptide probe combination is the peptide combination described in any of the foregoing embodiments, and the detection result includes the detection signal value of the peptide probe or the level of the protein bound to the peptide probe;

[0041] The prediction module compares the sum of the acquired detection results with a set threshold to obtain a prediction result. The set threshold is obtained by the following method: acquiring the sum of the detection results of the peptide probe combination in the training sample and the corresponding annotation results, wherein the annotation results are labels indicating the risk or prognostic risk of early-stage lung adenocarcinoma in the sample; using the Youden's J statistic method, obtaining the classification threshold of the risk or prognostic risk of early-stage lung adenocarcinoma in the training sample, and taking the median value of ≥100 results as the final classification threshold.

[0042] Optionally, the above modules can be stored in memory or embedded in the operating system (OS) of the electronic device provided in this application, and can be executed by the processor in the electronic device. Meanwhile, the data and program code required to execute the above modules can be stored in memory.

[0043] In some embodiments, the test sample or training sample includes a plasma sample, a serum sample, or a whole blood sample. Optionally, the test sample or training sample also includes an environmental sample containing at least one of the plasma sample, serum sample, or whole blood sample.

[0044] In some embodiments, the sample size of the training samples can be ≥10, specifically any one or any two of the following: 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180 and 200.

[0045] On the other hand, embodiments of the present invention provide an electronic device, the electronic device including a processor and a memory; the memory is used to store a program, which, when executed by the processor, causes the processor to implement the training method or the early lung adenocarcinoma prediction method described in any of the foregoing embodiments;

[0046] The methods for predicting early-stage lung adenocarcinoma include:

[0047] The cumulative sum of detection results of the peptide probe combination in the sample to be tested is obtained; the peptide probe combination is the peptide combination described in any of the foregoing embodiments, and the detection result includes the detection signal value of the peptide probe or the level of the protein bound to the peptide probe;

[0048] The sum of the acquired detection results is compared with a set threshold to obtain the prediction result; wherein, the set threshold is obtained by the following method: the sum of the detection results of the peptide probe combination in the training sample and the corresponding annotation result are obtained, wherein the annotation result is a label of the risk or prognostic risk of early lung adenocarcinoma in the sample; the Youden's Jstatistic method is used to obtain the classification threshold of the risk of early lung adenocarcinoma in the training sample, and the median value of ≥100 results is taken as the final classification threshold.

[0049] The memory can be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), etc.

[0050] A processor can be an integrated circuit chip with signal processing capabilities. This processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0051] In practical applications, the electronic device can be a server, cloud platform, mobile phone, tablet computer, laptop computer, ultra-mobile personal computer (UMPC), handheld computer, netbook, personal digital assistant (PDA), wearable electronic device, virtual reality device, etc. Therefore, the embodiments of this application do not limit the types of electronic devices.

[0052] On the other hand, embodiments of the present invention provide a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the method for predicting early lung adenocarcinoma as described in any of the foregoing embodiments.

[0053] Computer-readable media include various media that can store program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.

[0054] The features and performance of the present invention will be further described in detail below with reference to embodiments.

[0055] Example 1

[0056] 1. Queue Sample

[0057] The inventors of this application continuously collected plasma samples from 269 patients with surgically confirmed lung adenocarcinoma and healthy controls undergoing cancer screening from December 2019 to March 2021. These included 147 cases of early-stage lung adenocarcinoma (Early-LUAD) and 122 normal healthy controls (NHC).

[0058] The case group in this application consists of patients with early-stage lung adenocarcinoma whose surgically removed tissue was confirmed by pathological biopsy. The definition of early stage is based on stage 0, IA, and IB in the 8th edition of the TNM staging system of the American Joint Committee on Cancer (AJCC) and the Union for International Cancer Control (UICC).

[0059] The healthy controls were individuals with no history of malignant tumors, whose annual routine cancer screenings included a series of imaging and laboratory tests that showed no significant malignant abnormalities.

[0060] Plasma samples from patients with early-stage lung adenocarcinoma were isolated from blood specimens collected before surgical treatment. Blood samples from healthy controls were collected during cancer screening. All whole blood specimens were initially separated by centrifugation at 3000 rpm for 10 minutes on the day of acquisition, and the supernatant plasma sample was immediately frozen and stored at -80°C for later use.

[0061] Table 2. Sample Information

[0062]

[0063]

[0064] 2. Peptide chip detection

[0065] This study utilizes peptide chip detection technology developed by Shenzhen Carbon Cloud Intelligent Technology Co., Ltd. The detection chip combines advanced semiconductor manufacturing processes and amino acid solid-phase synthesis technology, synthesizing a high-density array of 130,000 peptides designed to maximize amino acid combination diversity. This simulates the binding of antigen epitopes to antibodies in the sample. The antibody spectrum is unbiasedly reflected by detecting fluorescence signals in a fluorescence microscopy imaging system, making it a high-throughput immunoaffinity assay. This technology can analyze differentially expressed antibody profiles between individuals, enabling the establishment of disease classification models and the prediction of disease-related autoantigens.

[0066] 2.1 Experimental Design

[0067] Each chip represents one detection unit. Before the experiment begins, a well-designed experimental plan should be developed, ensuring that the proportion of samples from each disease group is roughly the same on each chip, with samples from each group randomly arranged within the chip, and fixed calibrators used. The number of chips required should be calculated, and the chip numbering and sample arrangement determined.

[0068] 2.2 Experimental Procedure

[0069] (1) Sample preparation: Plasma samples were stored at -80°C. When preparing for testing, the aliquoted samples were thawed at 4°C, and then each sample was diluted 500 times with 1% mannitol phosphate buffer and added to a 96-well plate in the order of detection for analysis.

[0070] (2) Chip hydration and assembly: Before use, the peptide chip is immersed in ultrapure water for 20 minutes for hydration. After removal, the surface of the peptide chip is briefly sprayed with 90% isopropanol 3 times to ensure complete coverage. Then, it is centrifuged to remove excess liquid and placed into the peptide chip holder.

[0071] (3) Chip sealing: Add 0.1% casein blocking solution to each sample well and incubate at 37°C for 1 hour.

[0072] (4) Primary antibody incubation and elution: Add the above diluted sample and incubate at 37°C for 1 hour in a constant temperature shaker. After incubation, place the sample in a plate washer and wash it 3 times with phosphate buffer to remove unbound molecules and reagents.

[0073] (5) Secondary antibody incubation and elution: Prepare a 2 nM fluorescent secondary antibody solution using 0.75% casein solution, add 40 μL / well to the assay cassette, and incubate at 37°C for 1 hour using a constant temperature shaker. After incubation, place the cassette in a plate washer and wash three times with phosphate buffer to thoroughly remove unbound molecules and reagents.

[0074] (6) Imaging: Remove the peptide chip from the cartridge, clean it with 90% isopropanol, and then centrifuge it to dry. Scan the peptide chip in the ImageXpress scanning system to obtain one TIFF image file for each sample, which is the raw data.

[0075] 3. Data Preprocessing

[0076] Perform logarithmic transformation and standardization on the sample values ​​to output a standardized matrix, specifically including:

[0077] (1) Extract the fluorescence intensity values ​​of the features and output one GPR5 data file and one corner images file. The GPR5 file contains all the information of a sample and the fluorescence intensity information of all features.

[0078] (2) Extract the characteristic fluorescence intensity information from the GPR5 data files of all samples to generate the original fluorescence intensity (FG, foreground) data matrix. Then, perform a logarithmic transformation on the data of each sample to obtain the LFG (log-transferred foreground) data matrix, retaining only the library probes, which is the original probe signal data. Subtract the median value of all probe signals in each sample from the probe signal of that sample, which is used as the data for subsequent calculation and analysis. This value is used for differential analysis and modeling (which can be defined as the peptide probe signal value).

[0079] The HealthTell V13 chip used in this application contains over 130,000 probes. Of these, 125,509 are library probes, ranging in length from 5 to 13 amino acids. These library probes cover 99.9% of all possible 4-mer combinations and 48.3% of all possible 5-mer combinations of the 16 amino acids. The remaining probes are used for data quality control, etc.

[0080] 4. Screening Method

[0081] (1) Differential Analysis: Student's t-Test was used to perform differential analysis on the discovery set samples to identify peptide probes that could distinguish between the Early-LUAD and NHC groups. Peptide probes that simultaneously met the criteria of P-value < 0.05, fold change (FC) > 0.05, and positive rate > 0.15 were selected. The positive rate was calculated based on the threshold of mean + 2 × sd for each probe in the control group (NHC group), representing the proportion of samples exceeding this threshold in the Early-LUAD group.

[0082] (2) Differential stability analysis: For the above differentially expressed peptide probes, 20% of the discovery set samples were randomly discarded. In the remaining 80% of the discovery set samples, the P-value of each peptide probe was calculated using Student's t-Test. This process was repeated 100 times, and the number of times each peptide probe met the P-value < 0.05 was recorded, which is the differential stability of the peptide probe.

[0083] Finally, the intersection of peptide probes with P-value < 0.01 in (1) differential analysis and peptide probes with a number of occurrences of 100 in (2) differential stability analysis was selected as an important biomarker for distinguishing Early-LUAD and NHC groups.

[0084] 5. Screening Results

[0085] Comparing the Early-LUAD and NHC groups, 143 peptide probes were identified in the differential analysis. After filtering for differential stability, 14 peptide probes were finally selected as peptide probe combinations, such as... Figure 1 As shown, the following are listed: RYLWANKRS, WRKESDNVEG, LYKSGNKYNVD, VKFDKYNLE, WLPLKRPAKG, RSGGFSWLG, RFQVSAWAG, PFNSQLGG, YRDKLRAKFS, RFNGDGWG, QHSAWNRYSYNKR, ADLSKFAQSD, KARSFFSG, and VHKDYWQHG.

[0086] 6. Predictive Model

[0087] For the differentially stable peptide probes (14 peptide probes obtained by comparing Early-LUAD and NHC groupings), the validation set samples were randomly split, with 70% of the samples used as the training set and the remaining 30% used as the test set. This random splitting was performed 100 times. A classification model was built on the training set, and the model performance was evaluated on the test set using AUC, sensitivity, and specificity metrics.

[0088] In the above process: (1) Based on the signal values ​​of the above 14 peptide probes in the training set samples, various classification models were tried for machine learning training, such as logistic regression, support vector machine and random forest; it was found that none of them achieved good classification results; (2) The feature summation was tried as the distinguishing index. The summation of the values ​​of all peptide probe signals was used as the feature. Youden's J statistic method was used to obtain the best classification threshold on the training set. The median value of 100 results was taken as the final classification threshold. The classification sensitivity of this method was 67%, the specificity was 84%, and the AUC reached 0.79. The sensitivity and AUC were better than the results of Reference 1 (sensitivity 36%~39%, specificity about 90%, AUC: 0.63~0.71). The effects of different classification models are shown in Table 3.

[0089] Table 3. Performance of different classification models in Early-LUAD vs NHC

[0090] Classification model AUC Sensitivity Specificity Logistic Regression 0.40±0.10 0.75±0.12 0.15±0.12 Support Vector Machine 0.51±0.13 0.99±0.02 0.01±0.02 Random Forest 0.53±0.12 0.93±0.06 0.06±0.09 Based on feature accumulation and differentiation 0.79±0.03 0.67±0.09 0.84±0.11

[0091] Based on the average value of the evaluation indicators on the test set, the method of feature accumulation and differentiation was finally selected to predict and diagnose early-stage lung adenocarcinoma, as shown in Table 4.

[0092] Table 4. Final model performance distinguishing Early-LUAD vs NHC

[0093]

[0094] Applications of disease diagnosis:

[0095] Serum samples from the subjects were collected as test samples and detected using a polypeptide chip containing these polypeptide probes (Table 1). The sum of the signal values ​​of these polypeptide probes was compared with a set classification threshold. If the sum of the signal values ​​of the test sample was greater than or equal to the classification threshold, it was classified as a disease group. If the sum of the signal values ​​of the test sample was less than the classification threshold, it was classified as a healthy group.

[0096] In this embodiment, the classification threshold for Early-LUAD vs NHC is set to 1.924.

[0097] However, if the sample to be tested is significantly different from the samples in this batch, such as having other diseases, the classification threshold may need to be redefined. The method for determining the threshold can be as described above, using Youden's Jstatistic method to obtain the best classification threshold on the training set, and taking the median value of 100 (minimum 100) results as the final classification threshold.

[0098] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A polypeptide combination, characterized in that, It includes: The amino acid sequence of the polypeptide is shown in SEQ ID No. 1 to 14.

2. The polypeptide combination according to claim 1, characterized in that, The polypeptide is modified with a detectable marker.

3. The polypeptide combination according to claim 2, characterized in that, The markers include at least one of fluorescent dyes and nanoparticle markers.

4. The application of polypeptide combinations in the preparation of products for the diagnosis or auxiliary diagnosis of early-stage lung adenocarcinoma, characterized in that, The polypeptide combination is any one of the polypeptide combinations according to claims 1 to 3.

5. The application according to claim 4, characterized in that, The products include reagents or kits.

6. A product characterized in that, It includes the polypeptide combination of any one of claims 1 to 3, and the product is a reagent or kit for diagnosing or assisting in the diagnosis of early lung adenocarcinoma.

7. The product according to claim 6, characterized in that, The polypeptide combination is coupled to a solid support.

8. The product according to claim 7, characterized in that, The solid support includes at least one of magnetic beads, chips, glass, nylon membrane, nitrocellulose and PVDF membrane.

9. A device for predicting early-stage lung adenocarcinoma, characterized in that, It includes: An acquisition module is used to acquire the cumulative sum of detection results of peptide probe combinations in the sample to be tested; the peptide probe combination is any one of claims 1 to 3, and the detection result includes the detection signal value of the peptide probe or the level of the protein bound to the peptide probe; The prediction module compares the sum of the acquired detection results with a set threshold to obtain a prediction result. The set threshold is obtained by the following method: acquiring the sum of the detection results of the peptide probe combination in the training sample and the corresponding annotation results, wherein the annotation results are labels indicating the risk or prognostic risk of early-stage lung adenocarcinoma in the sample; using the Youden's Jstatistic method, obtaining the classification threshold of the risk or prognostic risk of early-stage lung adenocarcinoma in the training sample, and taking the median value of ≥100 results as the final classification threshold.

10. An electronic device, characterized in that, The electronic device includes a processor and a memory; the memory stores a program that, when executed by the processor, enables the processor to implement a method for predicting early-stage lung adenocarcinoma; the prediction method includes: The cumulative sum of detection results of peptide probe combinations in the sample to be tested is obtained; the peptide probe combination is any one of claims 1 to 3, and the detection result includes the detection signal value of the peptide probe or the level of the protein bound to the peptide probe; the cumulative sum of the obtained detection results is compared with a set threshold to obtain a prediction result; wherein, the set threshold is obtained by the following method: obtaining the cumulative sum of detection results of peptide probe combinations in the training sample and the corresponding labeling result, wherein the labeling result is a label of the risk or prognostic risk of early lung adenocarcinoma in the sample; using Youden's Jstatistic method, the classification threshold of the risk or prognostic risk of early lung adenocarcinoma in the training sample is obtained, and the median value of ≥100 results is taken as the final classification threshold.

11. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method for predicting early-stage lung adenocarcinoma as described in claim 10.

Citation Information

Patent Citations

  • Amino acid sequence for detecting lung cancer marker MYC epitope and application

    CN105037534A

  • Joint detection serum marker for early diagnosis of lung adenocarcinoma and application of joint detection serum marker

    CN113687076A