Peptide combination for distinguishing early lung adenocarcinoma from benign lung diseases and its application

By screening out 27 peptide probes and using their signal values to accumulate and establish a classification model, the problem of insufficient sensitivity and specificity of early lung adenocarcinoma and benign lung diseases in the prior art was solved, and a more efficient identification effect was achieved.

CN116087517BActive Publication Date: 2025-08-15PEKING UNION MEDICAL COLLEGE HOSPITAL +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211203468.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-29
Publication Date
2025-08-15
Estimated Expiration
2042-09-29

AI Technical Summary

Technical Problem

The prior art distinguishes early lung adenocarcinoma from benign lung diseases, the detection sensitivity and specificity are insufficient, and the existing detection methods are complicated to operate, making it difficult to effectively identify the properties of lung nodules.

Method used

By studying serum samples from the early lung adenocarcinoma and benign lung disease groups, 27 peptide probes were screened out, and the signal values of these peptide probes were accumulated as indicators, combined with machine learning methods, a classification model was established, and the classification threshold was determined by Youden’s J statistic method to achieve effective identification of early lung adenocarcinoma and benign lung diseases.

Benefits of technology

It improves the classification sensitivity and specificity of early stage lung adenocarcinoma and benign lung diseases, provides more accurate identification methods, and reduces detection costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116087517B_ABST
    Figure CN116087517B_ABST
Patent Text Reader

Abstract

The present invention discloses a polypeptide combination and its application for distinguishing early-stage lung adenocarcinoma from benign lung diseases, relating to the field of biological detection technology. By studying and analyzing plasma samples from an early-stage lung adenocarcinoma group and a benign lung disease group, the present invention discovered 133 polypeptide probes with significant signal differences between the two groups. Based on this, 27 polypeptide probes were further screened to produce more stable signal differences between the two groups. By trying multiple machine learning methods to establish a classification model, it was found that using the cumulative sum of the signal values of the 27 polypeptide probes as an indicator for model construction could establish a good classification method, providing a way to effectively identify or distinguish early-stage lung adenocarcinoma from benign lung diseases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biological detection technology, and in particular to a polypeptide combination for distinguishing early lung adenocarcinoma from benign lung diseases and applications thereof. Background Art

[0002] Lung cancer ranks second in global incidence among malignant tumors, and its mortality rate remains the leading cause of cancer death. Lung adenocarcinoma is the most common pathological type, accounting for approximately 40% of all pathological classifications. Early detection can reduce lung cancer-related mortality and improve prognosis. LDCT, currently the primary detection method, has high sensitivity, but its high false-positive rate leads to the detection of a large number of uncertain lung nodules, requiring long-term follow-up. Therefore, the development of new, easily detectable blood biomarkers is needed to better differentiate between benign and malignant lung nodules. Due to gene mutations, abnormal protein structure, modification, or expression, tumor cells may release abnormal proteins, including neoantigens, which stimulate the body's immune response and produce autoantibodies through clonal proliferation of immune B cells. Compared to autoantigens, autoantibodies are more stable and, after being amplified by the immune response, are easier to detect.

[0003] Numerous studies have reported that autoantibodies can be detected in the early stages of various cancers, often appearing before clinical signs or imaging findings. Therefore, tumor-associated autoantibodies are promising tumor biomarkers. Because blood samples are readily available and autoantibody signals are richer and less variable than tumor proteins, these readily detectable markers are attractive for early diagnosis.

[0004] Based on the one-to-one recognition model of antibodies against antigens, ELISA methods have been used both domestically and internationally to detect autoantibody combinations for lung cancer diagnosis targeting 6 or 7 protein antigens. The earliest study in 2010 reported that the ELISA method was used to detect autoantibodies targeting 6 tumor-associated antigens including p53, CAGE, NY-ESO-1, Annexin 1, GBU4-5 and SOX2. Its sensitivity in lung cancer diagnosis was 36%-39%, specificity was about 90%, and AUC value was 0.63-0.71 (Reference 1: Boyle, P, et al. (2011). Clinical validation of an autoantibody test for lung cancer. Ann Oncol. 22 (2): 383-389.). The team subsequently adjusted the autoantibody combination and finally formed a commercial ELISA detection kit for 7 autoantibody combinations (p53, NY-ESO-1, CAGE, GBU4-5, SOX2, HuD and MAGE A4). -Lung. Evidence from multiple studies shows that -Lung can assist CT to improve the risk assessment of pulmonary nodules.

[0005] Similarly, the composition of the commercial ELISA kits for seven autoantibodies of lung cancer in China (p53, CAGE, GBU4-5, GAGE 7, SOX2, PGP9.5 and MAGE A1) is similar to that of Similar to the LDCT-Lung test, its performance in the Chinese population is 61% sensitivity and 90% specificity (as indicated on the Keboro autoantibody kit). Many other autoantibody studies have explored the performance of these two categories of seven autoantibodies in different populations based on ELISA. Identifying new, easily detectable autoantibodies to further screen lung nodules detected by LDCT may avoid overdiagnosis and treatment during long-term LDCT follow-up.

[0006] All of the aforementioned autoantibody studies have used ELISA to detect autoantibodies targeting known common cancer-associated proteins or embryonic testis antigens. However, there are relatively few studies exploring unknown autoantibodies with diagnostic efficacy superior to those reported. This is due to the tedious and complex experimental procedures required to discover new autoantibody detection technologies, such as serum proteome analysis (SERPA), serological analysis of recombinantly expressed cDNA clones (SEREX), and phage display technology. In light of this, the present invention was proposed. Summary of the Invention

[0007] The purpose of the present invention is to provide a polypeptide combination and its application for distinguishing early lung adenocarcinoma from benign lung diseases.

[0008] The present invention is achieved in that:

[0009] In a first aspect, an embodiment of the present invention provides a polypeptide combination, comprising: polypeptides having sequences as shown in SEQ ID No. 1 to 27.

[0010] In a second aspect, embodiments of the present invention provide use of a polypeptide combination in preparing a product for distinguishing early-stage lung adenocarcinoma from benign lung diseases, wherein the polypeptide combination is the polypeptide combination described in the aforementioned embodiments.

[0011] In a third aspect, an embodiment of the present invention provides a product comprising the polypeptide combination described in the preceding embodiments.

[0012] In a fourth aspect, an embodiment of the present invention provides a prediction device for early lung adenocarcinoma and benign lung disease, which includes an acquisition module and a prediction module. The acquisition module is used to obtain the cumulative sum of the detection results of the polypeptide probe combination in the test sample; wherein the polypeptide probe combination is the polypeptide combination described in the above embodiment, and the detection result includes the detection signal value of the polypeptide probe or the level of the protein bound to the polypeptide probe; the prediction module compares the cumulative sum of the obtained detection results with a set threshold to obtain a prediction result; the set threshold is obtained by the following method: obtaining the cumulative sum of the detection results of the polypeptide probe combination in the training sample and the corresponding annotation result, the annotation result is a label for the sample as early lung adenocarcinoma or benign lung disease; using Youden's J statistic method to obtain the classification threshold of early lung adenocarcinoma and benign lung disease in the training sample, and taking the median value of ≥100 results as the final classification threshold.

[0013] In a fifth aspect, an embodiment of the present invention provides an electronic device comprising a processor and a memory; the memory is used to store a program, and when the program is executed by the processor, the processor implements a method for distinguishing between early lung adenocarcinoma and benign lung disease. The distinguishing method comprises: obtaining the cumulative sum of the detection results of the polypeptide probe combination in the test sample; wherein the polypeptide probe combination is the polypeptide combination described in the above embodiment, and the detection result includes the detection signal value of the polypeptide probe or the level of the protein bound to the polypeptide probe; comparing the cumulative sum of the obtained detection results with a set threshold to obtain a prediction result; the set threshold is obtained by the following method: obtaining the cumulative sum of the detection results of the polypeptide probe combination in the training sample and the corresponding annotation result, the annotation result is a label for the sample as early lung adenocarcinoma or benign lung disease; using Youden's J statistic method to obtain the classification threshold of early lung adenocarcinoma and benign lung disease in the training sample, and taking the median value of ≥100 results as the final classification threshold.

[0014] In a sixth aspect, an embodiment of the present invention provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the method for distinguishing between early lung adenocarcinoma and benign lung diseases described in the aforementioned embodiment.

[0015] The present invention has the following beneficial effects:

[0016] By studying and analyzing serum samples from an early-stage lung adenocarcinoma group and a benign lung disease group, the present invention discovered 133 polypeptide probes with significant signal differences between the two groups. Based on this, 27 polypeptide probes that can produce more stable signal differences between the two groups were further screened. Multiple machine learning methods were used to establish a classification model. Ultimately, it was found that using the cumulative sum of the signal values of the 27 polypeptide probes as an indicator and the Youden's J statistic method to obtain the classification threshold for early-stage lung adenocarcinoma and benign lung disease in training samples, a good classification method could be established, providing a way to effectively identify or differentiate early-stage lung adenocarcinoma and benign lung disease. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 The following are peptide probe combinations for Early-LUAD vs. BLD. The horizontal axis represents the group label, and the vertical axis represents the peptide probe signal value. Each point represents a sample. The contour curve depicts the kernel density distribution of the peptide probe signal values for all samples in the corresponding group. Vertically, the denser the sample density, the higher the peak of the contour curve, indicating that more samples have the same signal value. Comparing the peak of the Early-LUAD group's distribution curve with the peak of the BLD distribution curve, the Early-LUAD group's peak is higher. DETAILED DESCRIPTION

[0019] To make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are described clearly and completely below. Where specific conditions are not specified in the embodiments, conventional conditions or conditions recommended by the manufacturer are used. Where the manufacturer of the reagents or instruments is not specified, all are conventional products that can be purchased commercially.

[0020] The present invention focuses on the very early stages of lung adenocarcinoma (LUAD) (TNM stage 0 and I). The study examines the autoantibody spectrum of early-stage lung adenocarcinoma from the perspective of autoantibodies that bind to antigenic epitopes mimicked by multiple peptides. This is of great significance for early identification of these patients with the best long-term survival.

[0021] This application uses a peptide chip containing 130,000 peptide probes to study and analyze serum samples from an early lung adenocarcinoma group and a benign lung disease group, and found 133 peptide probes with significant signal differences between the two groups. On this basis, 27 peptide probes that can produce more stable signal differences between the two groups were further screened. Multiple machine learning methods were tried to establish a classification model. It was found that based on the signal values of the 27 peptide probes, the Youden's J statistic method was used to obtain the classification thresholds for early lung adenocarcinoma and benign lung disease in the training samples. This can establish a good classification method and indicator, providing a way to effectively identify or distinguish early lung adenocarcinoma and benign lung disease.

[0022] Specifically, the present invention provides a polypeptide combination, which includes: polypeptides with sequences as shown in SEQ ID No. 1 to 27.

[0023] Table 1. Peptide sequences

[0024]

[0025]

[0026] The invention of this application is primarily to identify polypeptide probes that can specifically distinguish early-stage lung adenocarcinoma from benign lung diseases. Detection methods based on these polypeptide probes can be implemented based on existing detection methods for polypeptide probes, which will not be described in detail. Compared to existing technologies or other detection proteins / peptides, under certain detection methods, the use of the polypeptide combination proposed in this application as a detection probe can effectively capture autoantibodies that specifically distinguish early-stage lung adenocarcinoma from benign lung diseases. The detection signal value of the detection probe can effectively and cost-effectively distinguish early-stage lung adenocarcinoma from benign lung diseases.

[0027] In some embodiments, the polypeptide is modified with a detectable label.

[0028] Optionally, the marker includes at least one of a fluorescent dye and a nanoparticle marker.

[0029] Optionally, the fluorescent dye is selected from at least one of fluorescein dyes, rhodamine dyes, Cy series dyes, Alexa series dyes and protein dyes.

[0030] Optionally, the nanoparticle marker includes any one of nanoparticles and colloids.

[0031] Optionally, the nanoparticles include at least one of organic nanoparticles, magnetic nanoparticles, quantum dot nanoparticles and rare earth complex nanoparticles.

[0032] Optionally, the colloid is selected from at least one of latex, colloidal selenium, colloidal metal, disperse dyes and dye-labeled microspheres.

[0033] Optionally, the detectable marker is labeled at the N-terminus and / or C-terminus of the polypeptide sequence, preferably at the N-terminus.

[0034] On the other hand, an embodiment of the present invention provides the use of a polypeptide combination in preparing a product for distinguishing early lung adenocarcinoma from benign lung diseases, wherein the polypeptide combination is the polypeptide combination described in any of the aforementioned embodiments.

[0035] On the other hand, an embodiment of the present invention provides a product comprising the polypeptide combination described in any of the aforementioned embodiments.

[0036] In some embodiments, the product comprises at least one of a reagent, a kit, and a chip. The polypeptide combination described in any of the preceding embodiments can be used as a detection probe to specifically bind to autoantibodies in a sample, and the content / level of the binding protein in the sample can be obtained by obtaining the signal value of the marker on the detection probe.

[0037] In some embodiments, the chip can be a polypeptide microarray chip.

[0038] In some embodiments, the polypeptide combination is coupled to a solid support, wherein the solid support comprises at least one of magnetic beads, chips, glass, nylon membrane, nitrocellulose, and PVDF membrane.

[0039] In another aspect, an embodiment of the present invention provides a device for predicting early-stage lung adenocarcinoma and benign lung diseases, comprising:

[0040] an acquisition module, configured to obtain the cumulative sum of the detection results of the polypeptide probe combination in the sample to be tested; the polypeptide probe combination is the polypeptide combination described in any of the aforementioned embodiments, and the detection result includes the detection signal value of the polypeptide probe or the level of the protein bound to the polypeptide probe;

[0041] The prediction module compares the cumulative sum of the obtained test results with a set threshold to obtain a prediction result; the set threshold is obtained by the following method: obtaining the cumulative sum of the test results of the polypeptide probe combination in the training sample and the corresponding annotation results, wherein the annotation results are labels indicating that the sample is early lung adenocarcinoma or benign lung disease; using Youden's Jstatistic method to obtain the classification threshold of early lung adenocarcinoma and benign lung disease in the training sample, and taking the median value of ≥100 results as the final classification threshold.

[0042] Optionally, the above modules can be stored in a memory in the form of software or firmware or fixed in the operating system (OS) of the electronic device provided in this application, and can be executed by a processor in the electronic device. At the same time, the data, program code, etc. required to execute the above modules can be stored in the memory.

[0043] In some embodiments, the test sample or training sample includes a plasma sample, a serum sample, or a whole blood sample. Optionally, the test sample or training sample also includes an environmental sample containing at least one of the plasma sample, the serum sample, or the whole blood sample.

[0044] In some embodiments, the sample size of the training samples may be ≥10, specifically any one of 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180 and 200, or a range between any two of them.

[0045] In another aspect, an embodiment of the present invention provides an electronic device comprising a processor and a memory; the memory is used to store a program, and when the program is executed by the processor, the processor performs a method for distinguishing early-stage lung adenocarcinoma from benign lung diseases;

[0046] The distinguishing method includes:

[0047] Obtaining a detection result of the polypeptide probe combination in the test sample; the polypeptide probe combination is the polypeptide combination described in any of the preceding embodiments, and the detection result includes a detection signal value of the polypeptide probe or a level of protein bound to the polypeptide probe;

[0048] The cumulative sum of the obtained test results is compared with a set threshold to obtain a prediction result; wherein the set threshold is obtained by the following method: the cumulative sum of the test results of the polypeptide probe combination in the training sample and the corresponding annotation results are obtained, and the annotation results are labels indicating that the sample is early lung adenocarcinoma or benign lung disease; the Youden's J statistic method is used to obtain the classification threshold of early lung adenocarcinoma and benign lung disease in the training sample, and the median value of ≥100 results is taken as the final classification threshold.

[0049] The memory may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.

[0050] A processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, including a central processing unit (CPU) or a network processor (NP). It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0051] In actual applications, the electronic device can be a server, a cloud platform, a mobile phone, a tablet computer, a laptop computer, an ultra-mobile personal computer (UMPC), a handheld computer, a netbook, a personal digital assistant (PDA), a wearable electronic device, a virtual reality device, etc. Therefore, the embodiments of the present application do not limit the type of electronic device.

[0052] On the other hand, an embodiment of the present invention provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the method for distinguishing between early lung adenocarcinoma and benign lung diseases described in any of the aforementioned embodiments.

[0053] Computer-readable media include: USB flash drives, mobile hard drives, read-only memories, random access memories, magnetic disks, or optical disks, among other media that can store program codes.

[0054] The features and performance of the present invention are further described in detail below with reference to the embodiments.

[0055] Example 1

[0056] 1. Cohort Sample

[0057] The inventors of this application continuously collected a total of 255 plasma samples from patients with suspected lung cancer (including lung adenocarcinoma and benign lung nodules confirmed by surgical pathology) from December 2019 to March 2021, including 147 cases of early lung adenocarcinoma (Early-LUAD) and 108 cases of benign lung disease (Benign lung disease, BLD).

[0058] The case group of this application is all patients with early-stage lung adenocarcinoma diagnosed by pathological biopsy of surgically resected tissue. The definition of early stage is based on the 8th edition TNM staging system of the American Joint Committee on Cancer (AJCC) and the Union for International Cancer Control (UICC) 0, IA, and IB. Benign lung disease refers to patients suspected of lung cancer by imaging examination, but with benign lung nodules confirmed by surgical resection tissue pathology. The types of benign nodules mainly include pulmonary tuberculosis and suspicious tuberculoid nodules, organizing pneumonia, hamartoma, sclerosing alveolar cell tumor of the lung, and inflammatory pseudotumor of the lung.

[0059] Plasma samples from patients with early-stage lung adenocarcinoma and benign lung diseases were isolated from blood specimens collected before surgical treatment. All whole blood specimens were initially separated by centrifugation at 3000 rpm for 10 minutes on the day of collection. The supernatant plasma sample was immediately frozen at -80°C for storage.

[0060] Table 2. Sample information

[0061]

[0062] 2. Peptide chip detection

[0063] This assay utilizes peptide chip detection technology developed by Shenzhen Carbon Cloud Intelligent Technology Co., Ltd. The chip, manufactured using advanced semiconductor manufacturing processes and amino acid solid-phase synthesis technology, features a high-density array of 130,000 peptides designed to maximize amino acid combination diversity. These peptides mimic antigen epitopes binding to antibodies in the sample, and the fluorescence signal detected by a fluorescence microscopy system provides an unbiased reflection of the antibody spectrum in the sample. This high-throughput immunoaffinity assay can analyze differentially expressed antibody profiles between individuals, leveraging these differentially expressed antibody profiles to establish disease classification models and predict disease-associated autoantigens.

[0064] 2.1 Experimental design

[0065] One chip is a single test unit. Before the experiment begins, carefully design the experiment, ensuring that the proportion of samples from each disease group on each chip is roughly consistent, that the samples from each group are randomly arranged within the chip, and that a fixed calibrator is used. Calculate the number of chips required and determine the chip numbering and sample layout.

[0066] 2.2 Experimental Procedure

[0067] (1) Sample preparation: Plasma samples were stored at −80°C. When preparing for testing, the aliquoted samples were thawed at 4°C. Each sample was then diluted 500-fold with 1% mannitol in phosphate buffer and loaded into a 96-well plate in the order of testing for analysis.

[0068] (2) Chip hydration and assembly: Before use, the peptide chip was immersed in ultrapure water for 20 minutes for hydration. After removal, the peptide chip surface was briefly sprayed three times with 90% isopropyl alcohol to ensure complete coverage. The chip was then centrifuged to remove excess liquid and loaded into the peptide chip holder.

[0069] (3) Chip blocking: Add 0.1% casein blocking solution to each sample well and incubate at 37°C for 1 hour.

[0070] (4) Primary antibody incubation and elution: Add the diluted sample and incubate in a constant temperature oscillator at 37°C for 1 hour. After incubation, place the plate in a plate washer and wash three times with phosphate buffer to fully remove unbound molecules and reagents.

[0071] (5) Secondary antibody incubation and elution: Prepare a 2 nM fluorescent secondary antibody solution with 0.75% casein solution, add 40 μL / well to the assay cassette, and incubate in a thermostat at 37°C for 1 hour. After incubation, place the plate in a plate washer and wash three times with phosphate buffer to fully remove unbound molecules and reagents.

[0072] (6) Imaging: Remove the peptide chip from the cartridge, spray clean with 90% isopropyl alcohol, and spin dry in a centrifuge. Scan the peptide chip using the ImageXpress scanning system. A TIFF image file is generated for each sample tested, which is the raw data.

[0073] 3. Data Preprocessing

[0074] Perform logarithmic transformation and standardization on the sample values and output a standardized matrix, including:

[0075] (1) Extract the fluorescence intensity values of the features and output a GPR5 data file and a corner images file. The GPR5 file contains all the information of a sample and the fluorescence intensity information of all features.

[0076] (2) Extract the characteristic fluorescence intensity information from the GPR5 data files of all samples to generate the raw fluorescence intensity (FG, foreground) data matrix. Then, perform logarithmic transformation on the data of each sample to obtain the LFG (log-transferred foreground) data matrix. Only the library probe is retained, which is the raw probe signal data. The probe signal of each sample is subtracted from the median value of all probe signals of that sample, which is used as the data for subsequent calculation and analysis. This value is used for differential analysis and modeling (which can be defined as the peptide probe signal value).

[0077] The HealthTell V13 chip used in this application contains over 130,000 probes. Of these, 125,509 are library probes, ranging in length from 5 to 13 amino acids. These library probes encompass 99.9% of all possible 4-mer combinations and 48.3% of all possible 5-mer combinations of the 16 amino acids. The remaining probes are used for data quality control and other purposes.

[0078] 4. Screening Methods

[0079] (1) Difference analysis: The Student's t-Test method was used to perform a difference analysis on the discovery set samples to identify peptide probes that can distinguish the Early-LUAD and BLD groups. Peptide probes that meet the requirements of P-value < 0.05, fold change (FC) > 0.05, and positive rate > 0.15 were screened. The positive rate was calculated based on the mean + 2 × sd of each probe in the control group (BLD group) as the threshold, and the proportion of samples exceeding this threshold in the Early-LUAD group to the total number of samples in that group.

[0080] (2) Differential stability analysis: For the differentially expressed peptide probes, 20% of the discovery set samples were randomly discarded. In the remaining 80% of the discovery set samples, the Student's t-test method was used to calculate the P-value of each peptide probe. This process was repeated 100 times, and the number of times each peptide probe met the P-value < 0.05 was recorded, which was the differential stability of the peptide probe.

[0081] Finally, the intersection of (1) peptide probes with a P-value < 0.001 in the differential analysis and (2) peptide probes with a number of times reaching 100 in the differential stability analysis was selected as the important biomarkers for distinguishing Early-LUAD and BLD groups.

[0082] 5. Filter results

[0083] Comparing the Early-LUAD and BLD groups, 133 peptide probes were found in the differential analysis. After filtering by differential stability, 27 peptide probes were finally screened, which were peptide probe combinations, such as Figure 1 shown, including PNAQWAHDG, HPVEWYVD, DDWQEDYKSD, SGDHLDQKFD, PDENDYKDG, PSVGDSG, PNHAADRSD, WNWSDPAAKVS, PNQVWDPPEG, SNPFADG, DWGNEAQHD, PSVWEPARSD, PDDNYHVDG, PQPYHR AFED, PQPEFELEG, PFNNVESE, WNDSEVPKFD, PFREREPDG, PNAEWVHVD, NRHSKRDNLNRG, NAVERHVWD, PEAKAARLG, PQLDQLENNRFS, PNAQPQFPRE, PFNKDPNYQED, PSVQWQEG and PSWEEKLD.

[0084] 6. Prediction method or differentiation method

[0085] For the peptide probes filtered for differential stability (27 peptide probes obtained by comparing the Early-LUAD and BLD groups), the validation set samples were randomly split, with 70% of the samples used as the training set and the remaining 30% as the test set. This random split was repeated 100 times, and a classification model was constructed on the training set. The model performance was evaluated on the test set using AUC, sensitivity, and specificity.

[0086] In the above process: (1) Based on the signal values of the above 27 peptide probes in the training set samples, various classification models were tried for machine learning training, such as logistic regression, support vector machine, and random forest. It was found that none of them achieved good classification results. (2) The feature accumulation sum was tried as the discrimination index. The numerical accumulation sum of all peptide probe signals was used as the feature. The Youden's J statistic method was used to obtain the optimal classification threshold on the training set. The median value of 100 results was taken as the final classification threshold. The classification sensitivity of this method was 45%, the specificity was 96%, and the AUC reached 0.77, which were all better than the results of reference 1 (sensitivity 36% to 39%, specificity about 90%, AUC: 0.63 to 0.71). The effects of different classification models are shown in Table 3.

[0087] Table 3. Effects of different classification models for Early-LUAD vs BLD

[0088] Classification Model AUC Sensitivity (%) Specificity (%) Logistic regression 0.50±0.11 0.64±0.13 0.34±0.16 Support Vector Machine 0.49±0.11 0.98±0.04 0.01±0.04 Random Forest 0.58±0.09 0.83±0.10 0.16±0.13 Accumulate and distinguish features 0.77±0.02 0.45±0.09 0.96±0.05

[0089] Based on the average values of the evaluation indicators on the test set, the feature accumulation and differentiation method was finally used to distinguish early lung adenocarcinoma from benign lung diseases, as shown in Table 4.

[0090] Table 4. Final model performance in distinguishing Early-LUAD from BLD

[0091]

[0092]

[0093] Application of disease diagnosis:

[0094] The serum of the subject is obtained as the test sample, and chip detection is performed using a polypeptide chip containing these polypeptide probes (Table 1). The cumulative sum of the signal values of these polypeptide probes is compared with the set classification threshold. If the cumulative sum of the signal values of the test sample is ≥ the classification threshold, the test sample is judged to be in the disease group; if the cumulative sum of the signal values of the test sample is < the classification threshold, the test sample is judged to be in the benign disease group.

[0095] For the model trained in this embodiment, the threshold for the Early-LUAD vs BLD grouping is 4.228.

[0096] However, if the sample to be tested is very different from the samples in this batch, such as if it is accompanied by some other diseases, this threshold may need to be re-determined. The method for determining the threshold can be as described above, using Youden's Jstatistic method to obtain the optimal classification threshold on the training set, and taking the median value of 100 (at least 100) results as the final classification threshold.

[0097] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A polypeptide combination for distinguishing early lung adenocarcinoma from benign lung diseases, characterized in that: It includes: The amino acid sequences are polypeptides shown in SEQ ID No. 1 to 27.

2. The polypeptide combination according to claim 1, characterized in that The polypeptide is modified with a detectable marker; the marker includes at least one of a fluorescent dye and a nanoparticle marker.

3. Use of a polypeptide combination in the preparation of a product for distinguishing early lung adenocarcinoma from benign lung diseases, characterized in that: The polypeptide combination is the polypeptide combination according to claim 1 or 2.

4. The use according to claim 3, characterized in that The product includes a reagent or a kit.

5. A product, characterized in that It comprises the polypeptide combination according to claim 1 or 2.

6. The product according to claim 5, characterized in that The polypeptide combination is coupled to a solid phase carrier; the solid phase carrier comprises at least one of magnetic beads, chips, glass, nylon membrane, nitrocellulose and PVDF membrane.

7. The product according to claim 5 or 6, characterized in that The products include reagents or kits for distinguishing early-stage lung adenocarcinoma from benign lung diseases.

8. A prediction device for early lung adenocarcinoma and benign lung diseases, characterized in that: It includes: an acquisition module, configured to obtain the cumulative sum of the detection results of the polypeptide probe combination in the sample to be tested; wherein the polypeptide probe combination is the polypeptide combination according to claim 1 or 2, and the detection result includes the detection signal value of the polypeptide probe or the level of the protein bound to the polypeptide probe; The prediction module compares the cumulative sum of the obtained test results with a set threshold to obtain a prediction result; the set threshold is obtained by the following method: obtaining the cumulative sum of the test results of the polypeptide probe combination in the training sample and the corresponding annotation results, wherein the annotation results are labels indicating that the sample is early lung adenocarcinoma or benign lung disease; using the Youden's J statistic method to obtain the classification threshold for early lung adenocarcinoma and benign lung disease in the training sample, and taking the median value of ≥100 results as the final classification threshold.

9. An electronic device, characterized in that: The electronic device includes a processor and a memory; the memory is used to store a program, and when the program is executed by the processor, the processor implements a method for distinguishing early lung adenocarcinoma from benign lung diseases; The differentiation method comprises: obtaining a cumulative sum of the detection results of the polypeptide probe combination in the sample to be tested; wherein the polypeptide probe combination is the polypeptide combination according to claim 1 or 2, and the detection result comprises the detection signal value of the polypeptide probe or the level of the protein bound to the polypeptide probe; The cumulative sum of the obtained test results is compared with a set threshold to obtain a prediction result; the set threshold is obtained by the following method: the cumulative sum of the test results of the polypeptide probe combination in the training sample and the corresponding annotation results are obtained, and the annotation results are labels indicating that the sample is early lung adenocarcinoma or benign lung disease; the Youden's J statistic method is used to obtain the classification threshold for early lung adenocarcinoma and benign lung disease in the training sample, and the median value of ≥100 results is taken as the final classification threshold.

10. A computer-readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for distinguishing early lung adenocarcinoma from benign lung diseases according to claim 9 is implemented.

Citation Information

Patent Citations

  • Lung cancer early diagnostic marker based on metabonomics and artificial intelligence technology and application thereof

    CN109884302A

  • Joint detection serum marker for early diagnosis of lung adenocarcinoma and application of joint detection serum marker

    CN113687076A