Biomarkers for early diagnosis of lung adenocarcinoma and / or classification of indeterminate pulmonary nodules and uses thereof

By using piR-hsa-8429916 and piR-hsa-8393202 piRNA biomarkers and a random forest classification model in lung adenocarcinoma, the accuracy problem of lung nodule classification in existing technologies has been solved, achieving efficient and low-cost detection for early diagnosis and nodule classification of lung adenocarcinoma.

CN120945053BActive Publication Date: 2026-03-17CANCER INST & HOSPITAL CHINESE ACADEMY OF MEDICAL SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Current technologies lack reliable non-invasive biomarkers to assist in the accurate classification of pulmonary nodules in low-dose computed tomography screening, resulting in high false positive rates and psychological burden. Furthermore, existing circulating tumor DNA testing has low sensitivity in early-stage lung adenocarcinoma and cannot effectively distinguish between malignant nodules and inflammatory lesions.

Method used

Two piRNAs, piR-hsa-8429916 and piR-hsa-8393202, were developed as biomarkers. Combined with a random forest classification model, the expression levels of these piRNAs in serum or lung tissue were detected for the early diagnosis of lung adenocarcinoma and the classification of unexplained lung nodules.

Benefits of technology

It significantly improves the sensitivity and specificity of early diagnosis of lung adenocarcinoma, reduces the false positive rate of low-dose computed tomography, and provides a low-cost, highly standardized detection method suitable for large-scale population screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120945053B_ABST
    Figure CN120945053B_ABST
Patent Text Reader

Abstract

The application discloses biomarkers for early diagnosis of lung adenocarcinoma and / or classification of unknown lung nodules and application thereof, and relates to the technical field of biological medicine.The biomarkers are piR-hsa-8429916 or piR-hsa-8393202.Two core biomarkers, piR-hsa-8393202 and piR-hsa-8429916, are successfully identified by screening and verifying piRNAs differentially expressed in lung adenocarcinoma tissues and serum, and the expression levels of the biomarkers are significantly positively correlated with tumor load.The diagnostic value of piR-hsa-8393202 and piR-hsa-8429916 in early diagnosis of lung adenocarcinoma and classification of lung nodules is significantly better than that of a traditional CEA marker.The application provides an innovative solution for early diagnosis of lung adenocarcinoma and risk stratification of lung nodules.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biomedical technology, and in particular to biomarkers for the early diagnosis of lung adenocarcinoma and / or the classification of unidentified pulmonary nodules and their applications. Background Technology

[0002] Lung adenocarcinoma (LUAD) is the most common histological subtype of non-small cell lung cancer (NSCLC), accounting for more than 40% of all lung cancer cases. It is estimated that approximately 2.2 million new cases and 1.8 million deaths occur annually. Survival rates for LUAD patients vary significantly depending on the clinical stage. The 10-year survival rate for early-stage (stage IA) patients is 92%, while the 5-year survival rate for advanced-stage (stage IV) patients is less than 10%. This significant disparity underscores the urgent need for early detection strategies during the pre-metastatic window to halt disease progression.

[0003] Low-dose computed tomography (LDCT) screening has revolutionized the detection of LUAD by identifying lung nodules in high-risk populations. However, the clinical application of LDCT is limited by a 96% false positive rate, leading to unnecessary invasive examinations of benign nodules and imposing a significant psychological burden on patients. Current guidelines lack reliable non-invasive biomarkers to assist LDCT in accurately classifying nodules, highlighting an urgent unmet need in the field of precision oncology.

[0004] Biomarkers based on liquid biopsy, particularly circulating tumor DNA (ctDNA) and cell-free DNA (cfDNA), have become promising tools for early cancer detection. In non-small cell lung cancer (NSCLC), ctDNA detection targeting driver mutations (such as EGFR L858R) has a sensitivity of 70%–85% in advanced patients, but in stage IA lung adenocarcinoma (LUAD), its sensitivity drops significantly to 30%–50% due to low tumor DNA shedding. This limitation is even more pronounced in subcentimeter nodules (<8 mm), where ctDNA detection rates are below 10%. Furthermore, cfDNA-based methylation detection plates have shown improved sensitivity (63%) for early LUAD, but their specificity (82%) remains insufficient to differentiate malignant nodules from inflammatory lesions. Recently, multi-omics approaches combining exosomal miRNAs and protein biomarkers have shown synergistic diagnostic potential, but their clinical application is hampered by pre-analysis variability in extracellular vesicle isolation protocols. While machine learning models combining radiomics features and serum biomarkers have improved diagnostic accuracy, their generalization ability is limited by cohort-specific bias and a lack of standardized preprocessing procedures. These shared limitations underscore the need to discover novel biomarkers with greater organ specificity and analytical robustness.

[0005] Emerging evidence suggests that PIWI-interacting RNAs (piRNAs), a class of small non-coding RNAs, are key regulators of oncogenic pathways. Unlike microRNAs, piRNAs exhibit tissue-specific expression patterns and are stable and detectable in biological fluids, making them ideal candidates for non-invasive diagnosis. In colorectal and breast cancer, dysregulation of piRNAs is associated with tumor proliferation and metastasis by epigenetically silencing tumor suppressor factors. In lung adenocarcinoma (LUAD), preliminary studies have shown that tumor-derived piRNAs are selectively packaged into exosomes and secreted into the circulatory system, potentially serving as molecular fingerprints of early malignant tumors. Despite these insights, a systematic investigation of piRNA expression profiles in lung adenocarcinoma tissues and their paired blood samples remains insufficient. Therefore, this invention aims to develop novel biomarkers for the early diagnosis of lung adenocarcinoma and / or the classification of unexplained pulmonary nodules, thereby providing technical support for the early diagnosis of lung adenocarcinoma and the accurate classification of nodules. Summary of the Invention

[0006] The purpose of this invention is to provide biomarkers for the early diagnosis of lung adenocarcinoma and / or the classification of unexplained pulmonary nodules, and their applications, to address the problems existing in the prior art. The expression level of this biomarker is significantly positively correlated with the tumor burden of lung adenocarcinoma, and can be applied to the early diagnosis of lung adenocarcinoma and the classification of unexplained pulmonary nodules, thus providing an innovative solution for the early diagnosis of lung adenocarcinoma and risk stratification of pulmonary nodules.

[0007] To achieve the above objectives, the present invention provides the following solution:

[0008] This invention provides a biomarker for the early diagnosis of lung adenocarcinoma and / or the classification of unidentified pulmonary nodules, wherein the biomarker is piR-hsa-8429916 or piR-hsa-8393202.

[0009] The present invention also provides a combination of biomarkers for the early diagnosis of lung adenocarcinoma and / or the classification of unidentified pulmonary nodules, the combination of biomarkers including piR-hsa-8429916 and piR-hsa-8393202.

[0010] The present invention also provides the use of reagents for detecting the expression levels of the above-mentioned biomarkers or combinations of biomarkers in serum or lung tissue in the preparation of early diagnostic products for lung adenocarcinoma.

[0011] Furthermore, the early diagnostic product for lung adenocarcinoma is a reagent kit.

[0012] The present invention also provides the use of reagents for detecting the expression levels of the above-mentioned biomarkers or combinations of biomarkers in serum or lung nodule tissue in the preparation of detection products for distinguishing between benign and malignant lung nodules.

[0013] Furthermore, the testing product is a reagent kit.

[0014] The present invention also provides an early diagnostic product for lung adenocarcinoma, comprising reagents for detecting the expression levels of the above-mentioned biomarkers or combinations of biomarkers in serum or lung tissue.

[0015] Furthermore, the early diagnostic product for lung adenocarcinoma is a reagent kit.

[0016] The present invention also provides a detection product for distinguishing between benign and malignant pulmonary nodules, comprising a reagent for detecting the expression levels of the above-mentioned biomarkers or combinations of biomarkers in serum or pulmonary nodule tissue.

[0017] Furthermore, the testing product is a reagent kit.

[0018] The present invention discloses the following technical effects:

[0019] This invention successfully identified two core biomarkers, piR-hsa-8393202 and piR-hsa-8429916, by systematically screening and verifying differentially expressed piRNAs in lung adenocarcinoma tissues and serum. Their expression levels were significantly positively correlated with tumor burden, and they exhibited excellent stability in serum.

[0020] Based on this, the present invention constructs a random forest classification model that integrates piRNA expression, age, sex, and CEA level. The data balance is optimized by the SMOTE algorithm and validated by a multi-center cohort. It shows excellent performance in the early diagnosis of lung adenocarcinoma (stage IA-II): the AUC value of the training set reaches 0.918 and the AUC value of the validation set reaches 0.863, which is significantly better than traditional CEA markers. In particular, the AUC is improved to 0.930 in the EGFR mutation subgroup, providing a new direction for molecular subtyping-assisted diagnosis.

[0021] In the classification of pulmonary nodules, the model is particularly effective in identifying sub-centimeter nodules (AUC=0.897), and its classification accuracy for some solid nodules (AUC=0.906) and patients with abnormal CEA is further improved to 0.966. It can significantly reduce the false positive rate of LDCT screening by 11.3%, providing a key basis for clinical decision-making for unexplained pulmonary nodules.

[0022] Furthermore, the detection methods based on the biomarkers of this invention, such as RT-qPCR, have the advantages of low cost and high standardization, are easy to promote to large-scale population screening, and have the reliability and universality of clinical application, providing an innovative solution for early diagnosis of lung adenocarcinoma and risk stratification of lung nodules. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 A schematic diagram illustrating the workflow for acquiring, profiling, and differentially analyzing piRNAs using publicly available datasets;

[0025] Figure 2 Differentially expressed piRNAs between lung adenocarcinoma tumor tissues (n=48) and paired adjacent non-cancerous tissues (n=48);

[0026] Figure 3 Differentially expressed piRNAs between serum samples from lung adenocarcinoma patients (n=4) and healthy donors (n=5);

[0027] Figure 4 Venn diagram of piRNAs that are simultaneously upregulated in serum samples and tumor tissues of patients with lung adenocarcinoma;

[0028] Figure 5 A graph showing the top 10 candidate piRNAs and their area under the curve (AUC) values ​​used to distinguish lung adenocarcinoma tumor tissue from paired adjacent normal tissue;

[0029] Figure 6 Heatmap of expression of the top 10 candidate piRNAs in a lung adenocarcinoma tissue cohort;

[0030] Figure 7 Heatmap of expression of the top 10 candidate piRNAs in a serum cohort of lung adenocarcinoma;

[0031] Figure 8 The graph shows the expression level detection results of 10 candidate piRNAs in the tissue cohort (n=48);

[0032] Figure 9 Subject operating characteristic curves for candidate piRNAs in the organization cohort;

[0033] Figure 10 The expression levels of 10 candidate piRNAs in the serum cohort (n=48) are shown in the figure.

[0034] Figure 11 Receiver operating characteristic curves for candidate piRNAs in the serum cohort;

[0035] Figure 12Statistical graphs showing the expression levels of piR-hsa-8429916(A) and piR-hsa-8393202(B) in the healthy, benign, and LUAD subgroups;

[0036] Figure 13 Workflow diagram for diagnostic model development;

[0037] Figure 14 A comparative evaluation chart of diagnostic performance metrics for nine different machine learning paradigms;

[0038] Figure 15 The receiver operating characteristic curve for each fold in the 10-fold cross-validation of a random forest;

[0039] Figure 16 A bee colony diagram showing the impact of each feature on the model results using Shapely values;

[0040] Figure 17 Receiver operating characteristic curves for the diagnosis of LUAD patients in training set (A), test set 1 (B), and test set 2 (C);

[0041] Figure 18 Confusion graph of the diagnostic model binary results in training set (A), test set 1 (B), and test set 2 (C);

[0042] Figure 19 Principal component analysis plots showing the predictive performance of the diagnostic model in training set (A), test set 1 (B), and test set 2 (C) in distinguishing between LUAD (red) and non-LUAD (green);

[0043] Figure 20 The following are performance analysis graphs for the lung nodule classification model: A is a schematic diagram of the development of the lung nodule classifier based on 2-piRNAs; B is the receiver operating characteristic (ROC) curve for each fold of 10-fold cross-validation using random forest; C is the result graph of relative variable importance assessed by mean reduced Gini index; D and E are box plots of the scores of the 2-piRNAs-based classifier for patients with benign and malignant nodules in the training set (D) and test set (E), respectively; F and G are the ROC curves of the 2-piRNAs-based classifier and carcinoembryonic antigen (CEA) for lung nodule classification in the training set (F) and test set (G); H and I are the confusion graphs of binary results of the lung nodule classification model in the training set (H) and test set (I); J and K are the principal component analysis graphs of the predictive performance of the classification model in distinguishing between malignant nodules (red) and benign nodules (green) in the training set (J) and test set (K).

[0044] Figure 21Figure 1 shows the subgroup performance analysis results of the 2-piRNAs-based lung nodule classifier. Specifically, a) represents the receiver operating characteristic (ROC) curves of the 2-piRNAs-based classifier under different carcinoembryonic antigen (CEA) states; b) represents the ROC curves of the 2-piRNAs-based classifier under different nodule sizes; c) represents the ROC curves of the 2-piRNAs-based classifier under different nodule radiographic subtypes; and d) is a schematic diagram illustrating the risk stratification process of lung nodules based on the 2-piRNAs-based classifier.

[0045] Figure 22 Venn diagram of upregulated piRNAs in tumor tissues and serum cohorts of lung adenocarcinoma;

[0046] Figure 23 The image shows the results of detecting the expression levels of piR-hsa-8393202(A) and piR-hsa-8429916(B) in serum samples that have undergone repeated freeze-thaw cycles.

[0047] Figure 24 A graph showing the comparison results of serum piR-hsa-8393202 (A) and piR-hsa-8429916 (B) in paired samples before and after surgery (n=48);

[0048] Figure 25 Correlation analysis plot of relative expression of piR-hsa-8393202(A) and piR-hsa-8429916(B) in paired serum and tissue samples (n=45);

[0049] Figure 26 Correlation analysis plot of decreased serum piR-hsa-8393202(A) and piR-hsa-8429916(B) expression with maximum tumor size (n=53);

[0050] Figure 27 Statistical graphs of piR-hsa-8393202 levels in cells (A) and culture medium (B) in normal lung epithelial cell line BEAS-2B and some lung adenocarcinoma cell lines;

[0051] Figure 28 Statistical graphs showing the levels of piR-hsa-8429916 in cells (A) and culture medium (B) of normal lung epithelial cell line BEAS-2B and some lung adenocarcinoma cell lines. Detailed Implementation

[0052] Various exemplary embodiments of the present invention will now be described in detail. This detailed description should not be considered as a limitation of the present invention, but rather as a more detailed description of certain aspects, features, and embodiments of the present invention.

[0053] It should be understood that the terminology used in this invention is merely for describing particular embodiments and is not intended to limit the invention. Furthermore, with respect to numerical ranges in this invention, it should be understood that each intermediate value between the upper and lower limits of the range is also specifically disclosed. Any stated value or intermediate value within a stated range, as well as each smaller range between any other stated value or intermediate value within said range, is also included in this invention. The upper and lower limits of these smaller ranges may be independently included or excluded from the range.

[0054] Unless otherwise stated, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. While only preferred methods and materials have been described herein, any methods and materials similar or equivalent to those described herein may be used in the implementation or testing of this invention. All references to this specification are incorporated by way of citation to disclose and describe methods and / or materials associated with those references. In the event of any conflict with any incorporated reference, the content of this specification shall prevail.

[0055] Various modifications and variations can be made to the specific embodiments described in this specification without departing from the scope or spirit of the invention, as will be apparent to those skilled in the art. Other embodiments derived from this specification will also be apparent to those skilled in the art. This specification and embodiments are merely exemplary.

[0056] The terms “include,” “including,” “have,” “contain,” etc., used in this article are all open-ended terms, meaning that they include but are not limited to.

[0057] Example 1

[0058] I. Experimental Methods

[0059] 1. Research Design

[0060] A prospective, retrospective, blinded evaluation (PRoBE) cohort study was conducted from 2021 to 2024. The study recruited 1653 participants from three medical centers in three separate provinces and municipalities in China. A total of 1556 blood samples and 53 pairs of lung adenocarcinoma (LUAD) tissue and matched adjacent normal tissue samples were collected. The study protocol was approved by the Institutional Review Boards of the National Cancer Center of China (NCC2021C-527), Henan Cancer Hospital, and Shanxi Cancer Hospital, and complies with the Declaration of Helsinki (2013 revised version). All participants signed written informed consent forms before enrollment.

[0061] 2. Participant Recruitment

[0062] All diagnoses of lung adenocarcinoma (LUAD) followed the World Health Organization (WHO) Classification of Thoracic Tumors (Fifth Edition) and were staged using the International Association for the Study of Lung Cancer (IASLC) Eighth Edition TNM staging system. All lung adenocarcinoma patients were newly diagnosed, and blood samples were collected prior to biopsy, surgery, or systemic treatment.

[0063] Exclusion criteria included: (1) concurrent malignancy; (2) prior chemotherapy / immunotherapy; (3) severe liver and kidney dysfunction (AST / ALT > 3 times the upper limit of normal, eGFR < 30 ml / min / 1.73 m²); (4) active infection (HIV, HBV) or autoimmune disease; and (5) age under 18 years. The control group was matched to the case group in age, sex, and smoking status, and had normal inflammatory markers (CRP < 5 mg / L, white blood cell count 4-10 × 10⁻⁶). 9 / Lift).

[0064] Benign nodule cases were confirmed by histopathology, excluding individuals with a family history of cancer or occupational exposure to carcinogens. Pathological types of benign pulmonary nodules included bronchial adenomas, chronic granulomatous lesions, atypical adenomatous hyperplasia, and chronic obstructive pulmonary disease. Subjects underwent at least three consecutive CT scans. Healthy controls underwent comprehensive screening, including tumor marker testing (carcinoembryonic antigen, alpha-fetoprotein), chest X-ray, and abdominal / pelvic ultrasound to rule out occult malignancies or inflammatory diseases.

[0065] 3. Public data mining and candidate piRNA screening

[0066] Small RNA sequencing data were obtained from lung adenocarcinoma (LUAD) tumor tissue and serum samples, and were acquired from the Gene Expression Comprehensive Database (GEO), access numbers GSE110907 and GSE151963, respectively. Raw sequencing reads were converted to FASTQ format using the SRA toolkit (v3.0.0) (https: / / www.ncbi.nlm.nih.gov / sra / docs / toolkitsoft / ) and quality was assessed using FastQC. Reads meeting the following criteria were retained: (1) length distribution consistent with piRNA biological characteristics (26–32 nucleotides); (2) base call quality ≥Q30 in over 80% of cycles; and (3) no adapter contamination.

[0067] Adapter trimming and quality filtering were performed using Cutadapt (v3.4) with the following parameters: -q 20 --minimum-length 18 --maximum-length 35 --trim-n. The piRNA annotation file (piRBase v2.0) was aligned with the human reference genome (GRCh38) using BWA (v0.7.17) with strict alignment parameters: -n 1 -k 11 -l 11 (allowing ≤1 mismatch). The aligned reads were processed using SAMtools (v1.12) to generate sorted BAM files.

[0068] Read counts were calculated using HTseq-count (version 0.9.1) in joint counting mode. Piwi-interacting RNAs (piRNAs) with a count per million (CPM) ≥1 in at least 20% of tumor samples were retained for subsequent analysis. Data were normalized using the truncated mean (TMM) method based on M-values. Differentially expressed piRNAs were identified using DESeq2 (v1.34.0) with a threshold of |log2 (fold change)| ≥1 and a false discovery rate (FDR) of less than 0.05 as determined by Benjamini-Hochberg correction.

[0069] 4. Collection of tissue samples

[0070] Between September 2021 and March 2022, samples of 53 pairs of lung adenocarcinoma (LUAD) and their paired adjacent normal tissues were collected from patients who underwent surgical resection at the Cancer Hospital of the Chinese Academy of Medical Sciences (Beijing).

[0071] 5. Serum Collection and Processing

[0072] After a 12-hour fast, peripheral venous blood (3-4 ml) was drawn using a serum vacuum tube. The sample was centrifuged at 1500×g for 10 minutes within 30 minutes of collection to minimize hemolysis. Serum samples were aliquoted and stored at -80°C until analysis. Strict quality control was maintained: samples showing hemolysis (absorbance at 414 nm > 0.3) or improper storage were excluded.

[0073] 6. Total RNA extraction and RT-qPCR

[0074] Total RNA from clinical tissue samples was extracted using TRIzol reagent (Invitrogen, catalog number 15596018), while serum-derived RNA was extracted using TRIzol LS (Invitrogen, catalog number 10296010) to optimize the recovery of low-abundance RNA. Reverse transcription was performed using PrimeScript based on poly(A) tails. TMPerform RT Master Mix (Sangon Biotech, catalog number B532451) and operate according to the manufacturer's instructions.

[0075] Quantitative PCR analysis was performed using Roche. The 480Ⅱ system was used, employing SYBR Green assay reagents (Roche, catalog number 04887352001). Tissue RNA quantification used U6 small nuclear RNA as an internal control, while serum samples, in addition to endogenous miR-16-5p, also included exogenous cel-miR-39 as a control (Thermo Fisher Scientific, catalog number 4395408) for data standardization. Each sample underwent three independent biological replicate analyses. Relative expression levels were calculated using the ΔΔCt method, and melting curve analysis confirmed reaction specificity.

[0076] 7. Cell lines and culture

[0077] Human normal alveolar epithelial cell line BEAS-2B and lung adenocarcinoma cell lines (A549, NCI-H1650, HCC827, and NCI-H358) were purchased from the Cell Bank of Shanghai Institutes for Biological Sciences, Chinese Academy of Sciences (Shanghai, China). BEAS-2B cells were purchased from the National Cell Bank of China. All cell lines were passaged for no more than 6 generations after thawing and identified by short tandem repeat (STR) analysis. Mycoplasma contamination was excluded using routine PCR-based detection methods. BEAS-2B cells were cultured in Durbeco Modified Eagle Medium (DMEM; Gibco, Grand Island, NY, USA) containing 10% fetal bovine serum (FBS; Gibco). A549 cells were cultured in DMEM containing 10% FBS and 100 U / mL penicillin / streptomycin (Gibco). NCI-H1650, HCC827, and NCI-H358 cells were cultured in RPMI-1640 medium (Gibco) containing 10% FBS and antibiotics. All cells were cultured under humid conditions at 37°C and 5% CO2. For piRNA analysis, 1×10⁻⁶ cells were cultured... 6 Cells were seeded in 10 cm culture dishes and allowed to adhere overnight. The cells were then starved in FBS-free medium for 4 hours. Cell pellet and corresponding cell-free culture supernatant were collected for subsequent quantitative analysis of piR-hsa-8429916 and piR-hsa-8393202.

[0078] 8. Construction of Diagnostic Models

[0079] The diagnostic prediction model was developed using the Tidymodels package (version 0.1.4) in R (version 4.3.0). Participants in Cohort 1 (n = 1028) were stratified into a training set (n = 718) and an internal test set (n = 310) through random sampling, while an external validation cohort (n = 349) from independent centers was included to evaluate its generalizability. A modular workflow was implemented for nine machine learning algorithms (logistic regression, support vector machine, random forest, gradient boosting, k-nearest neighbors, Naive Bayes, neural networks, elastic networks, and decision trees). All models were trained using four predictor variables: expression levels of two piRNAs, age, and sex. Data preprocessing included standardization (z-scores for continuous variables) and class imbalance correction via SMOTE. For the optimized random forest model, 1000 decision trees were generated through double randomization: bootstrap resampling (80% sample replacement) and feature subspace selection. Hyperparameter tuning employed a Latin hypercube grid search, evaluating 20 combinations of mtry(2-4) and min-n(5-20) using stratified 10-fold cross-validation. Model performance was assessed using the area under the receiver operating characteristic (AUC) (primary endpoint), sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV). Key predictors were identified through permutation importance analysis, and computational reproducibility was ensured through fixed random seeds and parallel processing (8-core CPU). Model interpretability was assessed using SHAP (Shapley Additive exPlanations) values, calculated using the Monte Carlo approximation method, iterated 50 times, and biased. Analysis was performed using the "fastshap" R package, ensuring additive consistency and clinical interpretability. Receiver operating characteristic curves were generated using the multiple receiver operating characteristic package, and the DeLong test was used to compare AUC differences between models.

[0080] 9. Construction of Classifier Model

[0081] A study was conducted on 350 participants with IPN to develop a binary classification model using five predictor variables. The random forest algorithm was implemented using the tidymodels framework (version 0.1.4) in R (version 4.3.0), following the Transparent Reporting (TRIPOD) guidelines for multivariate predictive models of individual prognosis or diagnosis. Stratified random sampling was used, allocating 75% (n = 262) to the training set and 25% (n = 88) to the test set. Missing values ​​(<2% of data) were imputed using median replacement, followed by z-score standardization. Hyperparameters were optimized using grid search in combinations of mtry(2–4) and min–n(5–20) and evaluated using 10-fold stratified cross-validation. The final random forest ensemble contained 1000 decision trees, utilizing bootstrap aggregation and feature subspace randomization to enhance generalization. Model selection prioritized receiver operating characteristic (AUC) while balancing sensitivity and specificity. Diagnostic thresholds were determined by maximizing the Youden index, and confidence intervals were calculated using a bootstrapping method (1000 iterations) to quantify the stability of the indicator. Performance was evaluated on the test set. Comprehensive receiver operating characteristic (ROC) curves were plotted using the "multi receiver operating characteristic" package. Feature importance was quantified by averaging the reduced Gini coefficient based on Gini impurity and visualized using a gradient color plot from the "vip" package. A multi-level random seed strategy ensured the reproducibility of data partitioning, model training, and validation. Principal component analysis (PCA) was performed on z-score-standardized features, and singular value decomposition was used to retain components with a cumulative variance explanation greater than 95%.

[0082] 10. Statistical Analysis

[0083] When comparing two unpaired samples, a two-sided Wilcoxon rank-sum test was used. For nonparametric analysis of paired samples, the Wilcoxon signed-rank test was used. When comparing three or more groups, a two-sided Kruskal-Wallis test was used. Spearman rank correlation (rho) analysis was applied to determine the correlation between the two variables, and the p-value was calculated using a two-sided test. Receiver operating characteristic (ROC) curves, precision and recall (PR) curves, sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), and overall accuracy were used to assess diagnostic performance. The DeLong test was used to assess the significance of differences in AUC. Data with outliers were analyzed using nonparametric tests. All statistical analyses were performed using SPSS 20.0 (IBM), GraphPad Prism (v.9.0), and R (v.3.6.0) software (https: / / www.r-project.org / ). P < 0.05 was considered statistically significant.

[0084] II. Experimental Results

[0085] 1. Participants and their characteristics

[0086] A total of 1,653 participants were included, including 1,050 patients with pathologically confirmed lung adenocarcinoma (LUAD), 109 cases of benign pulmonary nodules, and 494 healthy controls. The primary cohort consisted of participants retrospectively enrolled at the National Cancer Center (NCC) in Beijing between 2021 and 2024, comprising 711 LUAD patients, 72 patients with benign pulmonary nodules, and 317 healthy controls. The external cohort was established through a multicenter collaboration between 2023 and 2024, recruiting 168 LUAD patients, 17 benign cases, and 52 healthy controls from Henan Cancer Hospital, and 34 LUAD patients and 78 healthy controls from Shanxi Cancer Hospital.

[0087] This invention is divided into four phases: discovery, screening, modeling, and evaluation. A nested cohort analysis was performed on 350 patients with solitary pulmonary nodules (<30 mm) detected by low-dose computed tomography (LDCT).

[0088] 2. Identification of candidate piRNAs that are simultaneously upregulated in serum samples and tumor tissues of lung adenocarcinoma patients.

[0089] In this invention, a standardized bioinformatics workflow was used to systematically reanalyze small RNA sequencing datasets from the Gene Expression Omnibus Database (GEO). Figure 1 By employing stringent differential expression criteria (|log2 fold change (log2FC)| > 1, and a false discovery rate (FDR) < 0.05 after Benjamini-Hochberg correction), this invention reveals a significant piRNA expression dysregulation pattern in the pathogenesis of lung adenocarcinoma (LUAD). Analysis of tissue sample cohorts showed that 1699 piRNAs were upregulated and 786 were downregulated. Figure 2 Analysis of blood sample cohorts showed that 774 piRNAs were upregulated and 463 were downregulated. Figure 3 Considering that specifically upregulated piRNAs are easier to detect and have greater potential as clinical biomarkers, this invention focuses on screening 55 piRNAs that are upregulated in both tissue and blood samples. Subsequently, after rigorous screening (fold change > 5 and mean baseline expression > 10), this list was further refined to 10 candidate piRNAs with strong diagnostic potential. Figure 4 Receiver operating characteristic (ROC) curve analysis results showed that these selected biomarkers possessed strong discriminative power. Figure 5 Their differential expression patterns in organizations ( Figure 6 ) and blood ( Figure 7 The results were systematically validated in all sample cohorts, confirming that these piRNAs in lung adenocarcinoma samples were continuously upregulated compared to the normal control group.

[0090] 3. Validation in an internal screening cohort revealed two novel piRNAs for the detection of lung adenocarcinoma (LUAD).

[0091] To verify the differential expression of these 10 candidate piRNAs, relative quantitative analysis was performed in an independent local cohort. Preliminary tissue-based analysis showed that, compared with matched adjacent normal tissues, five piRNAs (piR-hsa-8393202, piR-hsa-8482068, piR-hsa-8429916, piR-hsa-8113890, and piR-hsa-4378145) were significantly upregulated in LUAD tumors (paired t-test). Figure 8 They exhibited good discriminative ability, with AUC values ​​of 0.829 (95% confidence interval: 0.743–0.916), 0.873 (0.801–0.945), 0.787 (0.693–0.880), 0.752 (0.651–0.852), and 0.708 (0.604–0.811), respectively. Figure 9 Subsequent serum analysis identified a distinct group of five upregulated piRNAs (piR-hsa-8393202, piR-hsa-4491510, piR-hsa-8429916, piR-hsa-8113890, piR-hsa-4378145) between LUAD patients and healthy controls (Mann-Whitney U test). Figure 10 The diagnostic AUC values ​​were 0.834 (0.721-0.956), 0.810 (0.683-0.938), 0.674 (0.518-0.829), 0.673 (0.516-0.831), and 0.687 (0.532-0.841), respectively. Figure 11 Notably, only piR-hsa-8393202 and piR-hsa-8429916 showed sustained upregulation in both tumor tissue and serum samples. This trans-tissue elevation highlights their potential as specific biomarkers for the diagnosis of LUAD.

[0092] 4. Construction of a LUAD diagnostic model based on two piRNAs

[0093] Relative quantitative assessments of piR-hsa-8393202 and piR-hsa-8429916 in serum samples from 711 patients with LUAD, 72 patients with benign lung disease (BPD), and 245 healthy controls (HC) were performed at the National Cancer Center. Figure 12 Violin plot analysis showed that serum expression of piR-hsa-8393202 and piR-hsa-8429916 was significantly increased in LUAD patients compared with BPD patients and HC individuals (all p < 0.001). Further intergroup comparisons revealed statistically significant differences in piR-hsa-8429916 and piR-hsa-8393202 expression between the HC and BPD groups. These results preliminarily validate the potential of these two piRNAs as diagnostic biomarkers for LUAD.

[0094] To verify the diagnostic capabilities of piR-hsa-8393202 and piR-hsa-8429916 in clinical applications, the included samples were randomly divided into a training cohort (n=718, n=497 / 221; LUAD / non-LUAD) and a non-overlapping test cohort (n=310, n=214 / 96; LUAD / non-LUAD). The two groups were well matched in terms of age and sex (P>0.05). Figure 13 It is worth noting that the samples mainly came from early stages, with patients with stage 0-II cancer accounting for 89.5% (445 / 497) and 86.9% (186 / 214) of the total patients in the training and testing cohorts, respectively.

[0095] This invention systematically employs nine machine learning algorithms, including logistic regression, support vector machine, random forest, gradient boosting, k-nearest neighbors, Naive Bayes, neural network, elastic network, and decision tree, to develop a predictive model. The algorithm performance was rigorously evaluated using AUC-response operating characteristic analysis. Figure 14 The results showed that the Random Forest algorithm had stronger discriminative ability, achieving an accuracy of 0.826, while the accuracy of other models ranged from 0.816 to 0.773 (p < 0.01). Therefore, due to its balanced precision-recall characteristics (F1 score = 0.891), the Random Forest algorithm was selected as the optimal modeling method.

[0096] To reduce overfitting and evaluate generalization ability, 10-fold cross-validation was performed in the training queue. The model performed consistently across all folds, with a mean AUC of 0.869 ± 0.02 (range: 0.762–0.941), which is highly consistent with the results on the test set. Figure 15The narrow confidence intervals (95% CI: 0.823–0.915) and extremely small standard deviations (±0.03) further confirm the stability of the model. Next, this invention applies SHAP values ​​to interpret the diagnostic model (…). Figure 16 Of the four variable features, piR-hsa-8429916 was the most important, with high expression levels (SHAP values: 0.2–0.6). The second most important feature, piR-hsa-8393202, showed a bidirectional effect (SHAP values: -0.2–0.4), indicating that it plays a context-dependent role in distinguishing between benign and malignant lesions. The effects of age and sex were negligible.

[0097] 5. Validation of LUAD diagnostic features based on 2-piRNA in a multicenter retrospective cohort.

[0098] This diagnostic model demonstrates strong discriminative ability for LUAD detection in a multi-center cohort. In the training set ( Figure 17 In test set A), receiver operating characteristic (ROC) analysis yielded an AUC of 0.923 (95% confidence interval: 0.905–0.942; P < 0.001). At the optimal cutoff of 0.157, the sensitivity was 78.1%, the specificity was 91.9%, the positive predictive value was 92.5%, and the negative predictive value was 55.6%. Internal validation in test set 1 (…) Figure 17 The external validation test set B yielded an AUC of 0.871 (95% confidence interval: 0.830–0.912; P < 0.001), with a sensitivity of 82.2%, specificity of 82.4%, positive predictive value of 93.5%, and negative predictive value of 57.3% (cutoff value = 0.363). Figure 17 The results were comparable (AUC = 0.865; 95% confidence interval: 0.825–0.906; P < 0.001), with a sensitivity of 90.8%, a specificity of 71.4%, a positive predictive value of 81.4%, and a negative predictive value of 84.7% (cutoff value = 0.140).

[0099] The classification accuracy was further quantified through confusion matrix analysis. Figure 18 The training set correctly identified 445 LUAD cases (true positives) and 159 non-LUAD controls (true negatives), with 62 false negatives and 52 false positives. Test set 1 showed 86 true positives, 144 true negatives, 10 false negatives, and 10 false positives, while test set 2 contained 43 true positives, 183 true negatives, 19 false positives, and 104 false negatives. Principal component analysis plot ( Figure 19The model's discriminative power was confirmed. On principal component 1 (PC1), there was a significant clustering between lung adenocarcinoma cases with high predictive value and non-lung adenocarcinoma control groups with low predictive value, supporting the model's biological interpretability.

[0100] Subsequent subgroup analyses presented in Table 1 show that the two piRNA-based model demonstrated superior diagnostic performance in patients with EGFR mutations compared to wild-type patients, across all cohorts. The mutation subgroups exhibited significantly higher AUC values: 0.930 vs. 0.916 in the training cohort, 0.875 vs. 0.868 in test set 1, and 0.929 vs. 0.882 in test set 2. These findings validate the model's robust diagnostic stability in genetically stratified populations. The consistently improved classification accuracy, particularly in EGFR-driven cases, further underscores the model's reliability in molecular subtyping of lung adenocarcinoma.

[0101] Table 1. Subgroup performance analysis of the model based on two piRNAs for lung adenocarcinoma diagnosis.

[0102]

[0103] Subgroup analysis based on carcinoembryonic antigen (CEA) status showed that in patients with elevated CEA levels, models based on two piRNAs exhibited significantly enhanced diagnostic accuracy: 0.958 vs. 0.924 in the training cohort, 0.928 vs. 0.857 in test set 1, and 0.870 vs. 0.866 in test set 2. Although CEA has high sensitivity for lung adenocarcinoma and non-neuroendocrine large cell lung cancer, its clinical application is limited due to its poor disease specificity, as elevated CEA levels are also seen in gastrointestinal malignancies and pulmonary fibrosis. The complementary properties of the two piRNAs with CEA suggest their potential for synergistic use in the early detection and histologically specific diagnosis of LUAD.

[0104] 6. Construction of an indeterminate lung nodule classification model based on 2-piRNA

[0105] This invention aims to evaluate the ability of a 2-piRNA classifier to effectively stratify IPNs. 350 samples showing lung nodules on low-dose computed tomography (LDCT) images were selected from study participants. Figure 20(A). This invention utilizes two piRNAs (piR-hsa-8429916 and piR-hsa-8393202), along with age, sex, and carcinoembryonic antigen (CEA) levels, to develop a random forest algorithm. The model is constructed using a training cohort containing 244 samples (62 benign nodules and 182 malignant nodules) and independently validated using another cohort of 106 samples (27 benign nodules and 79 malignant nodules). The model's stability is evaluated through 10-fold cross-validation on the training cohort, exhibiting a consistent receiver operating characteristic (ROC) trajectory across all iterations. Figure 20 (B) The pooled AUC was 0.885 (95% confidence interval: 0.863–0.907), indicating strong discriminative power and little tendency to overfit. Mean reduced Gini index analysis showed that the two piRNA markers were the main discriminative features, with variable importance significantly higher than demographic characteristics or traditional biomarkers. Figure 20 (C). Box plot distribution shows a significant difference in scores between the benign and malignant groups, both in the training set (C). Figure 20 (D) or test set ( Figure 20 In the 2-piRNA classifier, the diagnostic efficacy was superior in both cohorts, with an area under the curve (AUC) of 0.885 (95% confidence interval: 0.859–0.927) derived from receiver operating characteristic analysis. In the training cohort, it achieved a sensitivity of 77.5% and a specificity of 84.1%, significantly outperforming carcinoembryonic antigen (CEA), which had an AUC of 0.573. Figure 20 External validation showed that its performance remained excellent (AUC = 0.842, 95% confidence interval: 0.786-0.897; sensitivity = 78.5%, specificity = 82.4%), significantly better than CEA (AUC = 0.445). Figure 20 (G). Principal component analysis (PCA) visualization showed that pathologically confirmed benign nodules (green) and malignant nodules (red) were clearly separated spatially. Figure 20 (J and K). Furthermore, confusion matrix analysis further confirmed the diagnostic accuracy of this classifier for benign and malignant subclasses (J and K). Figure 20 The presence of H and I in the middle of the lung nodules highlights their clinical application value in lung nodule stratification.

[0106] 7. Validation and subgroup assessment of the 2-piRNA-based lung nodule classification model

[0107] To further evaluate the diagnostic performance of the 2-piRNA-based lung nodule classifier in different subgroups, further research was conducted. In subgroups stratified by carcinoembryonic antigen (CEA) levels ( Figure 21In subgroup a), the group with abnormal CEA levels achieved a remarkable area under the curve (AUC) of 0.966, a sensitivity of 93.1%, and a specificity of 88.9%. This finding indicates that its diagnostic efficiency was significantly higher than that of the group with normal CEA levels, whose AUC was 0.875. In subgroups grouped by nodule size (… Figure 21 In subgroup analysis (b), the AUC of nodules ≥0.8 cm in diameter reached 0.897, with a sensitivity of 78.6% and a specificity of 89.7%. In contrast, the diagnostic performance was lower in the nodule diameter group (less than 0.8 cm), with an AUC of 0.796. In subgroup analysis categorized by nodule type (…),… Figure 21 In the partial solid nodule group (c), the highest AUC value of 0.906 was achieved, with a sensitivity of 83.5% and a specificity of 87.0%. The area under the curve (AUC) of the pure ground-glass nodule group was 0.829, while that of the pure solid nodule group was 0.832, highlighting the difference in diagnostic ability among different nodule types. Carcinoembryonic antigen (CEA) and cytokeratin 19 fragment (CYFRA21-1) are the most well-studied laboratory markers in lung nodule classification. To further evaluate its diagnostic performance, this invention compared a classifier based on two piRNAs with established laboratory markers. Importantly, in all subgroup analyses, the classifier consistently outperformed CEA and CYFRA21-1 detection in distinguishing between benign and malignant lung nodules (Table 2). In summary, non-invasive detection of two piRNAs and low-dose computed tomography (LDCT) can reduce false positives in LDCT and accurately stratify lung adenocarcinoma (LUAD) for treatment. Figure 21 (d).

[0108] Table 2. Subgroup performance of intraductal papillary mucinous neoplasm (IPN) classification based on 2-piRNAs classifier, serum carcinoembryonic antigen (CEA) level, and cytokeratin 19 fragment (CYFRA21-1) level.

[0109]

[0110] 8. piR-hsa-8393202 and piR-hsa-8429916 exhibited tumor-derived characteristics and enhanced stability in serum samples.

[0111] This invention evaluated the stability of two piRNAs in repeated freeze-thaw cycles. Figure 22 Serum samples that underwent three freeze-thaw cycles contained piR-hsa-8393202. Figure 23 (A) and piR-hsa-8429916 ( Figure 23 The levels of B in the middle were consistent by RT-qPCR quantitative analysis.

[0112] To determine their tumor-specific origin, this invention analyzed paired serum samples from 48 patients with LUAD before and after surgery. Quantitative analysis showed that the levels of these two piRNAs were significantly reduced after tumor resection (piR-hsa-8393202: P<0.001; piR-hsa-8429916: P<0.001). Figure 24 A slight positive correlation was found between serum piRNA levels and levels in matched tumor tissues (piR-hsa-8393202: R). 2 =0.287, P<0.001; piR-hsa-8429916: R 2 =0.187, P=0.002)( Figure 25 It is noteworthy that the degree of reduction after surgery was positively correlated with tumor volume (piR-hsa-8393202: R). 2 =0.687, P<0.001; piR-hsa-8429916: R 2 =0.338, P<0.001)( Figure 26 This further supports their tumor origin.

[0113] Cell culture experiments confirmed the specific expression and secretion by tumor cells. 2×10⁶ cells were cultured. 6 Adherent cells were cultured in 10 cm culture dishes for 12 hours, after which cells and cell-free culture medium were collected. Compared with the normal bronchial epithelial cell line (BEAS-2B, n=3), two piRNAs were significantly overexpressed in the lung adenocarcinoma cell line. Figure 27 China A and Figure 28 (A). Notably, the levels of piR-hsa-8393202 and piR-hsa-8429916 in the conditioned medium of lung adenocarcinoma cells were significantly higher than those in the supernatant of BEAS-2B cells. Figure 27 China B and Figure 28 (B) indicates that tumor cells have active secretion.

[0114] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made by those skilled in the art to the technical solutions of the present invention without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. Use of reagents for detecting biomarker or biomarker combination expression levels in serum or lung tissue in the manufacture of a product for the early diagnosis of lung adenocarcinoma, characterized in that, The biomarker is piR-hsa-8429916 or piR-hsa-8393202. The biomarker combination is piR-hsa-8429916 and piR-hsa-8393202.

2. Use according to claim 1, characterized in that, The lung adenocarcinoma early diagnosis product is a kit.

3. Use of reagents for detecting biomarker or biomarker combination expression levels in serum or lung nodule tissue in the manufacture of a test product for differentiating between benign and malignant lung nodules, characterized in that, The biomarker is piR-hsa-8429916 or piR-hsa-8393202. The biomarker combination is piR-hsa-8429916 and piR-hsa-8393202.

4. Use according to claim 3, characterized in that, The detection product is a kit.