A snRNA and snoRNA biomarker combination for diagnosing or aiding in the diagnosis of gastric cancer and uses thereof
Patent Information
- Application Number
- CN202610700509.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-20
- Publication Date
- 2026-09-25
AI Technical Summary
然而,中国尚未有相关基于大规模验证数据构建的游离snRNA用于胃癌早期诊断的标志物模型
一、本发明提供的snRNA和snoRNA生物标志物组合作为胃癌诊断标志物,其特异性好,能准确鉴别胃癌患者;敏感性高,能早期检测出癌症患者且诊断效率高,为诊断胃癌提供了新的临床手段,具有良好的转化医学前景。
Smart Images

Figure CN122811362A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biomedical technology, specifically to a combination of snRNA and snoRNA biomarkers for the diagnosis or auxiliary diagnosis of gastric cancer and their applications. Background Technology
[0003] Currently, clinical guidelines recommend gastric cancer screening methods including serum tumor marker testing, imaging examinations (such as barium meal X-ray, CT, MRI, etc.), and endoscopy and tissue biopsy. These methods play an important role in the diagnosis of gastric cancer, but each has certain limitations. Traditional serological markers, such as carcinoembryonic antigen (CEA) and carbohydrate antigen 19-9 (CA19-9), although widely used in the auxiliary diagnosis and efficacy monitoring of gastric cancer, have low sensitivity and specificity for early gastric cancer. For example, studies have shown that the sensitivities of CEA (cutoff value 3 ng / mL) and CA19-9 (cutoff value 37 U / mL) for the diagnosis of gastric cancer are 20%-30% and 30%-40%, respectively, and the specificities are 70%-80% and 80%-90%, respectively. While imaging examinations and endoscopy can detect gastric lesions relatively accurately, the former carries the risk of radiation exposure, and the latter is an invasive procedure with poor patient compliance and high cost, making it difficult to use as a routine means for large-scale population screening. Therefore, in order to meet clinical needs, there is an urgent need to develop a new method for the early diagnosis of gastric cancer that is highly sensitive, specific, non-invasive, convenient, and economical.
[0004] Liquid biopsy is a breakthrough technology for tumor screening and diagnosis. It has unique advantages such as being non-invasive, providing real-time dynamic detection, and overcoming tumor heterogeneity. It enables gastric cancer patients to receive treatment as early as possible, achieve radical treatment through surgery, improve patient survival rates, and reduce treatment costs. It has significant clinical significance and social benefits.
[0005] Liquid biopsy, as an emerging tumor screening and diagnostic technology, has advantages such as being non-invasive, repeatable, allowing for real-time dynamic monitoring, and overcoming tumor heterogeneity. It has shown great application potential in early tumor diagnosis, efficacy evaluation, and prognostic monitoring.
[0006] In recent years, various liquid biopsy biomarkers, such as circulating tumor cells (CTCs), circulating tumor DNA (ctDNA), and exosomes, have been used in gastric cancer research. Among them, circulating small non-coding RNAs (sncRNAs) have attracted much attention due to their stability in the blood. sncRNAs, including microRNAs (miRNAs), have been shown to be closely related to tumor development and progression, and have the potential to become novel tumor biomarkers. Cell-free small nuclear RNAs (snRNAs) in the blood are a class of small RNA molecules, typically 60-300 nucleotides in length, mainly located in the cell nucleus. They are a core component of the spliceosome and participate in the splicing process of pre-mRNA. snRNAs bind to proteins to form complexes, recognize and cleave introns, and link exons, thereby generating mature mRNA. Although the potential of snRNAs as biomarkers has not yet been widely validated, their core function and high specificity in RNA splicing offer possibilities for their application in disease diagnosis and treatment. Similar to snRNA, snoRNA (small nucleolar RNA) is also a class of small non-coding RNAs with important functions. They are typically 60-300 nucleotides in length and are mainly located in the nucleolus. snoRNAs bind to specific proteins to form snoRNP complexes, participating in the processing and modification of ribosomal RNA (rRNA) (such as methylation and pseudouridineization), thereby maintaining the structure and function of ribosomes. Recent studies have found that snoRNAs can also regulate the expression of certain mRNAs and miRNAs and are closely related to tumorigenesis and development. Abnormal expression of snoRNAs in certain tumor tissues is thought to be related to the proliferation, migration, and invasion capabilities of cancer cells.
[0007] Although snRNA and snoRNA have different functions, they both play crucial roles in RNA metabolism and processing, and both function as RNA-protein complexes. More importantly, the stability and specificity of snRNA and snoRNA in blood make them potential biomarkers for liquid biopsy. Currently, studies have shown that certain snoRNAs exhibit specific expression patterns in the blood of cancer patients, while research on free snRNA is still limited. Exploring the diagnostic value of snRNA and snoRNA in gastric cancer by combining their functions and characteristics will not only help elucidate their biological mechanisms but may also provide novel biomarkers for the early diagnosis of gastric cancer. However, China currently lacks biomarker models for the early diagnosis of gastric cancer based on large-scale validation data using free snRNA.
[0008] In view of this, a combination of snRNA and snoRNA biomarkers for the diagnosis or auxiliary diagnosis of gastric cancer is proposed and its application is described. Summary of the Invention
[0009] To address the aforementioned technical problems, this invention provides a combination of snRNA and snoRNA biomarkers for the diagnosis or auxiliary diagnosis of gastric cancer and their applications.
[0010] To achieve the above objectives, the present invention provides the following technical solution: A combination of snRNA and snoRNA biomarkers for the diagnosis or auxiliary diagnosis of gastric cancer, wherein the nucleotide sequences of the snRNA biomarker combination are shown in SEQ ID No. 1-2, and the nucleotide sequence of the snoRNA biomarker combination is shown in SEQ ID No. 3.
[0011] The application of a combination of target snRNA and snoRNA biomarkers in a sample in the preparation of a kit for the early diagnosis or auxiliary diagnosis of gastric cancer, wherein the nucleotide sequences of the snRNA biomarker combination are shown in SEQ ID No. 1-2, and the nucleotide sequence of the snoRNA biomarker combination is shown in SEQ ID No. 3.
[0012] Preferably, the reagents for detecting the target snRNA and snoRNA in the sample include primers and probes for detecting the target snRNA and snoRNA.
[0013] A kit for diagnosing or assisting in the diagnosis of gastric cancer, the kit comprising primers and probes for detecting target snRNA and target snoRNA; the nucleotide sequences of the snRNA biomarker combination are shown in SEQ ID No. 1-2, and the nucleotide sequences of the snoRNA biomarker combination are shown in SEQ ID No. 3.
[0014] A method for gastric cancer tumor screening and risk calculation includes the following steps: S1: Serum sample collection and processing; S2: RNA sequencing library construction and quantification of snRNA and snoRNA; S3: Differential analysis and classification performance evaluation and screening of characteristic snRNAs and snoRNAs; S4: Integrated model building; S5: Independent validation set performance verification and risk calculation.
[0015] As a preferred option, in step S1, clinical samples are collected in two stages, CF1 and CF2, to obtain the experimental set and the independent validation set, respectively. CF1 stage: Blood samples were collected using EDTA anticoagulant tubes, and plasma was separated by centrifugation at 12000g for 10 minutes within 1 hour. The plasma was aliquoted into RNase-free cryovials as experimental sets and stored at -80°C until RNA extraction. CF2 phase: Blood samples were collected using the same method as an independent validation set and stored until RNA extraction.
[0016] Preferably, step S2 includes the following steps: S21: The 3' end of RNA is ligated using a 5'-adenosylated / 3'-blocked single-stranded DNA adapter. An RT primer containing UMI is introduced to bind the residual adapter, followed by 5' end ligation. RNA is then extracted from the serum sample from step S1. S22: cDNA was synthesized and amplified using high-sensitivity polymerase reverse transcription, and the 110-140 bp PCR product was recovered by PAGE gel. S23: Further circularized into a single-stranded DNA library, and after quality control, single-end 50 bp high-throughput sequencing was performed based on the DNBSEQ platform; S24: After sequencing, the data undergoes quality control to obtain clean read data. The clean read sequences are aligned to the human reference genome GRCh38.p14. After the data alignment is completed, the data is quantified to calculate the expression abundance of snRNA and snoRNA, retaining the characteristic of average expression level (expressed as read count) greater than 5.
[0017] Preferably, in step S3... Differentially expressed snRNAs and snoRNAs were screened using an empirical Bayesian method. The p-value was corrected using the Benjamini-Hochberg method, and candidate biomarkers were screened based on the following criteria: AUC > 0.75, |LogFC| > 1, and average expression level > 3 CPM. We further screened for characteristic snRNAs and snoRNAs using Lasso regression to obtain the final characteristic biomarkers.
[0018] Preferably, in step S4, the construction of the stacked ensemble model involves: first, generating a prediction probability matrix based on the base learner; then, using the generalized linear model (GLM) as a meta-learner, integrating the classification outputs of the base model through a weighted fusion strategy.
[0019] As a preferred option, in step S5, a blinded validation is performed using a CF2 stage independent sample cohort to evaluate the detection performance of the ensemble model for gastric cancer, and a risk value is calculated based on the ensemble model.
[0020] Compared with existing technologies, this invention provides a combination of snRNA and snoRNA biomarkers for the diagnosis or auxiliary diagnosis of gastric cancer and their applications, which have the following beneficial effects: I. The combination of snRNA and snoRNA biomarkers provided by this invention, as diagnostic biomarkers for gastric cancer, has good specificity and can accurately identify gastric cancer patients; it also has high sensitivity, can detect cancer patients at an early stage, and has high diagnostic efficiency, providing a new clinical means for diagnosing gastric cancer and has good prospects for translational medicine.
[0021] II. This invention constructs a gastric cancer diagnostic model based on snRNA and snoRNA biomarkers. This model can be used for early screening and auxiliary diagnosis of gastric cancer, and can provide accurate assessment of gastric cancer risk, thereby promoting early treatment and management decisions.
[0022] The features and advantages of the present invention will be described in detail through embodiments and in conjunction with the accompanying drawings. Attached Figure Description
[0023] Figure 1 This is a diagram showing the expression of the snRNA and snoRNA biomarkers of this invention in CF1 samples; Figure 2 This is a graph showing the performance of the gastric cancer diagnostic model based on snRNA and snoRNA biomarkers of this invention on the cross-validation set. Figure 3 This is a graph showing the performance of the gastric cancer diagnostic model based on snRNA and snoRNA biomarkers of this invention on an independent validation set. Figure 4 This is a risk score distribution diagram of the model of this invention on the independent validation set. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. However, it should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of the invention. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concept of the invention.
[0025] Referring to the figure, this invention proposes a combination of snRNA and snoRNA biomarkers for the diagnosis or auxiliary diagnosis of gastric cancer and their applications: Example 1: A combination of snRNA and snoRNA biomarkers for the diagnosis or auxiliary diagnosis of gastric cancer, wherein the nucleotide sequences of the snRNA biomarker combination are shown in SEQ ID No. 1-2, and the nucleotide sequences of the snoRNA biomarker combination are shown in SEQ ID No. 3.
[0026] Example 2: Application of a combination of target snRNA and snoRNA biomarkers in a sample in the preparation of a kit for early diagnosis or auxiliary diagnosis of gastric cancer, wherein the nucleotide sequences of the snRNA biomarker combination are shown in SEQ ID No. 1-2, and the nucleotide sequence of the snoRNA biomarker combination is shown in SEQ ID No. 3.
[0027] Preferably, the reagents for detecting the target snRNA and snoRNA in the sample include primers and probes for detecting the target snRNA and snoRNA.
[0028] Example 3: A kit for diagnosing or assisting in the diagnosis of gastric cancer, the kit comprising primers and probes for detecting target snRNA and target snoRNA; the nucleotide sequences of the snRNA biomarker combination are shown in SEQ ID No. 1-2, and the nucleotide sequences of the snoRNA biomarker combination are shown in SEQ ID No. 3.
[0029] Example 4: A method for gastric cancer tumor screening and risk calculation, comprising the following steps: S1: Serum sample collection and processing; S2: RNA sequencing library construction and quantification of snRNA and snoRNA; S3: Differential analysis and classification performance evaluation and screening of characteristic snRNAs and snoRNAs; S4: Integrated model building; S5: Independent validation set performance verification and risk calculation.
[0030] As a preferred option, in step S1, clinical samples are collected in two stages, CF1 and CF2, to obtain the experimental set and the independent validation set, respectively. CF1 stage: Blood samples were collected using EDTA anticoagulant tubes, and plasma was separated by centrifugation at 12000g for 10 minutes within 1 hour. The plasma was aliquoted into RNase-free cryovials as experimental sets and stored at -80°C until RNA extraction. CF2 phase: Blood samples were collected using the same method as an independent validation set and stored until RNA extraction.
[0031] It should be noted that, after approval by the Ethics Committee of the First Affiliated Hospital of Wenzhou Medical University and obtaining informed consent from the participants, clinical samples were collected in two phases: CF1 and CF2. Phase CF1 sampling involved collecting peripheral blood samples from 30 pathologically confirmed gastric cancer patients before surgery, and peripheral blood samples from 30 age / sex-matched healthy controls (without a history of gastric disease or malignancy). Phase CF2 sampling involved collecting blood samples from both 30 gastric cancer patients and 30 healthy controls.
[0032] Preferably, step S2 includes the following steps: S21: The 3' end of RNA is ligated using a 5'-adenosylated / 3'-blocked single-stranded DNA adapter. An RT primer containing UMI is introduced to bind the residual adapter, followed by 5' end ligation. RNA is then extracted from the serum sample from step S1. S22: cDNA was synthesized and amplified using high-sensitivity polymerase reverse transcription, and the 110-140 bp PCR product was recovered by PAGE gel. S23: Further circularized into a single-stranded DNA library, and after quality control, single-end 50 bp high-throughput sequencing was performed based on the DNBSEQ platform; S24: After sequencing, clean reads were obtained through quality control. These clean reads were aligned to the human reference genome GRCh38.p14. After alignment, the data was quantified to calculate the expression abundance of snRNA and snoRNA, retaining features with an average expression level (expressed as read count) greater than 5. Features with an average expression level below this threshold were generally considered to have high sequencing noise or low biological significance and were therefore discarded. The quantification results of snRNA and snoRNA in the two-stage samples are shown in Table 1.
[0033] In an optional implementation, high-quality data output is ensured by filtering low-quality reads (Phred<13 bases >0.1%), pruning adapter sequences, filtering reads with lengths other than 15-44 bp, removing full PolyA sequences, limiting high adenine content, and eliminating reads containing non-standard bases through quality control. Using Bowtie software (version v1.3.1), the clean read sequences were aligned to the human reference genome GRCh38.p14. The FASTA file of the genome assembly and the gene annotation files in GTF format were downloaded from the Ensembl database. Quantization was performed using the featureCounts software (Subread toolkit v2.0.6), and the annotation files were obtained from the GENCODE database and the snoDB database.
[0034] Table 1. Quantification results of snRNA and snoRNA in two-stage samples Preferably, in step S3, differentially expressed snRNAs and snoRNAs are screened based on empirical Bayesian methods, p-values are corrected using the Benjamini-Hochberg method, and candidate biomarkers are screened using the following criteria: AUC>0.75, |LogFC|>1, and average expression level>3 CPM; The AUC value and 95% confidence interval of the ROC curve were calculated using the pROC package (1000 bootstrap resampling). The optimal threshold was determined using the Youden index. Sensitivity, specificity, positive / negative predictive value, precision, and accuracy were calculated simultaneously.
[0035] We further screened for characteristic snRNAs and snoRNAs using Lasso regression to obtain the final characteristic biomarkers.
[0036] Expression profiles of differentially expressed snRNAs and snoRNAs were used as independent variables, with the gastric cancer group versus the control group as the dependent variable. Five-fold cross-validation was employed to select the optimal penalty parameter λ. A Lasso regression model was constructed based on the optimal λ value to obtain the regression coefficient for each snRNA and snoRNA. SnRNAs and snoRNAs with non-zero regression coefficients were selected as the final feature snRNAs and snoRNAs. To ensure stability, the selection was repeated 100 times, and snRNAs and snoRNAs selected more than 95% of the time were retained as feature biomarkers.
[0037] Table 2. Summary of the classification performance of snRNA and snoRNA biomarkers Preferably, in step S4, the construction of the stacked ensemble model involves: first, generating a prediction probability matrix based on the base learner; then, using the generalized linear model (GLM) as a meta-learner, integrating the classification outputs of the base model through a weighted fusion strategy.
[0038] In an optional implementation, to improve the generalization ability and prediction accuracy of the gastric cancer diagnostic model, Random Forest, Naïve Bayes, Gradient Boosting Decision Tree (XGBoost), Deep Neural Network, and Logistic Regression are used as base learners; and the hyperparameters of each model are optimized through five-fold cross-validation.
[0039] like Figure 2 As shown, the results indicate that the ensemble model achieved an AUC of 0.999 (95% CI: 0.995-1.000) in cross-validation, which is 6.8% better than the best-performing single base model (AUC=0.931).
[0040] As a preferred option, in step S5, a blinded validation is performed using a CF2 stage independent sample cohort to evaluate the detection performance of the ensemble model for gastric cancer, and a risk value is calculated based on the ensemble model.
[0041] like Figure 4 As shown, the ensemble model achieved an AUC of 0.933 (95% CI: 0.866–0.999) on the independent validation set, higher than any single model. This result demonstrates that the ensemble model, through collaborative optimization using multiple algorithms, effectively overcomes the performance degradation problem of single models across cohorts of data, exhibiting high robustness in distinguishing gastric cancer patients from healthy individuals in real-world clinical scenarios and possessing the potential for large-scale screening applications.
[0042] It should be noted that the snRNA and snoRNA biomarker combinations of the present invention include, but are not limited to, the following specific snRNA and snoRNA molecules: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions or improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A combination of snRNA and snoRNA biomarkers for diagnosing or assisting in the diagnosis of gastric cancer, characterized in that: The nucleotide sequences of the snRNA biomarker combination are shown in SEQ ID No. 1-2, and the nucleotide sequences of the snRNA biomarker combination are shown in SEQ ID No.
3.
2. The application of the combination of target snRNA and snoRNA biomarkers in a sample in the preparation of a kit for the early diagnosis or auxiliary diagnosis of gastric cancer, characterized in that: The nucleotide sequences of the snRNA biomarker combination are shown in SEQ ID No. 1-2, and the nucleotide sequences of the snRNA biomarker combination are shown in SEQ ID No.
3.
3. The application as described in claim 2, characterized in that: The reagents used to detect the target snRNA and snoRNA in the sample include primers and probes for detecting the target snRNA and snoRNA.
4. A reagent kit for diagnosing or assisting in the diagnosis of gastric cancer, characterized in that: The kit includes primers and probes for detecting target snRNA and target snoRNA; the nucleotide sequences of the snRNA biomarker combination are shown in SEQ ID No. 1-2, and the nucleotide sequence of the snoRNA biomarker combination is shown in SEQ ID No.
3.
5. A method for gastric cancer tumor screening and risk calculation, characterized in that, Includes the following steps: S1: Serum sample collection and processing; S2: RNA sequencing library construction and quantification of snRNA and snoRNA; S3: Differential analysis and classification performance evaluation and screening of characteristic snRNAs and snoRNAs; S4: Integrated model building; S5: Independent validation set performance verification and risk calculation.
6. The gastric cancer tumor screening and risk calculation method as described in claim 5, characterized in that: In step S1, clinical samples are collected in two phases, CF1 and CF2, to obtain the experimental set and the independent validation set, respectively. CF1 stage: Blood samples were collected using EDTA anticoagulant tubes, and plasma was separated by centrifugation at 12000g for 10 minutes within 1 hour. The plasma was aliquoted into RNase-free cryovials as experimental sets and stored at -80°C until RNA extraction. CF2 phase: Blood samples were collected using the same method as an independent validation set and stored until RNA extraction.
7. The gastric cancer tumor screening and risk calculation method as described in claim 5, characterized in that: Step S2 includes the following steps: S21: The 3' end of RNA is ligated using a 5'-adenosylated / 3'-blocked single-stranded DNA adapter. An RT primer containing UMI is introduced to bind the residual adapter, followed by 5' end ligation. RNA is then extracted from the serum sample from step S1. S22: cDNA was synthesized and amplified using high-sensitivity polymerase reverse transcription, and the 110-140 bp PCR product was recovered by PAGE gel. S23: Further circularized into a single-stranded DNA library, and after quality control, single-end 50 bp high-throughput sequencing was performed based on the DNBSEQ platform; S24: After sequencing, the data undergoes quality control to obtain clean read data. The clean read sequences are aligned to the human reference genome GRCh38.p14. After the data alignment is completed, the data is quantified to calculate the expression abundance of snRNA and snoRNA, retaining the characteristic of average expression level (expressed as read count) greater than 5.
8. The gastric cancer tumor screening and risk calculation method as described in claim 5, characterized in that: In step S3, Differentially expressed snRNAs and snoRNAs were screened using an empirical Bayesian method. The p-value was corrected using the Benjamini-Hochberg method, and candidate biomarkers were screened based on the following criteria: AUC > 0.75, |LogFC| > 1, and average expression level > 3 CPM. We further screened for characteristic snRNAs and snoRNAs using Lasso regression to obtain the final characteristic biomarkers.
9. The gastric cancer tumor screening and risk calculation method as described in claim 5, characterized in that: In step S4, the construction of the stacked ensemble model involves: first, generating a prediction probability matrix based on the base learner; then, using the generalized linear model (GLM) as a meta-learner, integrating the classification outputs of the base model through a weighted fusion strategy.
10. The method for gastric cancer tumor screening and risk calculation as described in claim 6, characterized in that: In step S5, a blinded validation was performed using the CF2 stage independent sample cohort to evaluate the detection performance of the ensemble model for gastric cancer, and the hazard value was calculated based on the ensemble model.