Plasma pirna combination for early diagnosis of gastric cancer and use thereof

By combining plasma piRNA and using a Lasso Logistic regression model, the invasiveness and accuracy issues of gastric cancer screening and diagnosis were resolved, achieving efficient and accurate early diagnosis of gastric cancer while reducing testing costs.

WO2025217982A1PCT designated stage Publication Date: 2025-10-23THE FIRST AFFILIATED HOSPITAL OF WENZHOU MEDICAL UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/095385
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-16
Filing Date
2024-05-27
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Existing gastric cancer screening and diagnosis methods are highly invasive, costly, and lack accuracy. In particular, traditional pathological examinations are inconvenient for patients, and the development of non-invasive biomarkers has not been fully utilized.

Method used

Using a combination of plasma piRNAs, including hsa-piR-32885, hsa-piR-3440, and hsa-piR-786, a Lasso Logistic regression model was established to predict gastric cancer risk through blood testing, reducing testing costs and improving accuracy.

Benefits of technology

It achieves high sensitivity and high specificity in the early diagnosis of gastric cancer, reduces detection costs, avoids the risks of invasive examinations, and the model's prediction accuracy reaches over 90%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024095385_23102025_PF_FP_ABST
    Figure CN2024095385_23102025_PF_FP_ABST
Patent Text Reader

Abstract

A plasma piRNA combination for early diagnosis of gastric cancer is obtained on the basis of the plasma piRNA expression profile of the Chinese population, and a model for early diagnosis of gastric cancer is established on the basis of the plasma piRNA combination. The model can relatively accurately predict the risk of gastric cancer onset, and help to reduce the detection cost. Moreover, a Lasso Logistic regression model is used in the present invention, which greatly reduces the number of variables included in the model, and facilitates the application and popularization of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Plasma piRNA combination for early diagnosis of gastric cancer and application TECHNICAL FIELD

[0001] The present application relates to the technical field of biological information, in particular to a plasma piRNA combination for early diagnosis of gastric cancer and application. BACKGROUND

[0002] Gastric cancer is one of the common malignant tumors in China. In recent years, the incidence of gastric cancer in China is gradually increasing, which seriously threatens the life and health of residents. Screening, early diagnosis and early treatment of high-risk groups of gastric cancer can effectively reduce the incidence and mortality of gastric cancer.

[0003] The traditional screening and diagnosis mode of gastric cancer is: hematology preliminary screening, and the positive preliminary screening is diagnosed by the gold standard method of endoscopic biopsy. Although the traditional pathological examination has certain accuracy, it is an invasive method, which brings many inconveniences to patients, and even causes complications and sequelae such as bleeding and infection. Compared with pathological examination, the development of blood markers brings convenience to the screening and diagnosis of gastric cancer, and the results are more stable, and have the characteristics of patient compliance. The wide application of existing pepsinogen, gastrin 17, HP and tumor markers shows the potential of biomarkers in the screening and diagnosis of gastric cancer. Therefore, it is urgent to explore more effective, accurate and sensitive non-invasive clinical biomarkers for early screening and diagnosis of gastric cancer.

[0004] PiRNA is a non-coding RNA with a length of 25-33 nt, which is closely related to genome stability, epigenetic regulation, reproductive stem cell differentiation, embryonic development and the occurrence and development of various diseases. Therefore, piRNA has the potential to be a biomarker. At present, the related content of piRNA in gastric cancer patients still needs further research, and better understanding of the role of piRNA in the occurrence and development of gastric cancer can excavate new biomarkers for early screening and diagnosis of gastric cancer, which plays an important role in early screening and diagnosis of gastric cancer patients.

[0005] SUMMARY

[0006] The purpose of the present application is to provide a plasma piRNA combination for early diagnosis of gastric cancer and application. The present application provides a plasma piRNA combination for early diagnosis of gastric cancer, and establishes and verifies a model for early diagnosis of gastric cancer according to the results, which is convenient for early screening.

[0007] The technical scheme of the present application: a plasma piRNA combination for early diagnosis of gastric cancer, wherein the plasma piRNAs in the plasma piRNA combination include hsa-piR-32885, hsa-piR-3440, hsa-piR-786, hsa-piR-12390, hsa-piR-414, hsa-piR-23197, hsa-piR-32911, hsa-piR-32945, hsa-piR-28060, hsa-piR-7096, hsa-piR-32870 and hsa-piR-30778.

[0008] The application of the above-mentioned plasma piRNA combination for early diagnosis of gastric cancer in constructing an early diagnosis model of gastric cancer.

[0009] The application, wherein the early diagnosis model of gastric cancer is a Lasso Logistic regression model.

[0010] The application, wherein the mathematical expression of the Lasso Logistic regression model is as follows: gastric cancer incidence risk score = ∑ (plasma piRNA expression value x regression coefficient).

[0011] The application, wherein the regression coefficients are as follows:

[0012] The regression coefficient of hsa-piR-32885 is -0.001;

[0013] The regression coefficient of hsa-piR-3440 is -0.039;

[0014] The regression coefficient of hsa-piR-786 is -0.007;

[0015] The regression coefficient of hsa-piR-12390 is -0.001;

[0016] The regression coefficient of hsa-piR-414 is -0.002;

[0017] The regression coefficient of hsa-piR-23197 is 0.037;

[0018] The regression coefficient of hsa-piR-32911 is 0.008;

[0019] The regression coefficient of hsa-piR-32945 is 0.004;

[0020] The regression coefficient of hsa-piR-28060 is 0.001;

[0021] The regression coefficient of hsa-piR-7096 is 0.002;

[0022] The regression coefficient of hsa-piR-32870 is 0.002;

[0023] The regression coefficient of hsa-piR-30778 is 0.026.

[0024] The application of the foregoing, the Lasso Logistic regression model is constructed as follows:

[0025] (1) Collecting plasma of gastric cancer patients and healthy control population, respectively extracting plasma free piRNA; (2) obtaining plasma piRNA expression profile by using piRNA transcriptome sequencing; (3) randomly dividing the gastric cancer patients and healthy control population into training set and test set, establishing Lasso Logistic regression model in the training set, and obtaining the piRNA regression coefficient included in the model; (4) based on the Lasso Logistic regression model established in the training set, evaluating the prediction accuracy of the model in the test set by using ROC curve, sensitivity and specificity index.

[0026] Compared with the prior art, the innovation of the present application is that based on the plasma piRNA expression profile of Chinese population, a plasma piRNA combination for early diagnosis of gastric cancer is obtained, and an early diagnosis model of gastric cancer is established based on the plasma piRNA combination, the area under the ROC curve (area under the ROC curve, ATC) of the model for predicting gastric cancer is 0.96, the sensitivity is 90%, and the specificity is 96%, which can effectively distinguish gastric cancer patients from healthy controls. In addition, the present application uses Lasso Logistic regression model, which greatly reduces the number of variables included in the model, which will help to reduce the cost of detection and promote the application of the model. The present application is detected by blood, which is convenient, fast, small in trauma area, easy to use and good in stability. Only 200ul of plasma is needed each time, which avoids the risk of physical damage caused by multiple tissue biopsies to obtain tumor tissue for pathological identification by gastroscope, greatly reducing the detection cost BRIEF DESCRIPTION OF DRAWINGS

[0027] Fig. 1 is a flowchart of the technical scheme of the present application.

[0028] Fig. 2 is a graph of the relationship between the regularization parameter λ and the partial likelihood estimation bias in the Lasso Logistic regression model.

[0029] Fig. 3 is a schematic diagram of the ROC curve of the model in the training set.

[0030] Fig. 4 is a schematic diagram of the ROC curve of the model in the test set. DETAILED DESCRIPTION

[0031] The application will be further described in connection with the accompanying drawings and examples, but not as the basis for limiting the application.

[0032] Example: Construction and verification of gastric cancer early diagnosis model based on plasma piRNA expression profile.

[0033] This example is divided into four parts: collection of plasma samples, Qiagen miRNeasy plasma kit for extraction of plasma small RNA, small RNA transcriptome sequencing, and establishment of Lasso Logistic regression model based on plasma piRNA expression matrix. The flow of this example is shown in Figure 1.

[0034] (1) Collect plasma samples;

[0035] Collect 5ml of whole blood from 200 gastric cancer patients and 100 healthy controls into blood collection tubes containing EDTA anticoagulant. After collection, repeatedly invert the blood collection tube to mix the EDTA anticoagulant and blood thoroughly. Centrifuge at 3000rpm, 4℃ for 10min, and the supernatant is the plasma. Take 2ml of plasma and store it in an EP tube at -80℃.

[0036] (2) Qiagen miRNeasy plasma kit for extraction of plasma small RNA;

[0037] Cell lysis and small RNA extraction: Add 1 ml of QIAzol lysis reagent to 200 μl of sample, vortex or invert the liquid several times, and incubate at room temperature (15-25 °C) for 5 min. Add 200 μl of chloroform, shake vigorously for 15 s, and incubate at room temperature for 2-3 min. After incubation, centrifuge at 4 °C and 12,000 g for 15 min. After centrifugation, transfer the upper aqueous phase to a new EP tube, add 1.5 times the volume of the aqueous phase of 100% ethanol, and mix them thoroughly by inverting. Take 700 μL of liquid to the RNeasy MinElute spin column (column), centrifuge at 8000 g for 15 s at room temperature, and discard the waste in the collection tube. Repeat the above steps once with the remaining liquid. Add 700 μl of Buffer RWT to the RNeasy MinElute spin column (column), centrifuge at 8000 g for 15 s, and discard the waste in the collection tube. Pipette 500 μl of buffer RPE into the RNeasy MinElute spin column (column), centrifuge at 8000 g for 15 s, and discard the waste in the collection tube. Add 500 μl of 80% ethanol to the RNeasy MinElute spin column (column), centrifuge at 8000 g for 2 min, and discard the waste in the collection tube. Place the RNeasy MinElute spin column (column) in a new 2 ml EP tube, open the spin column cap, set the centrifuge to maximum speed, centrifuge for 5 min, discard the waste in the collection tube and the collection tube. Place the RNeasy MinElute spin column (column) in a new 1.5 ml collection tube.

[0038] Small RNA elution: Add 15 μL of RNase-free water to the filter membrane, gently cover the tube cap, stand at room temperature for 2 min, and centrifuge at maximum speed for 1 min. The bottom of the tube is the separated RNA.

[0039] RNA concentration and RNA integrity evaluation: Take 1 μL for Aglient 2100 RNA Pico chip detection, and the peak value is generally below 200 nt. Only high-quality RNA samples (RIN≥7, >50 ng / μL, OD260 / 280 between 1.8-2.2) are used to construct sequencing libraries.

[0040] (3) Small RNA transcriptome sequencing;

[0041] Small RNA quantification: The small RNA sample used for library construction is first quantified using a library quantification kit. Use 1 μg to generate a sequencing library.

[0042] Adapter sequence: 3' and 5' adapter sequences were ligated, respectively.

[0043] cDNA synthesis: Under the action of MMLV-derived PrimeScript reverse transcriptase (RT), a single-stranded cDNA was reversely synthesized using random primers as a template for the RNA after ligation, followed by double-strand synthesis to form a stable double-stranded structure.

[0044] Library enrichment: PCR amplification (11-12 cycles) was performed using sequencing primers to enrich the library concentration.

[0045] Library purification: According to the length distribution characteristics of small RNA, the target fragments were recovered by gel cutting (6% Novex TBE PAGE gel, 1.0 mm, 10 well).

[0046] Sequencing and data analysis: Qubit 4.0 quantification, mixed according to the data ratio for sequencing; bridge PCR amplification was performed on cBot to generate clusters; Illumina NovaSeq 6000 platform sequencing.

[0047] (4) Bioinformatics analysis part;

[0048] Raw sequence data statistics: Illumina sequencing belongs to the second generation sequencing technology, and a single run can produce billions of reads. Such a large amount of data cannot display the quality of each read. Statistical methods are used to statistically analyze the base distribution and quality fluctuation of each cycle of all sequencing reads, which can intuitively reflect the sequencing quality and library construction quality of the sample from a macroscopic point of view. For each sample, the raw sequencing data is sequenced and related quality is evaluated, including: A / T / G / C base content distribution statistics, base quality distribution statistics, and base error rate distribution statistics.

[0049] Raw sequencing data quality control: The raw sequencing data contains sequencing adapter sequences or low-quality reads. In order to ensure the accuracy of subsequent bioinformatics analysis, the raw sequencing data is first filtered to obtain high-quality sequencing data to ensure the smooth progress of subsequent analysis. The specific steps and order are as follows: 1) Remove the 3' adapter sequence in the reads, remove the reads without inserted fragments due to adapter self-ligation, etc.; 2) Cut the 3' end sequencing low-quality bases (quality value less than 20); 3) Remove reads containing unknown bases N; 4) Remove reads with too short length (<18 nt); 5) Remove reads with too long length (>32 nt); After quality control, the length of clean reads is analyzed, and according to the characteristics of small RNA, reads with length of 18-32 nt are selected as useful reads for subsequent analysis.

[0050] Aligning with reference genome: using Bowtie to align the quality controlled useful reads with the designated reference genome (human genome), and then aligning the reads to piRBase database to obtain the plasma piRNA expression matrix.

[0051] Establishing Lasso Logistic regression model based on plasma piRNA expression matrix;

[0052] After obtaining the plasma piRNA expression matrix, all samples were randomly divided into training set and test set according to the ratio of 50%, 50%, and Lasso Logistic regression model was constructed in the training set, and then ATC, sensitivity, specificity and other indicators were used to evaluate the prediction accuracy of the model in the training set and test set. The software used is the glmnet package of R language program.

[0053] The biggest difference between Lasso Logistic regression model and traditional Logistic regression model is that Lasso Logistic regression model introduces the regularization parameter λ of regression coefficient. By adjusting the value of parameter λ, the regression coefficient of some variables can be equal to 0 (so that the regression coefficients of other piRNAs except the piRNAs shown in Table 1 are equal to 0), which achieves the purpose of variable selection and is conducive to the application and popularization of the model.

[0054] The optimal value of λ is determined according to the method of 20-fold cross-validation in the training set. When the value of λ is taken, the partial likelihood estimation bias of Lasso Logistic regression model is the smallest, as shown in Figure 2, and it is obtained that when the value of λ is taken, the regression coefficients of 783 piRNAs are equal to 0, and the regression coefficients of 12 piRNAs are not equal to 0. The sequences of the 12 piRNAs and their regression coefficients are shown in Table.

[0055] Table 1

[0056] The regression coefficient value of each piRNA expression value represents the change value of the risk score of gastric cancer in the subject when the expression of the piRNA changes by 1 unit. If the regression coefficient is positive, it means that the risk of gastric cancer increases when the expression of the piRNA increases; similarly, if the regression coefficient is negative, it means that the risk of gastric cancer decreases when the expression of the piRNA increases. The mathematical calculation formula of the risk score of gastric cancer is:

[0057] Gastric cancer risk score (Lasso_Logistic_Score) = ∑ (plasma piRNA expression value x regression coefficient).

[0058] After a Lasso Logistic regression model was used to construct a gastric cancer incidence risk prediction model in the training set, the ATC of the model in the training set was 0.989, the sensitivity was 91%, and the specificity was 98%, as shown in Figure 3. The above model was applied to the test set, and the ATC of the model in the test set was 0.96, the sensitivity was 90%, and the specificity was 96%, as shown in Figure 4. The above results show that the method and the constructed model of the application can accurately predict the incidence risk of gastric cancer.

[0059] In summary, based on the plasma piRNA expression profile of Chinese population, the application obtains a plasma piRNA combination for early diagnosis of gastric cancer, and establishes a gastric cancer early diagnosis model based on the plasma piRNA combination, which can accurately predict the incidence risk of gastric cancer, and is helpful to reduce the cost of detection. Meanwhile, the application uses a Lasso Logistic regression model, which greatly reduces the number of variables in the model, and is convenient for application and promotion of the model.

[0060] The above describes one embodiment of the application, and those skilled in the art can understand that various changes, modifications, replacements and supplements can be made to the embodiments, methodologies and models without departing from the principles and purposes of the application, and these changes, modifications, replacements and supplements should also be considered as the protection scope of the application.

Claims

1. A plasma piRNA combination for early diagnosis of gastric cancer, characterized in that: The plasma piRNAs in the plasma piRNA combination include hsa-piR-32885, hsa-piR-3440, hsa-piR-786, hsa-piR-12390, hsa-piR-414, hsa-piR-23197, hsa-piR-32911, hsa-piR-32945, hsa-piR-28060, hsa-piR-7096, hsa-piR-32870 and hsa-piR-30778.

2. Application of the plasma piRNA combination for early diagnosis of gastric cancer according to claim 1 in constructing a model for early diagnosis of gastric cancer.

3. Use according to claim 2, characterized in that: The model for early diagnosis of gastric cancer is a Lasso Logistic regression model.

4. Use according to claim 3, characterized in that: The mathematical expression of the Lasso Logistic regression model is as follows: Gastric cancer incidence risk score = ∑ (plasma piRNA expression value x regression coefficient).

5. Use according to claim 4, characterized in that: The regression coefficients are as follows: The regression coefficient of hsa-piR-32885 is -0.001; The regression coefficient of hsa-piR-3440 is -0.039; The regression coefficient of hsa-piR-786 is -0.007; The regression coefficient of hsa-piR-12390 is -0.001; The regression coefficient of hsa-piR-414 is -0.002; The regression coefficient of hsa-piR-23197 is 0.037; The regression coefficient of hsa-piR-32911 is 0.008; The regression coefficient of hsa-piR-32945 is 0.004; The regression coefficient of hsa-piR-28060 is 0.001; The regression coefficient of hsa-piR-7096 is 0.002; The regression coefficient of hsa-piR-32870 is 0.002; The regression coefficient of hsa-piR-30778 is 0.

026.

6. Use according to claim 4, characterized in that: The Lasso Logistic regression model construction method is as follows: (1) Collect plasma of gastric cancer patients and healthy control population, and extract plasma free piRNA respectively; (2) Obtain plasma piRNA expression profile by piRNA transcriptome sequencing; (3) Randomly divide the gastric cancer patients and healthy control population into training set and test set, establish Lasso Logistic regression model in the training set, and obtain piRNA regression coefficient included in the model; (4) Based on the Lasso Logistic regression model established in the training set, evaluate the prediction accuracy of the model by using ROC curve, sensitivity and specificity index in the test set.

Citation Information

Patent Citations

  • Low-expression piRNA probe for detecting gastric tissue and detection process thereof

    CN101289692A

  • High-expression piRNA probe for detecting gastric tissue and detection process thereof

    CN101289693A

  • COMPOSITIONS AND METHODS OF USING piRNAS IN CANCER DIAGNOSTICS AND THERAPEUTICS

    CN109072240A

  • Short non-coding protein regulatory rnas (sprrnas) and methods of use

    WO2016065349A2