Application of plasma specific piRNA marker in preparation of reagent for detecting gastric cancer

A gastric cancer detection model constructed using plasma-specific piRNA markers hsa-piR-21921, hsa-piR-9725, hsa-piR-19960, and hsa-piR-23621 and the Lasso Logistic regression algorithm solves the problems of invasiveness and missed diagnosis in gastric cancer detection, achieving efficient and accurate early screening and reducing detection costs.

CN121780693APending Publication Date: 2026-04-03THE FIRST AFFILIATED HOSPITAL OF WENZHOU MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-29
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing methods for detecting gastric cancer are highly invasive and prone to missing early-stage cancer, necessitating the development of more effective, accurate, sensitive, and non-invasive detection methods.

Method used

A gastric cancer detection model was constructed using plasma-specific piRNA markers hsa-piR-21921, hsa-piR-9725, hsa-piR-19960, and hsa-piR-23621, combined with the Lasso Logistic regression algorithm. The risk of developing the disease was calculated by detecting the expression level of piRNA in plasma.

Benefits of technology

It achieves high accuracy and high sensitivity in gastric cancer detection, with an area under the curve (AUC) of up to 0.959, a sensitivity of 94.3%, and a specificity of 91.1%. It avoids the pain of invasive examinations, reduces detection costs, and facilitates large-scale population screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121780693A_ABST
    Figure CN121780693A_ABST
Patent Text Reader

Abstract

The invention discloses application of a plasma specific piRNA marker in preparation of a reagent for detecting gastric cancer. The piRNA marker is composed of the following plasma piRNAs: hsa-piR-21921, hsa-piR-9725, hsa-piR-19960, and hsa-piR-23621, and the piRNA marker is composed of the following plasma piRNAs: hsa-piR-21921, hsa-piR-9725, hsa-piR-19960, and hsa-piR-23621. The invention provides the piRNA marker for gastric cancer detection, and the piRNA marker can be used for preparing a reagent for gastric cancer detection and has the advantages of high accuracy, high sensitivity and high specificity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bioinformatics, and in particular to the application of a plasma-specific piRNA marker in the preparation of a reagent for detecting gastric cancer. Background Technology

[0002] Gastric cancer is a malignant neoplastic disease originating from the epithelial cells of the gastric mucosa. In recent years, the incidence of gastric cancer has been relatively high, ranking fifth among malignant tumors in my country. Screening, early diagnosis, and early treatment for high-risk groups can effectively reduce the incidence and mortality of gastric cancer.

[0003] The traditional method for detecting gastric cancer is upper gastrointestinal endoscopic biopsy and pathological examination. Although histopathology is the gold standard for diagnosing gastric cancer, it requires inserting an endoscope into the stomach and taking mucosal tissue samples for pathological evaluation, which is inconvenient. Furthermore, early cancerous lesions are small and atypical in shape, making them prone to being missed during biopsy. Gastroscopy is an invasive procedure, and some patients have poor tolerance and low compliance. Therefore, exploring more effective, accurate, sensitive, and non-invasive clinical biomarkers is urgently needed for the early screening and diagnosis of gastric cancer.

[0004] piRNAs are 25–33 nt long non-coding RNAs closely related to genome stability, epigenetic regulation, germline stem cell differentiation, embryonic development, and the occurrence and development of various diseases. Therefore, piRNAs have the potential to serve as biomarkers. Currently, further research is needed regarding the role of piRNAs in gastric cancer. A better understanding of the role of piRNAs in the development and progression of gastric cancer can help identify novel biomarkers for gastric cancer detection, which is crucial for early screening and diagnosis of gastric cancer patients. Summary of the Invention

[0005] The purpose of this invention is to provide the application of a plasma-specific piRNA marker in the preparation of reagents for detecting gastric cancer. This invention provides a piRNA marker for gastric cancer detection, which has the advantages of high accuracy, high sensitivity, and high specificity in the preparation of reagents for gastric cancer detection.

[0006] The technical solution of the present invention: the application of plasma-specific piRNA markers in the preparation of reagents for detecting gastric cancer, wherein the piRNA markers are composed of the following plasma piRNAs: hsa-piR-21921, hsa-piR-9725, hsa-piR-19960 and hsa-piR-23621.

[0007] For the above applications, the base sequence of hsa-piR-21921 is as follows: TGTGGACTGTTTTCTTTGCCTAGTGAC; The base sequence of hsa-piR-9725 is as follows: TGAATTGAATGAGTTCGGATTGGCCT; The base sequence of hsa-piR-19960 is as follows: TGGTAGTTGTATTGTCGTTTCAGAAAC; The base sequence of hsa-piR-23621 is as follows: CTCAACTGGAAGTTGGTCGCCTGCCCC.

[0008] In the aforementioned application, the reagent is used to detect the expression level of plasma piRNA in human plasma samples.

[0009] In the aforementioned applications, the reagent comprises primers, probes, or chips for the specific detection of the piRNA marker.

[0010] The aforementioned application involves inputting the expression level of plasma piRNA into a gastric cancer detection model constructed based on the Lasso Logistic regression algorithm to calculate a gastric cancer risk score.

[0011] The mathematical expression of the aforementioned gastric cancer detection model is as follows: Gastric cancer risk score = ∑(plasma piRNA expression value × plasma piRNA regression coefficient).

[0012] In the aforementioned application, the regression coefficients are as follows: The regression coefficient of hsa-piR-21921 is -0.02718091; The regression coefficient of hsa-piR-9725 is -0.04306971; The regression coefficient of hsa-piR-19960 is -0.003761725; The regression coefficient of hsa-piR-23621 is 0.0003593632.

[0013] The aforementioned application, and the method for constructing the gastric cancer detection model, is as follows: (1) Collect plasma from gastric cancer patients and healthy controls, and extract cell-free piRNA from the plasma respectively; (2) Obtain plasma piRNA expression profile by piRNA transcriptome sequencing; (3) Randomly divide gastric cancer patients and healthy controls into training set and test set, establish Lasso Logistic regression model in training set, and obtain piRNA regression coefficients included in the model; (4) Based on the Lasso Logistic regression model established in training set, evaluate the predictive accuracy of the model in test set using ROC curve, sensitivity and specificity indicators.

[0014] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention discovers and validates the value of piRNA markers composed of hsa-piR-21921, hsa-piR-9725, hsa-piR-19960, and hsa-piR-23621 in gastric cancer detection. The detection model constructed based on this combination performed excellently on an independent test set, with an area under the curve (AUC) as high as 0.959, a sensitivity of 94.3%, and a specificity of 91.1%, significantly outperforming some existing non-invasive diagnostic methods and effectively distinguishing gastric cancer patients from healthy controls.

[0015] 2. This invention is based on plasma testing, requiring only a small amount of blood (e.g., 200 μL of plasma) to complete the test, avoiding the pain and risks associated with invasive procedures such as endoscopy. Patient compliance is good, facilitating large-scale gastric cancer screening. This invention uses a Lasso Logistic regression model for variable selection, ultimately retaining only the four most representative plasma piRNAs, greatly simplifying the testing panel. This precise selection helps reduce testing costs and is more conducive to the clinical promotion and application of this technology. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of the technical solution of the present invention.

[0017] Figure 2 This is a graph showing the relationship between the regularization parameter λ and the partial likelihood estimation bias in a gastric cancer detection model.

[0018] Figure 3 This is a schematic diagram of the ROC curve of the gastric cancer detection model on the training set.

[0019] Figure 4 This is a schematic diagram of the ROC curve of the gastric cancer detection model on the test set. Detailed Implementation

[0020] The present invention will be further described below with reference to the accompanying drawings and embodiments, but this should not be construed as limiting the present invention.

[0021] Example 1: Construction and validation of a gastric cancer detection model based on plasma piRNA expression profile.

[0022] This embodiment consists of four parts: collecting plasma samples, extracting small RNAs from plasma using the Qiagen miRNeasy plasma kit, small RNA transcriptome sequencing, and establishing a Lasso Logistic regression model based on the plasma piRNA expression matrix. The procedure of this embodiment is as follows: Figure 1 As shown.

[0023] (1) Collecting plasma samples: Whole blood (5 ml) was collected from 200 gastric cancer patients and 100 healthy controls and placed in blood collection tubes containing EDTA anticoagulant. The blood collection tubes were repeatedly inverted to ensure thorough mixing of the EDTA anticoagulant with the blood. The mixed blood collection tubes were then placed in a centrifuge and centrifuged at 3000 rpm and 4°C for 10 minutes. The supernatant after centrifugation was the plasma.

[0024] Use a sterile pipette to draw 2 ml of plasma, transfer it into an EP tube, and store it frozen at -80°C. (2) Extraction of small RNA from plasma using the Qiagen miRNeasy plasma kit: 1. Cell lysis and small RNA extraction: Add 1 ml of QIAzol lysis reagent to 200 μl of sample, vortex or invert to mix thoroughly, and incubate at room temperature (15-25℃) for 5 min. Add 200 μl of chloroform to the mixture, shake vigorously for 15 seconds, and incubate at room temperature for 2-3 min. After incubation, centrifuge at 4℃ and 12000g for 15 min, and transfer the upper aqueous phase to a new EP tube. Add 1.5 times the volume of 100% ethanol to the transferred aqueous phase, and invert to mix thoroughly with the ethanol.

[0025] 2. Small RNA column purification: Take 700 μL of the above mixture and add it to an RNeasy MinElute spin column. Centrifuge at 8000 g for 15 s at room temperature and discard the waste liquid in the collection tube. Repeat the above operation once for the remaining liquid, and pass the remaining mixture through the column.

[0026] Add 700 μl of Buffer RWT to the column, centrifuge at 8000g for 15 seconds, and discard the waste liquid; Add 500 μl of buffer RPE to the column, centrifuge at 8000g for 15s, and discard the waste liquid; Add 500 μl of 80% ethanol to the column, centrifuge at 8000g for 2 min, and discard the waste liquid; Transfer the RNeasy MinElute spin column to a new 2 ml EP tube, open the rotating column cap, centrifuge at maximum speed for 5 min, and discard the residual liquid and collection tube. 3. Small RNA elution: Place the RNeasy MinElute spin column into a new 1.5 ml collection tube, add 15 μL of RNase-free water to the center of the filter membrane, gently cap the tube, and let it stand at room temperature for 2 minutes. Centrifuge at maximum speed for 1 minute, and collect the liquid at the bottom of the tube; this is the extracted small RNA. 4. RNA concentration and RNA integrity assessment: Take 1 μL of RNA sample and perform Aglient 2100 RNA Pico chip detection. The RNA fragment peak is generally below 200 nt.

[0027] The RNA quality standards were set as follows: RIN≥7, >50ng / μL, OD260 / 280 between 1.8 and 2.2. Only RNA samples meeting these standards were used for subsequent sequencing library construction. (3) Small RNA transcriptome sequencing: 1. Small RNA quantification: The qualified small RNA samples were quantified using a library quantification kit, and 1 μg of small RNA was used as the starting template for library construction.

[0028] 2. Connector connection: Specific adapter sequences were ligated to the 3' and 5' ends of the small RNA, respectively. 3. cDNA synthesis: Using the small RNA following the ligation of the adapter as a template, random primers and MMLV-derived PrimeScript reverse transcriptase (RT) were added to synthesize one-stranded cDNA via reverse transcription, followed by two-stranded cDNA synthesis to form a stable double-stranded cDNA structure.

[0029] 4. Rich library: PCR amplification was performed using sequencing primers, with 11-12 amplification cycles set to enrich the library concentration.

[0030] 5. Library purification: Gel electrophoresis was performed using a 6% Novex TBE PAGE gel (1.0 mm thickness, 10 wells). Based on the length distribution characteristics of small RNAs, the target fragments were excised and recovered to obtain purified sequencing libraries. 6. Sequencing and data analysis: The purified libraries were quantified using a Qubit 4.0 fluorescence quantitative analyzer, and multiple libraries were mixed according to a preset data ratio. The mixed library was transferred to cBot for bridge PCR amplification to generate clusters. The amplified sequencing chip was then placed on an Illumina NovaSeq 6000 platform for sequencing.

[0031] (4) Bioinformatics Analysis Section: 1. Statistics of raw sequence data: Illumina next-generation sequencing technology was used to sequence the samples. This technology can generate billions of reads in a single run. Due to the massive amount of data, it is impossible to display the quality of each read individually. Therefore, statistical methods were used to analyze the base distribution and quality fluctuation of each cycle of all sequencing reads, providing a macroscopic and intuitive reflection of the sequencing quality and library construction quality of the samples. Specific statistical indicators for the raw sequencing data of each sample included: A / T / G / C base content distribution statistics, base quality distribution statistics, and base error rate distribution statistics.

[0032] 2. Quality control of raw sequencing data: To ensure the accuracy of subsequent bioinformatics analysis, the raw sequencing data is filtered to obtain high-quality sequencing data. The specific steps and order are as follows: 1) Remove the 3' connector sequence from the reads, and also remove the reads that have no inserted segments due to connector self-connection or other reasons; 2) Cut the 3' end to sequence low-quality bases (quality value less than 20); 3) Remove reads containing the unknown base N; 4) Remove reads that are too short (<18 nt); 5) Remove reads that are too long (>32 nt); After quality control, the length of clean reads was analyzed. Based on the characteristics of small RNAs, reads with a length of 18-32 nt were selected as useful reads for subsequent analysis.

[0033] 3. Alignment with reference genome: Bowtie was used to align the quality-controlled useful reads with a specified human reference genome, and then the reads were aligned to the piRBase database to obtain the plasma piRNA expression matrix.

[0034] 4. Gastric cancer detection model based on Lasso Logistic regression algorithm: Based on the obtained plasma piRNA expression matrix, a gastric cancer detection model was constructed. The specific process is as follows: 1) Randomly divide all samples into training and test sets at a 50% to 50% ratio; 2) A gastric cancer detection model based on the Lasso Logistic regression algorithm was constructed using the glmnet package in the R language program on the training set; 3) The predictive accuracy of the model is evaluated using metrics such as AUC, sensitivity, and specificity in both the training and test sets.

[0035] Compared with the traditional Logistic regression model, the gastric cancer detection model introduces a regularization parameter λ for the regression coefficients. By adjusting the value of parameter λ, the regression coefficients of certain variables can be made equal to 0 (the regression coefficients of other piRNAs besides those shown in Table 1 can be made equal to 0), thus enabling variable screening and facilitating the application and promotion of the model.

[0036] The optimal value of λ was determined using 20-fold cross-validation on the training set. This value of λ minimizes the partial likelihood estimation bias of the Lasso Logistic Regression model. (See...) Figure 2 The results showed that when λ was set to a certain value, the regression coefficients of 463 piRNAs were equal to 0, while the regression coefficients of 4 piRNAs were not equal to 0. The sequences of these 4 piRNAs and their regression coefficients are shown in the table.

[0037] Table 1

[0038] The regression coefficient for each piRNA expression value represents the change in the subject's gastric cancer risk score for every 1 unit change in piRNA expression. A positive regression coefficient indicates an increased risk of gastric cancer associated with elevated piRNA expression; similarly, a negative regression coefficient indicates a decreased risk of gastric cancer associated with elevated piRNA expression. The mathematical formula for calculating the gastric cancer risk score is as follows: Lasso_Logistic_Score = ∑(plasma piRNA expression value × plasma piRNA regression coefficient).

[0039] After constructing a gastric cancer detection model using the Lasso Logistic Regression algorithm on the training set, the model achieved an AUC of 0.967, a sensitivity of 92.0%, and a specificity of 93.7% on the training set. Figure 3 As shown. Applying the above model to the test set, the model achieved an AUC of 0.959, a sensitivity of 94.3%, and a specificity of 91.1% on the test set. Figure 4As shown above, the results demonstrate that the method and model of this invention can accurately predict the risk of gastric cancer.

[0040] In summary, this invention, based on the plasma piRNA expression profile of the Chinese population, obtained a plasma piRNA combination for gastric cancer detection and established a gastric cancer detection model based on the plasma piRNA combination. This model can predict the risk of gastric cancer relatively accurately and helps to reduce the cost of detection. At the same time, this invention uses the Lasso Logistic regression algorithm, which significantly reduces the number of variables included in the model, making it easier to apply and promote the model.

[0041] Example 2: Preparation of a gastric cancer detection reagent containing specific primers, probes, and a chip I. Design and Preparation of Specific Primers Based on the base sequences of hsa-piR-21921, hsa-piR-9725, hsa-piR-19960, and hsa-piR-23621, specific upstream and downstream primers were designed following the principles of piRNA primer design. These piRNA-specific upstream and downstream primers can be directly designed by those skilled in the art based on the complete base sequences of the four piRNAs disclosed in the specification, combined with conventional knowledge and general rules for quantitative PCR primer design.

[0042] Primers with specific upstream and downstream sequences were synthesized using a nucleic acid synthesizer and purified by PAGE gel electrophoresis to a purity of ≥98%. The primers were dissolved in RNase-free water to prepare a 100 μM stock solution, which was then aliquoted and stored at -20°C. Before use, the solution was diluted to a 10 μM working solution.

[0043] II. Design and Preparation of Specific Probes Fluorescent probes were designed targeting the core conserved sequences of four piRNAs. The 5' end was labeled with a FAM fluorescent group, and the 3' end with a BHQ1 quencher group to ensure no cross-binding between the probe and primers. After chemical synthesis, the probes were purified by reversed-phase high-performance liquid chromatography (HPLC) to a purity ≥99%. The probes were dissolved in TE buffer to prepare a 50 μM stock solution, which was stored at -20℃ protected from light. Before use, the stock solution was diluted to 5 μM.

[0044] III. Fabrication of Specific Chips Aldehyde-modified glass chips were selected, rinsed with ultrapure water, dried under nitrogen, and then cross-linked in a UV cross-linker for 30 min to activate the surface aldehyde groups. The four piRNA-specific probes (50 μM) were mixed with spotting buffer at a 1:1 ratio, adjusting the final concentration to 25 μM. The probe mixture was spotted onto the treated glass chips using a chip spotting instrument, with three replicates for each probe, a spot diameter of 100 μm, and a spot spacing of 300 μm. After spotting, the chips were placed in an incubator at 70% humidity for 2 h to allow the probes to be immobilized on the chip surface through covalent binding of the aldehyde and amino groups. Unbound active sites were then blocked with blocking buffer, rinsed with ultrapure water, dried under nitrogen, and stored in a sealed container at 4°C.

[0045] The embodiments of the present invention have been described above. Those skilled in the art will understand that various changes, modifications, substitutions, and additions can be made to these embodiments, methodologies, and models without departing from the principles and spirit of the present invention, and these changes, modifications, substitutions, and additions should also be considered within the scope of protection of the present invention.

Claims

1. The application of plasma-specific piRNA markers in the preparation of reagents for detecting gastric cancer, characterized in that: The piRNA markers consist of the following plasma piRNAs: hsa-piR-21921, hsa-piR-9725, hsa-piR-19960, and hsa-piR-23621.

2. The application according to claim 1, characterized in that: The base sequence of hsa-piR-21921 is as follows: TGTGGACTGTTTTCTTTGCCTAGTGAC; The base sequence of hsa-piR-9725 is as follows: TGAATTGAATGAGTTCGGATTGGCCT; The base sequence of hsa-piR-19960 is as follows: TGGTAGTTGTATTGTCGTTTCAGAAAC; The base sequence of hsa-piR-23621 is as follows: CTCAACTGGAAGTTGGTCGCCTGCCCC.

3. The application according to claim 1, characterized in that: The reagent is used to detect the expression level of plasma piRNA in human plasma samples.

4. The application according to claim 1, characterized in that: The reagent contains primers, probes, or chips for the specific detection of the piRNA marker.

5. The application according to claim 3, characterized in that: The application involves inputting the expression level of plasma piRNA into a gastric cancer detection model constructed based on the Lasso Logistic regression algorithm to calculate a gastric cancer risk score.

6. The application according to claim 5, characterized in that: The mathematical expression for the gastric cancer detection model is as follows: Gastric cancer risk score = ∑(plasma piRNA expression value × plasma piRNA regression coefficient).

7. The application according to claim 6, characterized in that: The regression coefficients are shown below: The regression coefficient of hsa-piR-21921 is -0.02718091; The regression coefficient of hsa-piR-9725 is -0.04306971; The regression coefficient of hsa-piR-19960 is -0.003761725; The regression coefficient of hsa-piR-23621 is 0.0003593632.

8. The application according to claim 5, characterized in that: The method for constructing the gastric cancer detection model is as follows: (1) Collect plasma from gastric cancer patients and healthy controls, and extract cell-free piRNA from the plasma respectively; (2) Obtain plasma piRNA expression profile by piRNA transcriptome sequencing; (3) Randomly divide gastric cancer patients and healthy controls into training set and test set, establish Lasso Logistic regression model in training set, and obtain piRNA regression coefficients included in the model; (4) Based on the Lasso Logistic regression model established in training set, evaluate the predictive accuracy of the model in test set using ROC curve, sensitivity and specificity indicators.